Compare commits

..
698 Commits
Author SHA1 Message Date
Ryan Houdek 37f1e55ed5 Docs: Update for release FEX-2204 2022-04-19 01:19:00 -07:00
Ryan Houdek 8ad14728f6 Merge pull request #1644 from Sonicadvance1/ldiv_minor_opt
JITArm64: Get long divide out of the hot path
2022-04-01 18:23:30 -07:00
Ryan Houdek 6b3cd3d31d Merge pull request #1645 from Sonicadvance1/update_aarch64_fit
Scripts: Updates AArch64 fit for Clang 14
2022-04-01 18:23:13 -07:00
Ryan Houdek fba698cb74 Scripts: Updates AArch64 fit for Clang 14
Clang now supports these latest ARMv9 CPUs
2022-04-01 18:08:02 -07:00
Ryan Houdek 0946b123bb JITArm64: Get long divide out of the hot path
For 128-bit divides, we can very quickly check at runtime if we can
avoid the long divide and just do a 64-bit divide.

For unsigned just check if the top bits are all zero.
For signed just check if the top bits match bit 63 of the lower bits.

Additionally, keep the long divide handlers inside of the dispatcher.
This keeps the majority of the code bloat out of the code block itself,
significantly reducing block size for something doing these divides.
Also a fairly large icache improvement from this.

Hard performance number improvements here are hard to get since it
heavily depends on the application, also only occurs on x86-64.

Seems to have helped FTL and Dead Cells performance quite a bit though.
2022-03-31 09:33:19 -07:00
Ryan Houdek b43937a7a1 Merge pull request #1643 from Sonicadvance1/fix_termux
SignalDelegator: Adds missing include
2022-03-29 21:05:28 -07:00
Ryan Houdek 4564eba20d SignalDelegator: Adds missing include
Fixes Termux building.
Fixes #1642
2022-03-29 20:47:22 -07:00
Ryan Houdek 5cc0c0a3da Merge pull request #1641 from philpax/docs-remove-stale-text
docs: Remove stale text
2022-03-29 02:56:20 -07:00
Philpax f8e7c75f86 docs: Remove stale text 2022-03-29 11:27:43 +02:00
Ryan Houdek 042cd354dc Merge pull request #1633 from Sonicadvance1/disable_instructions_on_host_missing
OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
2022-03-23 13:48:04 -07:00
Ryan Houdek 977bda97b2 Merge pull request #1635 from Sonicadvance1/4000_0001h
CPUID: Adds 4000_0001h function
2022-03-23 13:42:45 -07:00
Ryan Houdek 4cf48ca9bb CPUID: Adds 4000_0001h function
Exposes the host architecture through this CPUID function. Only exposes
the architectures we support. Not burning 16-bits on using ELF machine
definitions here.

Uses 4 bits still for future expansion.
2022-03-22 16:53:44 -07:00
Mai M a247df50ea Merge pull request #1624 from Sonicadvance1/cleanup_ir_after_use
FEXCore: Delete IR after it is used
2022-03-22 13:23:05 -04:00
Mai M 1f1c214944 Merge pull request #1634 from Sonicadvance1/cpuid_documentation
Documentation: Adds hypervisor CPUID information
2022-03-22 12:57:15 -04:00
Ryan Houdek d16db4ebde OpcodeDispatcher: Fixes SIGILL on unsupported host instructions
If the host doesn't support the instructions required for implementing
an instruction then don't even add them to the opcodedispatcher.

This means that we will never try emitting instructions that the host
doesn't support (For these instructions anyway) and successfully passes
the guest SIGILL for these particular instructions.

Fixes #1631
2022-03-21 23:03:04 -07:00
Ryan Houdek ae1c563082 Documentation: Adds hypervisor CPUID information
Currently we only implement function 4000_0000h. This will expand in the
future but this is all we have right now.
2022-03-21 22:46:48 -07:00
Ryan Houdek ebd0edbab7 Merge pull request #1632 from FEX-Emu/skmp/flush-test-harness
TestHarnessRunner: Flush log on asserts
2022-03-21 12:50:43 -07:00
Stefanos Kornilios Misis Poiitidis e87e9d269a TestHarnessRunner: Flush log on asserts 2022-03-21 21:32:35 +02:00
Ryan Houdek 187c64182b Merge pull request #1628 from Sonicadvance1/fix_finit
OpcodeDispatcher: Fixes FNINIT
2022-03-17 20:37:57 -07:00
Ryan Houdek 60c7ea6e5f Merge pull request #1620 from Sonicadvance1/fix_1618
FEXCore: Fixes #1618
2022-03-17 20:36:06 -07:00
Ryan Houdek 6f1b4b0eee OpcodeDispatcher: Fixes FNINIT
Was incorrectly setting the FCW to 037h when it was supposed to be
037Fh.

Fixes a bug in a visual novel where its CPUID state wouldn't initialize
if this was set incorrectly.
2022-03-17 20:27:22 -07:00
Ryan Houdek fb69300397 FEXCore: Delete IR after it is used
For the JIT cores we don't need to keep IR around, it's only necessary
for the Interpreter. So once the AOT IR service is done dealing with the
IR, check to see if we can delete it.

This causes teeworld's title screen memory usage to go from 730MB to
566MB. 77.5% the memory usage there.

This is effectively an infinite memory leak if the codespace wasn't ever
overwritten or invalidated. So larger memory usage programs would end up
having a larger impact.
2022-03-13 19:01:40 -07:00
Ryan Houdek 5677924525 Merge pull request #1621 from Sonicadvance1/fix_1584
Softfloat: Fixes FSCALE
2022-03-13 18:57:06 -07:00
Ryan Houdek 8422fc632d Merge pull request #1623 from Sonicadvance1/remove_unused_debug_data
FEXCore: Removes unused debug data
2022-03-13 18:56:50 -07:00
Ryan Houdek d33cd744fb FEXCore: Removes unused debug data
This isn't used anywhere. Just remove these.
If we get the imgui debugger running again then we can add even more
stats to sort block costs by.
2022-03-13 18:40:49 -07:00
Ryan Houdek 3b0fb27ae9 Softfloat: Fixes FSCALE
I misread the implementation details of this instruction when
implementing.

The pseudocode says `ST(0) = ST(0) ∗ 2^rndint(ST(1))` so I understood
the instruction to use the current rounding mode of the host to extract
the integer portion of `ST(1)`.

The actual implementation is in the details of the statement `the
integer portion of the floating- point value in ST(1).`

This behaves like round towards zero/truncate, additional hardware
testing and documentation reading confirms this.

Fixes #1584
2022-03-13 14:11:31 -07:00
Ryan Houdek 4603e09a04 FEXCore: Fixes #1618 2022-03-13 13:42:37 -07:00
Ryan Houdek 7b0265ffe2 Merge pull request #1617 from Sonicadvance1/gdbstub_improvements4
GDBServer improvements: Three's a crowd
2022-03-13 13:24:41 -07:00
Ryan Houdek ec54560a38 GDBServer: Fixes memory reading
memory-map is not something we want to use. Adds a comment about it and
disables it.

Also changes core events to wait for an event from GDBStub for waking up
which fixes a hang.
2022-03-13 13:05:00 -07:00
Ryan Houdek 6a5abd3672 Merge pull request #1616 from Sonicadvance1/gdbstub_improvements3
Gdbstub improvements: The sequel
2022-03-13 13:03:25 -07:00
Ryan Houdek 3e6af39c42 GDBServer: Zero initialize some variables to fix connection stability
Otherwise you always had to attempt connecting twice in a row
2022-03-13 12:50:01 -07:00
Ryan Houdek cf82ffc052 GDBServer: Let gdb know when the library map has updated
We need to fetch the full list of map files from the memory map and hand
it over to gdb.
It will then fetch all the libraries from the remote host and give us
backtraces
2022-03-13 12:50:01 -07:00
Ryan Houdek b190150281 FEXCore: Merges redundant string trimming implementations 2022-03-13 12:50:01 -07:00
Ryan Houdek 53ffe5df43 Merge pull request #1613 from Sonicadvance1/gdbstub_improvements2
GDBServer improvements
2022-03-13 12:49:17 -07:00
Ryan Houdek d39df8d3ed Merge pull request #1614 from Sonicadvance1/add_comment
JIT: Adds comment to EmitDetectionString
2022-03-10 14:53:41 -08:00
Ryan Houdek eeb2b928b9 JIT: Adds comment to EmitDetectionString 2022-03-10 14:28:45 -08:00
Ryan Houdek 6cf24a748f GdbServer: Document what PassSignals is for 2022-03-10 14:23:21 -08:00
Ryan Houdek 3b4fd180de GDBServer: Pass auxv better
Fixes 32-bit auxv as well.
2022-03-10 14:21:46 -08:00
Ryan Houdek 7300c7a853 SignalDelegator: Remove anti-pattern usage 2022-03-10 14:00:51 -08:00
Ryan Houdek 376f6db3ac GDBServer: Reformat code to two space tabs
No functional change
2022-03-10 14:00:51 -08:00
Ryan Houdek ad3a960717 GDBServer: Support sending gdb the correct signal
Instead of just sending SIGSEGV, pass the real signal
2022-03-10 13:57:17 -08:00
Ryan Houdek 0a1ef867ee SignalDelegator: Support multiple backend host handlers
This will be necessary for gdbserver to handle signals indepedentally of
the FEX handling.
2022-03-10 13:57:17 -08:00
Ryan Houdek a7fe69deea CPU: Stop trying to initialize signal handlers per thread
These are static per process and only need to be initialized once.
We are going to support multiple signal handlers from the backend after
this, so can only install once.
2022-03-10 13:57:17 -08:00
Ryan Houdek a8c3b6d46f GDBServer: Capture signal capture numbers
This will allow us to wire this to a future signal handler for gdbserver
2022-03-10 13:57:17 -08:00
Ryan Houdek 60db2655bb GDBServer: Encode the return to pread correctly
This encodes the resulting data as raw binary rather than any special
escaped encoding
2022-03-10 13:57:10 -08:00
Ryan Houdek fd717b6995 GDBServer: Expose program offsets better
We were incorrectly returning programing offsets
Get the program offset from the frontend so we can know what to give gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek ba37388fe3 GDBServer: Expose auxv values
We already expose these in the code loader, pump it through gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 68f32d85c9 CodeLoader: Expose base ELF loaded offset
Useful for gdb
2022-03-09 19:07:45 -08:00
Ryan Houdek 2aa77e85de LinuxSyscalls: Expose CodeLoader through syscall interface 2022-03-09 19:07:45 -08:00
Ryan Houdek 91665fdf0e Merge pull request #1610 from wannacu/main
FileManager: Fix realpath failed on debian buster
2022-03-09 18:58:52 -08:00
Ryan Houdek 23a1c64bf7 Merge pull request #1612 from Sonicadvance1/tag_memory_allocations
JITs: Emit identification string in the code buffers
2022-03-09 18:52:46 -08:00
Ryan Houdek c2dcf06632 JITs: Emit identification string in the code buffers
At the start of each code buffer, emit a small string for letting memory
inspection know if a code region is for the JITs.
2022-03-09 18:29:37 -08:00
Ryan Houdek fad91bb818 Merge pull request #1609 from Sonicadvance1/fix_map_32bit
Linux: Fixes MAP_32BIT supported range
2022-03-08 17:33:34 -08:00
wannacu 0e769ece26 docs: Update Readme_CN.md 2022-03-08 18:10:08 +08:00
wannacu 9a780b40a2 Docs: Add Chinese README 2022-03-08 18:04:09 +08:00
wannacu 898873e9e3 FileManager: Fix realpath failed on debian buster
This happend on debian buster when run realpath(i386) on arm64 host.
2022-03-08 16:42:05 +08:00
Ryan Houdek 52292e5f7e Linux: Fixes MAP_32BIT supported range
I accidentally committed a 32-bit range that was significantly smaller
than what it should be.
While the minimal range worked for simple cases, it didn't work for
anything complex.
Give it the full range it needs.

Fixes #1600
2022-03-06 17:46:17 -08:00
Ryan Houdek 5de6c866b7 Merge pull request #1608 from Sonicadvance1/termux_build_option
Adds a cmake option for forcing a termux build
2022-03-06 13:54:20 -08:00
Ryan Houdek ec0cd3aec4 Adds a cmake option for forcing a termux build
This is necessary when cross-compiling rather than building on-device
2022-03-06 12:56:25 -08:00
Mai M f5f9512d9a Merge pull request #1606 from Sonicadvance1/fhu_page_size
Change page define usages over to self-defined
2022-03-06 15:53:15 -05:00
Mai M fb27cb4356 Merge pull request #1607 from Sonicadvance1/disable_guis_termux
Disables GUI applications in a Termux build
2022-03-06 15:52:37 -05:00
Mai M 94664580c8 Merge pull request #1605 from Sonicadvance1/update_docs_termux
Update ReleaseProcess docs for Termux
2022-03-06 15:52:10 -05:00
Ryan Houdek 99a93fa9ea Disables GUI applications in a Termux build 2022-03-06 08:09:55 -08:00
Ryan Houdek 4cb6918506 Change page define usages over to self-defined
In the case of an AArch64 builder is using 16kb or 64kb pages like is
common on servers then it would fail to compile, even if the resulting
application would only ever run on 4k page hosts.

Resolve this by removing the build check and hardcoding 4kb pages for
each of our uses. We still require 4kb pages to run, so this mostly just
removes the weirdness where it is 16kb builder + 4k runner. Would have
broken some of our assumptions when running.
2022-03-06 07:33:10 -08:00
Ryan Houdek 9cc743bf84 Update ReleaseProcess docs for Termux
FEX hardly works on Termux as-is, but we should make sure to document
how to update the packages otherwise we will quickly become outdated on
their package management.
2022-03-06 05:59:33 -08:00
Ryan Houdek a408749eef Docs: Update for release FEX-2203 2022-03-06 04:49:31 -08:00
Ryan Houdek d8a3687ac3 Merge pull request #1604 from Sonicadvance1/fix_cmpxchg_66h
OpcodeDispatcher: Fixes CMPXCHG8B/16B with 66h/72h/73h prefix
2022-03-06 04:10:58 -08:00
Ryan Houdek a32c7f6ce2 unittests: Adds cmpxchg unit tests for prefixes 2022-03-06 03:54:45 -08:00
Ryan Houdek 540feb857b OpcodeDispatcher: Fixes CMPXCHG8B/16B with 66h/72h/73h prefix
The documentation is incorrect about this instruction. It claims that
you use 66h prefix to choose between operating at 8B or 16B.
This is incorrect, real hardware only responds to REX.W for choosing the
operating size. These other prefixes are ignored but is still accepted as
an instruction decoding.
2022-03-06 03:54:45 -08:00
Ryan Houdek a3902a0d2d FEXCore: Fixes usage of GPRPair in operations
These were working around the previous quirks by accident
2022-03-06 03:54:45 -08:00
Ryan Houdek 5fbd01536f IR: Fixes some GPRPair IR op definitions
These were always wrong but how it the operations were handled meant
that it happened to work even though the IR representation was broken
2022-03-06 03:54:45 -08:00
Ryan Houdek bebcab0277 Merge pull request #1603 from Sonicadvance1/rng_support
FEXCore: Adds support for RDRAND/RDSEED
2022-03-06 03:54:20 -08:00
Ryan Houdek d0f17d400e unittests: Adds RDRAND/RDSEED unit tests 2022-03-06 03:40:25 -08:00
Ryan Houdek 11a07eb3f2 FEXCore: Adds support for RDRAND/RDSEED
This matches the AArch64 implementation fairly well.
Bundles RDRAND and RDSEED together for simplification, both instructions
are a single flag on AArch64.
2022-03-06 03:40:25 -08:00
Ryan Houdek cb13e1bdb9 X86Tables: Fixes secondary group decoding
If we're hitting these group tables then it needs the ignore overlay extension
since the prefixes are used to select ops inside this table.
2022-03-06 03:40:25 -08:00
Ryan Houdek 3e8c6d0be0 Merge pull request #1601 from Sonicadvance1/new_ir_json
IR: New IR JSON format
2022-03-06 03:40:05 -08:00
Ryan Houdek 847b6e5026 LinuxSyscalls: Fixes struct verifier on Ubuntu 20.04
The `linux/types.h` header needs to be included before the rest
otherwise we are missing types.
2022-03-04 17:44:26 -08:00
Ryan Houdek 7f6e9d3dae unittests/IR: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek e244e8f66b x86Jit: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 0adaa85d97 ArmJit: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek f1d34c7407 Interpreter: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 7a9492ceca IRPasses: Resolves fallout from recent JSON changes 2022-03-04 16:39:11 -08:00
Ryan Houdek 39192cd092 OpcodeDispatcher: Resolves fallout from recent JSON changes
Also fixes a bug in FCVTIntTo where it was getting passed an FPR when it
wants a GPR. Would cause RA problems on the JITs
2022-03-04 05:10:45 -08:00
Ryan Houdek 73abf9b9dd IRParser: Resolves fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 0387c24e21 IREmitter: Resolves the fallout from recent JSON changes 2022-03-04 04:07:25 -08:00
Ryan Houdek 9d14bc7846 IR: Updates generators and json to new format
This greatly simplifies the IR format by using string parsing for
gathering the information.
Tons of redundant information removed.
Significantly more difficult to mess up adding a new IR op.

Significantly improves the generator functions in IREmitter
2022-03-04 03:51:44 -08:00
Ryan Houdek ac32ecadbe Merge pull request #1597 from Sonicadvance1/3dnow_and_back_again
OpcodeDispatcher: Implements all the 3DNow! instructions
2022-03-04 03:36:40 -08:00
Mai M 08938ecb11 Merge pull request #1596 from Sonicadvance1/fix_old_kernel_bug
LinuxAllocator: Fixes bug with old kernels and hint allocation
2022-03-02 00:26:51 -05:00
Ryan Houdek 4a3cbf1b1f Merge pull request #1598 from Sonicadvance1/hypervisor_cpuid
CPUID: Implements leaf 4000_0000
2022-03-01 21:02:32 -08:00
Ryan Houdek 0a0bc21c88 CPUID: Implements leaf 4000_0000
This region is reserved for hypervisor uses. Let's follow other examples
and return a hypervisor vendor id signature as another way for software
to find if it is running under FEX-Emu.
2022-03-01 20:47:03 -08:00
Mai M a7ad7f4456 Merge pull request #1593 from Sonicadvance1/add_robin_map
Adds tsl::robin_map
2022-03-01 23:24:34 -05:00
Mai M 46ac05e3cf Merge pull request #1599 from Sonicadvance1/deprecated_distutils
Scripts: Stop using deprecated Distutils
2022-03-01 23:23:55 -05:00
Ryan Houdek 9911fe68d4 Scripts: Stop using deprecated Distutils
According to PEP 386: https://www.python.org/dev/peps/pep-0386/

distutils is deprecated and will be removed in an upcoming python
version.

Switch over to pkg_resources for version parsing and comparison
2022-03-01 07:10:41 -08:00
Ryan Houdek 57a56545b2 Merge pull request #1589 from Sonicadvance1/fix_vixl_assert
Update vixl to fix assert
2022-03-01 05:35:31 -08:00
Ryan Houdek b4e0565907 Merge pull request #1590 from Sonicadvance1/termux_fixes
Termux fixes
2022-03-01 04:08:48 -08:00
Ryan Houdek 5137af5bae Fixes epoxy include in FEXConfig and FEXLogServer 2022-03-01 03:55:19 -08:00
Ryan Houdek 385ed2c2ec Update imgui to fix autodetect 2022-03-01 03:55:16 -08:00
Ryan Houdek 3418cc8054 LinuxSyscalls: More type fixes 2022-03-01 03:55:15 -08:00
Ryan Houdek 1f835ea1f4 LinuxSyscalls: Fixes semid and ipc types
Newer headers redefine semid_ds and ipc_perm as semid64_ds and
ipc64_perm silently.
Use the new types directly since we are a 64-bit only application.
2022-03-01 03:55:12 -08:00
Ryan Houdek 0a40753624 More missing include fixes 2022-03-01 03:55:10 -08:00
Ryan Houdek 609538e758 Utils/Allocator: Use a namespace alias for pmr
This still lives under experimental in Termux environment
2022-03-01 03:55:08 -08:00
Ryan Houdek 09010e1292 LinuxSyscalls: Remove unused headers now
These don't even exist on termux.
2022-03-01 03:54:01 -08:00
Ryan Houdek bddf3871ba LinuxSyscalls/x32/Types: Fixes stat type definitions
The time argument definitions are defines in termux.
Rename our definition name of these so we don't get caught by define.
2022-03-01 03:50:18 -08:00
Ryan Houdek 1ca9e56502 LinuxSyscalls/x64/Types: Fix guest_stat definition
kernel types don't exist in termux. Use uint64_t and int64_t directly.

Also using reserved `__` causes compile failure.
2022-03-01 03:50:17 -08:00
Ryan Houdek fe58a9ae2c LinuxSyscalls/Thread: Don't use set_robust_list on Termux
Would get caught by seccomp and crash FEX
2022-03-01 03:50:15 -08:00
Ryan Houdek bc7c4dfa74 LinuxSyscalls/x32/Types: Don't redefine SIGEV defines 2022-03-01 03:50:12 -08:00
Ryan Houdek 7af7b1309d Msg: Switch msqd_t to FEX defined type 2022-03-01 03:50:10 -08:00
Ryan Houdek dd4630750f LinuxSyscalls/Types: Adds missing types for Termux 2022-03-01 03:50:07 -08:00
Ryan Houdek 302da029eb Work around Termux not supporting hardlinks
The Android filesystem they are on just doesn't support them
Instead of hardlinking FEXLoader to FEXInterpter, just build the
executable twice and eat the filesystem cost.
2022-03-01 03:50:05 -08:00
Ryan Houdek 3cc59bc68e LinuxSyscalls: Switches to a bunch of raw syscalls
For older and Termux build environments these helper libc functions
don't exist.
2022-03-01 03:49:31 -08:00
Ryan Houdek 30c27851ef FEXRootFSFetcher: Termux build environments 2022-03-01 03:30:47 -08:00
Ryan Houdek 9b29ff61e9 Adds some missing headers 2022-03-01 03:30:45 -08:00
Ryan Houdek 91f780223b Stop self-defining PAGE_SIZE
We only work on targets with 4096 byte page sizes.
Adds a cmake compile test to ensure this is adhered to.
2022-03-01 03:30:43 -08:00
Ryan Houdek 2933a00b12 Merge pull request #1579 from Sonicadvance1/optimize_syscalls_with_flags
Allow classifying syscalls with flags
2022-03-01 03:23:43 -08:00
Ryan Houdek a0efd2b01f Resolve comments. 2022-02-28 21:04:03 -08:00
Ryan Houdek 3e9dbda146 Classify syscalls 2022-02-28 21:04:03 -08:00
Ryan Houdek 6501715a3c Allow classifying syscalls with flags
In some cases we can generate more optimal code if we have more
information about a syscall which number gets const-propagated.

In particular optimizing through syscalls, not synchronizing state, and
never returning.

- Noreturn is used by a syscall that never returns, like exit.

This means that it never needs to try and synchronize state coming back

- Not synchronizing state and optimizing through syscalls

Useful for syscalls that don't read the state past arguments and only
returns a value.
2022-02-28 21:03:54 -08:00
Ryan Houdek 2af23d9bec Merge pull request #1591 from Sonicadvance1/new_cpus_in_native_fit
Scripts: Updates CPU fitting script for latest CPUs
2022-02-28 07:05:45 -08:00
Ryan Houdek 3ba2d6cfb4 Merge pull request #1586 from Sonicadvance1/testharness_env
TestHarnessRunner: Wire up environment variable option setting
2022-02-28 07:05:28 -08:00
Ryan Houdek abd266441c unittests: Implements 3DNow! unit tests
Covers the full space, of which there aren't many.

3DNow! unit tests are disabled on the CI runner since the x86 CPU in CI
doesn't support it.
2022-02-28 04:06:14 -08:00
Ryan Houdek b1b6078518 CPUID: Enables 3DNow! + Extensions
Now that we support these
2022-02-28 04:05:18 -08:00
Ryan Houdek a9a89bea9f OpcodeDispatcher: Implements all the 3DNow! instructions
This picks up all the instruction implementations, including 3DNow!
Extended and the Geode specific instructions that were added.

Most of these match preexisting SSE instructions except that they
operate at 64-bit and in the MMX registers.
2022-02-28 04:03:55 -08:00
Ryan Houdek 9b1b2e6496 X86Tables: Fills out 3DNow tables
Fully decoded the same way and adds the Geode specific instruction
decodings as well.
2022-02-28 04:02:57 -08:00
Ryan Houdek 638da92f45 Frontend: Fixes minor bug decoding 3DNow!
We already decoded the modrm `rm` bits, check while decoding modrm to
ensure we don't try decoding it again.
Was causing double decoding of SIB and displacement bytes, breaking
things
2022-02-28 04:01:32 -08:00
Ryan Houdek 0f3a169cc4 Opdispatcher: Minor bug fix with unimplemented op
If multiblock isn't enabled then on Unimplemented op we shouldn't create
a new block.

Was causing IR validation to get angry
2022-02-28 04:00:41 -08:00
Ryan Houdek c7dd176799 IR: Adds a VRev64 op
This directly matches the AArch64 instruction and will be used shortly
2022-02-28 04:00:13 -08:00
Ryan Houdek 23daf4cb72 LinuxAllocator: Fixes bug with old kernels and hint allocation
In the face of an application using MAP_FIXED_NOREPLACE AND the host
linux kernel doesn't understand this flag. Then we were falling down the
hint allocation path which would allocate a pointer in 64-bit space,
returning this pointer to a 32-bit userspace and breaking things.

Now when the hint fails with this flag, we know that it intersecting a
range and can early exit.
2022-02-26 21:12:44 -08:00
Ryan Houdek bb7fa84fb8 Scripts: Updates CPU fitting script for latest CPUs
Clang-13 doesn't yet understand the latest ARM CPUs so just document them
and set to the closest thing.
2022-02-26 02:14:52 -08:00
Ryan Houdek 3325ba52b9 Adds tsl::robin_map
This will be used with the code serialization service soon
2022-02-26 00:43:43 -08:00
Ryan Houdek 7f47fe6d73 Update vixl to fix assert
Any hardware using MTE will assert without this
2022-02-24 13:46:56 -08:00
Ryan Houdek ee165379c5 TestHarnessRunner: Wire up environment variable option setting
Wire up the environment variable option setting so asm files can set
these and it works
2022-02-24 13:37:39 -08:00
Ryan Houdek fa554d3096 Merge pull request #1588 from Azkali/main
Improve compatibility with older uapi kernel headers
2022-02-24 01:34:24 -08:00
The Great Wizard Azkali d60710d3f1 Define proper statx syscall depending on CPU architecture 2022-02-24 10:20:28 +01:00
Azkali 75988b2ae5 Improve compatibility with older uapi kernel headers
Following up the work previously done in 2079f6b3c7.
Adding more defines for older Linux uapi headers missing some defines.
2022-02-24 09:49:56 +01:00
Mai M 30803c66f7 Merge pull request #1587 from Sonicadvance1/fix_missing_telemetry_names
Telemetry: Fix missing telemetry names
2022-02-22 21:46:00 -05:00
Ryan Houdek a9d838fa27 Telemetry: Fix missing telemetry names
Didn't have names for tearing
2022-02-22 18:30:40 -08:00
Ryan Houdek 2b8f60c108 TestHarness: Support for asm files having the option to set config options
Allows some something like the following:
"Env": {
  "FEX_MAXINST": "500"
}

Not that I would recommend overriding MAXINST in the asm tests, as
command line overrides that
2022-02-21 14:53:27 -08:00
Mai M 5ec6ee5b69 Merge pull request #1578 from Sonicadvance1/update_vixl
Updates vixl for new cursor updating methods
2022-02-17 16:18:26 -05:00
Mai M 0db7205f61 Merge pull request #1582 from Sonicadvance1/ccache_option
Adds option to disable ccache
2022-02-16 22:49:48 -05:00
Ryan Houdek 6027494d69 Adds option to disable ccache
Can be useful when running static analysis tools
2022-02-16 19:24:24 -08:00
Mai M 252dcfe26f Merge pull request #1581 from Sonicadvance1/add_required_growsdown
FEXLoader: Adds back required MAP_GROWSDOWN
2022-02-16 19:07:40 -05:00
Ryan Houdek 944c93c10c FEXLoader: Adds back required MAP_GROWSDOWN
I was overzealous with my removal of MAP_GROWSDOWN.
We still require the primary thread to have this flag set.
2022-02-16 15:41:44 -08:00
Ryan Houdek a5fb7e7313 Merge pull request #1580 from Sonicadvance1/fix_hostthunks_install
Fixes Host and guest thunks install path
2022-02-15 15:46:10 -08:00
Ryan Houdek 364b3380fc Fixes Host and guest thunks install path
Hosts were using the cmake install path with $DESTDIR which duplicates
paths.

GuestThunks were doing some magic that wasn't actually necessary
2022-02-15 15:36:31 -08:00
Ryan Houdek a0edab8040 Updates vixl for new cursor updating methods
These will be required for code cache
2022-02-14 16:33:51 -08:00
Ryan Houdek b65194f433 Merge pull request #1576 from Sonicadvance1/move_x87_constant_helpers
JIT: Implements x87 fallback helpers as lookups in to state
2022-02-14 14:13:33 -08:00
Ryan Houdek 0d1c9cd7df JIT: Implements x87 fallback helpers as lookups in to state
This allows x87 fallbacks to be loaded from the upcoming code cache
without relocations.

Only 40 pointers necessary to store and means x87 code won't hit
relocations heavily.

Probably improves performance slightly on the x86 host side but should
be neglible.

Needs #1574 and #1575 merged first.
2022-02-14 14:00:10 -08:00
Ryan Houdek 99dcda7f8c Merge pull request #1575 from Sonicadvance1/move_constant_functions_x86
JITx86: Switches over to loading pointers from state
2022-02-14 13:59:07 -08:00
Ryan Houdek 2aa4d33de5 JITx86: Switches over to loading pointers from state
Just like the previous AArch64 JIT.
These pointers are process or thread specific depending on the pointer
and should be loaded from the State object.

Performance here might slightly increase.

This is required for code cache on x86

Needs #1574 merged first.
2022-02-14 13:49:54 -08:00
Ryan Houdek 1ef78d0a00 Merge pull request #1574 from Sonicadvance1/move_constant_functions
ARMJIT: Switches over to loading pointers from state
2022-02-14 13:39:28 -08:00
Stefanos Kornilios Mitsis Poiitidis 64d1840f1e Merge pull request #1571 from Sonicadvance1/fix_syscall_strace
Linux: Fix missing types for syscall strace
2022-02-14 19:05:48 +02:00
Stefanos Kornilios Mitsis Poiitidis 34530236f9 Merge pull request #1573 from Sonicadvance1/disable_int_tests_with_no_int
unittests: Disables Interpreter tests when its disabled
2022-02-14 19:00:13 +02:00
Ryan Houdek f5a9da082f ARMJIT: Switches over to loading pointers from state
These pointers are process or thread specific depending on which pointer
it is.
All of these pointers end up getting used inside of the JIT blocks
themselves and with code caching would result in a ton of relocations
occuring inside the code.

The pointers used within the dispatcher don't currently matter since I'm
not expecting to cache the dispatcher itself. It's only a page per
thread after all. This does move us significantly closer towards using a
single dispatcher for all threads though.

The performance impact of this change is unlikely to be felt at all,
some locations have less code generation which could improve perf
slightly. Some locations move from a 1-3 cycle constant calculation to a
4 cycle load, hard to be felt since it gets hidden by other
instructions.

x86-64 JIT will be added soon after this
2022-02-13 23:06:26 -08:00
Ryan Houdek 9b49e8cb59 FEXCore: Adds utility class for class member function casting
Adds validation for our class member casting to ensure we don't try
casting a virtual member.
2022-02-13 23:06:25 -08:00
Ryan Houdek 1b6d20b731 unittests: Disables Interpreter tests when its disabled
Would result in failures if you weren't expecting it.
2022-02-11 12:55:10 -08:00
Mai M e9c7c76174 Merge pull request #1572 from Sonicadvance1/LoadConstant_no_opt
Arm64Emitter: Allow non-optimizing LoadConstant
2022-02-11 00:50:56 -05:00
Ryan Houdek cbce06d012 Arm64Emitter: Allow non-optimizing LoadConstant
This is pulled from the code cache PR. Will be necessary for supporting
relocations.

Not currently being used but will be once we have code caching in place.
2022-02-10 20:04:02 -08:00
Ryan Houdek d95326b23b Linux: Fix missing types for syscall strace 2022-02-10 19:46:43 -08:00
Mai M 43fada7555 Merge pull request #1568 from Sonicadvance1/fix_musl_load
ELFCodeLoader: Fixes typo in AT_BASE calculation
2022-02-10 21:25:48 -05:00
Mai M afa7172cb1 Merge pull request #1570 from Sonicadvance1/remove_growsdown
Removes MAP_GROWSDOWN usage
2022-02-10 21:25:23 -05:00
Ryan Houdek b88d8a7cc4 Removes MAP_GROWSDOWN usage
This is just a memory leak waiting to happen.
Only the primary thread in an application really should have this set
since the kernel cleans it up.

We only ever allocate the primary thread of the guest application then
every host thread's stack on top of that. It's up to the guest when it
is cloning to set up new stack pointers, we don't manage that.

We are already allocating the first thread's size at the soft stack
limit with RLIMIT_STACK anyway.

Fixes #1556
2022-02-10 18:03:18 -08:00
Ryan Houdek 6add09b78b ELFCodeLoader: Fixes typo in AT_BASE calculation
Fixes executing musl applications with the dynamic linker.

It was using the main executable's p_offset instead of the
interpreter's.
Wasn't a problem with glibc since it uses a different symbol to find the
base (Don't ask me why it does this).

musl dynamic linker on the other hand just uses AT_BASE directly and
since it was calculated incorrectly it was crashing.

Testing application was `ls` which had a p_offset of 0x40, so it would
try and read some values from AT_BASE, starting at an offset below where
it was mapped.
2022-02-10 17:43:58 -08:00
Mai M 4bb3a54ccf Merge pull request #1566 from neobrain/refactor_thunk_misc
Miscellaneous thunk cleanups
2022-02-10 15:33:34 -05:00
Mai M 76f86e51a6 Merge pull request #1567 from Sonicadvance1/fix_fexgetconfig_rootfs
FEXGetConfig: Fix --current-rootfs option
2022-02-10 15:15:26 -05:00
Ryan Houdek 2a1b27df58 FEXGetConfig: Fix --current-rootfs option
If the configured rootfs wasn't a squashfs then it was failing to return
the directory.

Now it works for both squashfs and directory rootfs again.
2022-02-10 11:45:29 -08:00
Tony Wasserka 68426735a5 Thunks: Clean up ASTMatcher-based testing helpers
The run_thunkgen* helpers now parse generated source code and return its AST
representation, so HasASTMatching helper calls don't each need to redundantly
compile it themselves. This also ensures the generator output actually compiles
in tests where we didn't explicitly check that before.

This also allows printing the full AST of the generator output on test
failures. This must be enabled manually by changing a variable in the ostream
output operator for SourceWithAST.
2022-02-10 12:11:52 +01:00
Tony Wasserka ee6b558000 Thunks: Fix warning about unused field 2022-02-10 12:11:52 +01:00
Tony Wasserka 3c7872c6d5 Thunks: Rename FrontendAction to GenerateThunkLibsAction 2022-02-10 12:11:52 +01:00
Tony Wasserka 4aed6fc56c Thunks/gen: Remove now unneeded code 2022-02-10 12:11:52 +01:00
Tony Wasserka 5d555a10a2 Thunks: Explicitly put thunks into the text library section
Previously, defining zero-initialized variables right before LOAD_LIB
could cause the compiler to put thunk definitions into bss, hence triggering
errors during assembly ("attempt to store non-zero value in section `.bss'").
2022-02-10 12:11:52 +01:00
Ryan Houdek 5854d4ad1c Merge pull request #1565 from neobrain/refactor_thunk_ide_integration
Enable proper IDE integration of thunk libraries
2022-02-10 02:44:29 -08:00
Tony Wasserka bd6999eb87 CMake: Clean up build architecture for ThunkLibs
Host thunk libraries are always built as part of the main project now.
Guest thunk libraries are still cross-compiled in a CMake ExternalProject,
but *additionally* there are CMake targets in the main project to make
sure IDE engines can properly handle guest source files.
2022-02-10 11:23:43 +01:00
Tony Wasserka 3ebb2eaf0d Thunks: Fix guest libs build on clang 2022-02-10 11:23:41 +01:00
Mai M defd3be30c Merge pull request #1563 from Sonicadvance1/fix_auto_script
Updates Readme to fix install script
2022-02-09 18:18:39 -05:00
Mai M a8e5a0a68e Merge pull request #1562 from Sonicadvance1/remove_debug_memory_mapping
FEXLoader: Removes memory mapping check on startup
2022-02-09 18:18:18 -05:00
Ryan Houdek f27c43ff8a Updates Readme to fix install script
Fixes an issue where the FEXRootFSFetcher wouldn't get a any user input
and just fail out.
Save it to the tmp folder and execute from there instead.

Fixes #1557
2022-02-09 14:01:48 -08:00
Ryan Houdek 8840fc818c FEXLoader: Removes memory mapping check on startup
FEX always builds with PIE and we don't hit this issue anymore anyway.
If some application wants to inject a page in to the lower 32-bits then
we have no reason to complain about it anymore. Just let it go and
hopefully they know what they are doing.

Fixes #1559
2022-02-09 13:52:05 -08:00
Mai M ffcaf294e3 Merge pull request #1555 from Sonicadvance1/weirdo_edge_case
OpcodeDispatcher: Fixes weirdo edge case in segment moving
2022-02-08 00:26:55 -05:00
Ryan Houdek 6b7a84bef2 OpcodeDispatcher: Fixes weirdo edge case in segment moving
Just noticed this while casually reading the x86 architecture manuals.
The move segment registers instructions ignore the REX.R prefix on the
segment register.

Previously this was expected to create an invalid register selection.
A little bit silly but sure, support it.
2022-02-07 21:12:53 -08:00
Mai M 287c65dc64 Merge pull request #1554 from Sonicadvance1/fix_tricky_stat
Linux: x32: Fixes tricky stat64 defines
2022-02-06 19:52:31 -05:00
Mai M 62397bb19c Merge pull request #1553 from Sonicadvance1/fix_sigevent
Linux: Make sure to use correct accessors for sigevent
2022-02-06 19:52:14 -05:00
Ryan Houdek 0216bcf27f Linux: x32: Fixes tricky stat64 defines
Some build environments use a define to change stat64 and statfs64 to be
the same definition as stat and statfs.

Check if the define exists and if it does then remove the 64bit
constructors.
2022-02-06 16:05:27 -08:00
Ryan Houdek b981fcfe42 Linux: Make sure to use correct accessors for sigevent
Some of these are defined differently depending on environment
2022-02-06 15:33:07 -08:00
Mai M cb491a8acb Merge pull request #1552 from Sonicadvance1/fix_ucontext_copy
UContext: Fixes 32-bit siginfo_t copying definition
2022-02-06 18:23:04 -05:00
Mai M 1b99495b4b Merge pull request #1551 from Sonicadvance1/fix_older_env
Some fixes for older environments
2022-02-06 18:22:26 -05:00
Ryan Houdek 3c5a2cec90 UContext: Fixes 32-bit siginfo_t copying definition
The host provided siginfo_t definition can vary depending on the build
environment.
What doesn't change however is how the data is laid out.
It's always 128bytes, The first three 32-bit words are always known.
The 64-bit host side always has an additional 32-bit pad member.
Then the remaining bytes is the sifields.
2022-02-06 15:03:56 -08:00
Ryan Houdek 9a64e7f100 Linux: Renamed some 64-bit syscall names
Some build environments use defines to rename these. Which breaks our
naming
2022-02-06 14:19:30 -08:00
Ryan Houdek 8e8baec47a Linux: Use raw syscalls for pkey syscalls
For older libc environments
2022-02-06 14:19:30 -08:00
Ryan Houdek 59ca60e39f Fixes a bunch of header includes
Necessary for older build environments
2022-02-06 14:19:30 -08:00
Ryan Houdek 8c956e6ce1 Docs: Update for release FEX-2202 2022-02-05 22:48:35 -08:00
Ryan Houdek 832d013c92 Merge pull request #1513 from Sonicadvance1/reduce_flags_memory_usage
FEXCore: Defer a significant number of ALU flag calculation
2022-02-04 16:22:24 -08:00
Mai M 1c24206117 Merge pull request #1550 from Sonicadvance1/fix_weirdo_crc32
OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
2022-02-04 01:01:50 -05:00
Ryan Houdek f1979c15a2 unittests: Adds new CRC32 unittests
The instruction decode tables for crc32 introduced some dumb.
`F2h` and `F2h && 66h` prefixes both work for crc32.
This is a failure on Intel's part for sticking crc32 in to the vector
table.

MOVBE without any prefixes also does the same garbage where prefix `66h`
acts as an operand prefix size ONLY.
2022-02-03 21:11:50 -08:00
Ryan Houdek 556a1dab24 OpcodeDispatcher: Fixes CRC32 decoding in 0F_38 table
This table is particularly terrible. CRC32 is the first instruction in
this table that needs either prefix `72h` OR `66h && F2h`

For 8bit CRC32, this ignores the 66h operand size override prefix.
  - But our table decoding didn't handle this
For 16bit/32bit/64bit CRC32 this behaviour changes depending on 66h
prefix AND REX.W
  - 66h prefix is ignored when REX.W is set, always 64bit but it falls
    down the other table path

This is an absolutely weird edge case that nobody should hit, but here
we are.
2022-02-03 21:11:50 -08:00
Mai M caffad8562 Merge pull request #1549 from Sonicadvance1/implement_pcmpgtq
OpcodeDispatcher: Implements PCMPGTQ
2022-02-03 21:46:13 -05:00
Mai M 5978143141 Merge pull request #1547 from Sonicadvance1/remove_system_xxhash
CMake: Always use local xxhash to statically link
2022-02-03 21:45:58 -05:00
Ryan Houdek 594c70b5e0 OpcodeDispatcher: Implements PCMPGTQ
I thought we already had this implemented but I guess it was missed.

Required for SSE 4.2
2022-02-03 18:36:29 -08:00
Ryan Houdek 655e6989ca FEXCore: Defer a significant number of ALU flag calculation
This was mainly an optimization around memory usage. ALU ops tend to
bloat the IR quite heavily, but I also noticed a 2-4% uplift in
performance of some applications. So a nice side effect.

Should let us more aggressively target reducing our IR intrusive
allocator size since this is quite reduced.

In a pedantic heavy ALU op code block this reduces the number of IR ops
from 14,756 IR ops to 2,016 prior to optimization.
After optimization both had reduced down to 50 IR ops, proving the
output IR was the same.
2022-02-03 01:37:59 -08:00
Ryan Houdek afeb228a89 CMake: Always use local xxhash to statically link
Dynamically linking xxhash is causing problems with pressure-vessel.

With this in place we only have the typical C++ dependencies
```
$ ldd ./Bin/FEXLoader
        linux-vdso.so.1 (0x00007fff44d9d000)
        libstdc++.so.6 => /lib/x86_64-linux-gnu/libstdc++.so.6 (0x00007f4c4d884000)
        libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007f4c4d7a0000)
        libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1 (0x00007f4c4d786000)
        libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007f4c4d55e000)
        /lib64/ld-linux-x86-64.so.2 (0x00007f4c4e0fa000)
```
2022-02-03 01:31:43 -08:00
Mai M d308a438ea Merge pull request #1546 from Sonicadvance1/fix_fexconfig
Fixes FEXConfig build
2022-02-01 21:13:08 -05:00
Ryan Houdek 260fc8ba52 Fixes FEXConfig build
Oops. This was added late and didn't test it.
2022-02-01 15:44:06 -08:00
Mai M ade0d0f241 Merge pull request #1543 from Sonicadvance1/fixes_for_1423
Linux: Fixes for older build environments
2022-02-01 16:52:40 -05:00
Mai M 11a5105547 Merge pull request #1544 from Sonicadvance1/allow_disable_interpreter
Adds an option to disable the IR interpreter
2022-02-01 16:52:22 -05:00
Ryan Houdek 10ad5db686 Adds an option to disable the IR interpreter
By default we won't build with the interpeter to reduce user confusion.
The interpreter isn't really useful to end users so remove it.

Completely removes it from building except for the fallback operations.

This also removes the selection from FEXConfig to remove selection
confusion there.

File Stats:
FEXLoader Size with Interpreter:    3422768 bytes
FEXLoader Size without Interpreter: 3301944 bytes
Size difference:                    96.4699915%
Bytes removed:                      120824 bytes
4k pages removed:                   29.498046875 -> 30 rounded up

VM Stats (Reported from bloaty):
Memory Size with Interpreter:    6.50Mi
Memory Size without Interpreter: 6.38Mi
Size difference:                 98.1538462%
2022-02-01 13:00:29 -08:00
Ryan Houdek 68c441575d Linux: Fixes for older build environments
Should resolve the new building issues from #1423
2022-02-01 12:17:09 -08:00
Ryan Houdek 334a8ef87c Merge pull request #1542 from Sonicadvance1/fix_pressure_vessel_hangs
Fix pressure vessel hangs
2022-01-31 08:57:20 -08:00
Ryan Houdek b7a76af72f Merge pull request #1541 from Sonicadvance1/implement_crc
OpcodeDispatcher: Implements CRC32 instruction
2022-01-31 08:57:01 -08:00
Ryan Houdek 4c92b562b8 Merge pull request #1540 from Sonicadvance1/remove_extract
OpcodeDispatcher: Removes extraneous extract in VFCMP
2022-01-31 08:56:47 -08:00
Stefanos Kornilios Mitsis Poiitidis 9d08451903 Merge pull request #1536 from Sonicadvance1/fix_orbitals
Softfloat: Stop doing special handling for FREM
2022-01-31 16:51:18 +02:00
Stefanos Kornilios Mitsis Poiitidis c252f8bfc5 Merge pull request #1539 from Sonicadvance1/fix_wrong_offsets
IR: Fixes some wrong offsets in passes
2022-01-31 15:40:33 +02:00
Ryan Houdek dc7ec6377b Linux: Safely handle Filemanagement mutex on fork
If an application is forking heavily with threaded file accesses
happening then the mutex can end up in an unknown state.

On fork make sure to lock the mutex then immediately unlock after fork
occurs.

This final step resolves hanging that pressure-vessel hits on startup.
Since it is doing a ton of file opening and forking during
initialization.
2022-01-30 18:15:57 -08:00
Ryan Houdek ce6f4edaaa FileManagement: Use ScopedSignalMaskWithMutex
When using mutexes in syscall helpers we need to be extra careful around
signals.
2022-01-30 18:15:57 -08:00
Ryan Houdek 983c35ea3b Allocator: Use ScopedSignalMaskWithMutex
Instead of just a basic mutex, also mask the signals.
This fixes the problem where we can end up receiving a signal in the
middle of memory allocation. Thus leaving the locked mutex in a broken
state.

This more closely matches the Linux kernel behaviour.
Since if you're in the middle of a memory allocating syscall, you won't
get signaled.
2022-01-30 18:15:57 -08:00
Ryan Houdek 70aaa1117a FEXHeaderUtils: Adds ScopedSignalMaskWithMutex
This class allows a scoped region lock a mutex and mask signals.

This is necessary for thread and signal safety coming up
2022-01-30 18:15:57 -08:00
Ryan Houdek 59e9859087 unittests: Implements CRC32 unit tests 2022-01-30 15:38:26 -08:00
Ryan Houdek d9453ff639 OpcodeDispatcher: Implements CRC32 instruction
Now that the rest of the code matches behaviour, we just need to pass
this through.

Easy enough and get Horizon Zero Dawn running.
2022-01-30 15:38:26 -08:00
Ryan Houdek 70754991d1 CPUID: Fill out CPUID for SSE4.2 feature
Currently force disabled until the rest of SSE 4.2 is enabled
This is to remind us in the future that SSE4.2 can only be enabled in
CPUID with CRC32 instruction support.
2022-01-30 15:38:26 -08:00
Ryan Houdek 57ebfceb48 HostFeatures: Check for CRC32 op support
Available with CRC32 bit on Arm64 or SSE4.2 on x86-64
2022-01-30 15:38:26 -08:00
Ryan Houdek 9e224d2bb0 x86 JIT: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek e43bd04901 JITArm64: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 48762e03a6 Interpreter: Implements CRC32 op 2022-01-30 15:38:26 -08:00
Ryan Houdek 858924309e IR: Implements CRC32 op 2022-01-29 23:33:53 -08:00
Ryan Houdek cab02d1e65 OpcodeDispatcher: Removes extraneous extract in VFCMP
We don't need to extract the element to compare it.
2022-01-28 22:22:34 -08:00
Ryan Houdek 2a64f80567 IR: Fixes some wrong offsets in passes
GPR and FPR ending offsets were off by one here. Just a quick fix.
2022-01-28 22:19:51 -08:00
Ryan Houdek 174ddea99d Softfloat: Stop doing special handling for FREM
This isn't correct and breaks games.
This makes the FREM and REM1 implementation the same.
While not 100% correct, it is still better than before.
New issues will be created to handle the differences in the future.

Fixes #1374.
Also fixes most of the HL2 issues, just not the seam issue.
2022-01-28 19:42:57 -08:00
Ryan Houdek 6fb0e3c0cf Softfloat: Allow x87 fallback for all ops 2022-01-28 19:42:25 -08:00
Ryan Houdek 2b044bbdf4 Disable fprem unittests
These are about to be broken
2022-01-28 19:39:03 -08:00
Mai M ea76de0fd2 Merge pull request #1533 from Sonicadvance1/revise_posix_tests
unittests: Revise POSIX tests known failures and disabled
2022-01-25 15:19:24 -05:00
Ryan Houdek 9c8642e0dc unittests: Revise POSIX tests known failures and disabled
Some of these behaviours have changed now, particularly around signal
handling.

Some things still fail now of course. But most everything is now
documented as to why it is failing or disabled.

Fixes #955
2022-01-25 11:41:45 -08:00
Ryan Houdek 13f35f7b79 Merge pull request #1530 from Sonicadvance1/rootfs_fetcher_fixes
FEXRootFSFetcher: Fixes some edge case behaviours
2022-01-25 10:29:55 -08:00
Ryan Houdek e2798e370e Merge pull request #1518 from Sonicadvance1/fix_signed_branch
JIT: Fixes signed displacement wraparound on 32-bit
2022-01-25 10:29:46 -08:00
Ryan Houdek d9149548b5 Merge pull request #1531 from Sonicadvance1/fix_sockopt
Linux: Fixes 32-bit getsockopt and setsockopt
2022-01-25 09:09:16 -08:00
Ryan Houdek 4823933f79 Linux: Fixes 32-bit getsockopt and setsockopt
On Set, we have four options that need to be converted.
On Get, we have two options that need to be converted.

This fixes a crash that Tomb Raider 2013 was having on launch.
2022-01-24 17:14:50 -08:00
Ryan Houdek ee04067424 FEXRootFSFetcher: Fixes some edge case behaviours
Makes curl do its continue feature to give the users the best chance of
downloading a rootfs. We don't need to restart the full file transfer on
failure. Helps people with slower connections.

On failure to download, asks the user if they want to retry the download
rather than just exiting with a weird error about hash failure.

Once the image is downloaded, now changes options depending on if
squashfuse or unsquashfs works.

Prevents the user from selecting a bad option and getting unexpected
behaviour. Ideally we would do a squashfs mount test as well for
platforms that don't have working FUSE, like termux. This is harder to
get right and its for an unsupported platform, so I'm not going to
invest more time with it.

Fixes #1525
Fixes #1526
Fixes #1527
2022-01-23 22:54:05 -08:00
Ryan Houdek e4aef26ef5 FEXRootFSFetcher: Adds helper namespace for tool checking
Location to check if curl, squashfuse, and unsquashfs are working.

unsquashfs is a bit more complex where it needs to parse the help output
to see if zstd is supported
2022-01-23 22:44:31 -08:00
Ryan Houdek a41dc8eafa FEXRootFSFetcher: Fix pipe redirecting
In the case of launching without stdout/stderr then redirection could
have these constants be a redirected FD that sits in the same fd number.

Use -2 to indicate no redirection.
Use -1 to indicate closing traditional stderr/stdout
The rest will indicate if stdout and stderr should be replaced as
normal.
Making sure not to close the incoming fds if they matched the
stdout/stderr FD numbers.
2022-01-23 22:41:39 -08:00
Ryan Houdek f41cd8deff OpcodeDispatcher: Renamed GetDynamicPC to GetRelocatedPC
For clarity.
2022-01-23 18:52:53 -08:00
Ryan Houdek 2a0c3cce30 Core: Have GetDynamicPC mask based on operating size
This ensures on 32-bit we overflow correctly under relocation.
2022-01-23 18:50:16 -08:00
Ryan Houdek e817f5d98c unittests: Adds 32-bit tests for signed displacement wraparound
A bit meta since it needs to JIT some minor code but easy enough.
Ensures something like #1517 won't happen again.
2022-01-23 18:38:44 -08:00
Ryan Houdek 8b8cda9b80 JIT: Fixes signed displacement wraparound on 32-bit
This cropped up mostly with multiblock and `jmp <signed displacement>`
This also happened with non multiblock `jcc <signed displacement>`

Due to how IR relocations occur, this needs to happen fairly late but
isn't a big deal.

Fixes #1517
2022-01-23 18:38:43 -08:00
Ryan Houdek 8e3893df07 Merge pull request #1523 from lioncash/vixl-update
Externals: Update vixl
2022-01-20 15:13:36 -08:00
lioncash 51b335914c github: Synchronize submodules before checking them out
Ensures that we don't get stale remotes.
2022-01-20 17:57:45 -05:00
lioncash 8835d57ae3 Arm64Emitter: Adjust XRegister to Register
With the updated API, we need to make use of Register as opposed to
XRegister in our arrays.
2022-01-20 16:42:04 -05:00
lioncash eba1b65fb0 Externals: Update vixl to updated branch
Now we have access to some SVE goodies.
2022-01-20 16:42:02 -05:00
Ryan Houdek a3a138ef7e Merge pull request #1520 from lioncash/vixl
External: Point vixl submodule towards FEX's fork
2022-01-14 13:40:18 -08:00
lioncash 62d9a494cd External: Point vixl submodule towards FEX's fork
This allows it to be managed by all organization members
2022-01-14 12:21:19 -05:00
Stefanos Kornilios Mitsis Poiitidis 6744a06a53 Merge pull request #1519 from Sonicadvance1/aarch64_single_instruction_opt
AArch64: Single instruction optimization for AESKeyGenAssist
2022-01-14 15:41:57 +02:00
Ryan Houdek 140e9824b7 Merge pull request #1516 from lioncash/fmt
externals: Update fmt to 8.1.1
2022-01-14 01:55:50 -08:00
Ryan Houdek bad84f61fa AArch64: Single instruction optimization for AESKeyGenAssist
No need to do adr when loads can do a 1MB offset loadstore
2022-01-14 01:49:21 -08:00
lioncash 2296126af3 externals: Update fmt to 8.1.1
Brings along a bunch of enhancements and ensures we always build against
the latest version.

Also fixes up a few issues that arose due to changes in fmt
2022-01-13 14:48:35 -05:00
Stefanos Kornilios Mitsis Poiitidis 0a8717d9a8 Merge pull request #1515 from Sonicadvance1/fix_ptest
OpcodeDispatcher: Fixes ptest flags calculation.
2022-01-13 10:38:38 +02:00
Ryan Houdek 4e2220c27f unittests: Adds ptest unit test to ensure correct flag setting
ptest wasn't correctly setting OF, SF, AF, and PF to zero until now.
Do a unit test to ensure correct behaviour here
2022-01-11 16:45:56 -08:00
Ryan Houdek c87e11cee9 OpcodeDispatcher: Fixes ptest flags calculation.
We were missing four flags that require setting zero.
2022-01-11 16:45:11 -08:00
Ryan Houdek 7768f6965a Merge pull request #1501 from Sonicadvance1/finish_siginfo_32bit
Linux: Handles the remaining 32-bit siginfo_t usage
2022-01-11 00:02:55 -08:00
Ryan Houdek 023aaaae0c Merge pull request #1499 from Sonicadvance1/resolve_rootfs_path_in_interpreter
FEXLoader: Resolve the absolute path to rootfs if possible
2022-01-11 00:02:27 -08:00
Stefanos Kornilios Mitsis Poiitidis a2aa9f3fc1 Merge pull request #1512 from Sonicadvance1/fix_ssa_id_print
IR: Fixes SSA ID printing
2022-01-11 09:09:06 +02:00
Ryan Houdek e46ec9a0ce IR: Fixes SSA ID printing
These should print as decimal. They were ending up as hex
2022-01-10 18:09:03 -08:00
Ryan Houdek 784cbdd973 Merge pull request #1500 from Sonicadvance1/rootfsfetch_check_curl
FEXRootFSFetcher: Check if curl is installed and fail before running
2022-01-10 16:23:05 -08:00
Ryan Houdek 6022715a9b Merge pull request #1510 from Sonicadvance1/fix_asan_cpuid
CPUID: Fixes ASAN problem with reading midr
2022-01-10 02:12:10 -08:00
Ryan Houdek 9eb5ba5ad1 Merge pull request #1509 from Sonicadvance1/fix_logserver_sync
SocketLogging: Fixes MsgHandler not syncing with Assert level
2022-01-10 02:12:01 -08:00
Ryan Houdek 82e5977709 Merge pull request #1506 from Sonicadvance1/fix_apitest_syscalls
APITests: Fixes InterruptableConditionVariable test to use the syscal…
2022-01-10 02:11:44 -08:00
Ryan Houdek 7c08b67dff Merge pull request #1504 from Sonicadvance1/fix_unittest_rootfs_define
unittests: Fixes ROOTFS needing to be defined prior to cmake
2022-01-10 02:11:35 -08:00
Ryan Houdek 609587f9ee Merge pull request #1503 from Sonicadvance1/implement_bcd_tests
unittests: Adds a BCD unit test
2022-01-10 02:11:07 -08:00
Ryan Houdek 5dda3a1599 Merge pull request #1497 from Sonicadvance1/fix_alternative_links
Linux: Fixes emulatedpath with symlink following
2022-01-10 02:10:54 -08:00
Ryan Houdek 73aaa4c3a6 CPUID: Fixes ASAN problem with reading midr
Needs to be a string_view for the MIDR for the StrConv helper to work in
this instance.
There is no null terminator character when reading from the file is why.
2022-01-10 01:21:56 -08:00
Ryan Houdek 3ba5371d36 SocketLogging: Fixes MsgHandler not syncing with Assert level
AssertHandler by default synchronizes but MsgHandler with Assert level
should also synchronize.

Fixes an issue where LogMan::Msg::AFmt wasn't syncing so the
FEXLogServer would never see the messages.
2022-01-10 01:16:53 -08:00
Ryan Houdek bf581decde Merge pull request #1507 from Sonicadvance1/fix_warnings
Fixes some of the warnings that cropped up
2022-01-10 01:12:39 -08:00
Ryan Houdek 250504502a Fixes some of the warnings that cropped up 2022-01-10 00:46:10 -08:00
Ryan Houdek a0e826feaf APITests: Fixes InterruptableConditionVariable test to use the syscall wrappers.
Fixes a build error on old Ubuntu
2022-01-09 22:00:57 -08:00
Ryan Houdek eb17edec05 Merge pull request #1502 from Sonicadvance1/fix_fexlog_server_message
FEXLogServer: Stop duplicating and dropping messages
2022-01-09 03:22:24 -08:00
Ryan Houdek 228aed98c7 unittests: Fixes ROOTFS needing to be defined prior to cmake
cmake will bake in the environment variable in to the build scripts.
Instead have the guest_test_runner fetch it at runtime.

This means if you forget to set ROOTFS prior to running cmake, you can
now set it afterwards and rerun with just ctest instead of a cmake
dance.

Fixes #315
2022-01-09 01:56:13 -08:00
Ryan Houdek 6ba4aec88e unittests: Adds a BCD unit test
Nothing really fantastical found here. Just that sub-precision results
weren't rounded correctly on store

Fixes #770
2022-01-09 01:36:26 -08:00
Ryan Houdek b8d6b2cd4a F80: Ensures BCDStore rounds to the current rounding mode
BCD storing will round any subprecision results depending on the current
rounding mode.
2022-01-09 01:35:28 -08:00
Ryan Houdek d4e2f42f90 FEXLogServer: Stop duplicating and dropping messages
In the case that multiple messages appearing in a single packet then we
were repeating the first message and dropping any subsequent messages.

Fixes #1496
2022-01-09 00:22:12 -08:00
Ryan Houdek 3055c23365 Linux: Handles the remaining 32-bit siginfo_t usage
Just need to translate them between 32-bit and 64-bit versions.

Fixes #1254
2022-01-09 00:03:17 -08:00
Ryan Houdek cb7feaefbb Types: Allows passing 64-bit host siginfo_t to 32-bit siginfo_t
Needed for waitid
2022-01-09 00:00:13 -08:00
Ryan Houdek 5a5a498ed6 FEXRootFSFetcher: Check if curl is installed and fail before running
Before doing anything that requires curl, actually check if it is
installed.
Then instruct the user to install curl before using.

Doesn't try installing curl itself since we don't have a clean way to
execute sudo from potentially GUI.

Fixes #1498
2022-01-08 21:35:29 -08:00
Ryan Houdek 285ed8f1e0 FEXRootFSFetcher: Adds new Exec function with stdout,stderr redirection
Just so we can test for applications without spamming terminal
2022-01-08 21:34:59 -08:00
Ryan Houdek a68da468a5 FEXRootFSFetcher: ExecAndWaitForResponse sign extend program result
Only the lower 8bits of the execve result is the program result.
Makes sure to sign extend it so -1 is a true -1 instead of 255
2022-01-08 21:33:24 -08:00
Ryan Houdek d59aa6874e FEXLoader: Resolve the absolute path to rootfs if possible
If the user passes in an absolute path then check to see if it exists in
the rootfs before executing.

Useful for launching applications directly out of the rootfs with
FEXInterpreter.

In the case that the absolute path doesn't exist in the rootfs then
fallback to the host system as usual
2022-01-07 03:42:41 -08:00
Ryan Houdek 19fd89d2bf Linux: Fixes emulatedpath with symlink following
Some syscalls support `AT_SYMLINK_NOFOLLOW` In these instances we need
to follow the symlink on a couple of syscalls.

Fixes executing wine using the basic wine path
eg:
FEXBash "wine dxcapsviewer.exe"
2022-01-07 02:26:26 -08:00
Mai M e36beb8dbe Merge pull request #1495 from Sonicadvance1/add_tune_arch
CMake: Adds TUNE_ARCH option
2022-01-05 18:09:26 -05:00
Ryan Houdek 70f447b265 CMake: Adds TUNE_ARCH option
I forgot about this option working for tuning arch on AArch64. This will
be used in PPA releases in the future. Will leave the previous option
since it can be used in testing.
2022-01-05 13:48:35 -08:00
Ryan Houdek a0026c92a8 Merge pull request #1494 from Seas0/main
ThunkLibs: Add meta data to libvulkan_device
2022-01-04 23:37:54 -08:00
Seas0 8fc5f66a5b ThunkLibs: Add meta data to libvulkan_device 2022-01-05 14:28:26 +08:00
Ryan Houdek affbb40cc1 Docs: Update for release FEX-2201 2022-01-03 19:08:53 -08:00
Ryan Houdek b276750c8a Merge pull request #1488 from Sonicadvance1/python_auto_install
Scripts: Adds a python script that can hand hold a user through FEX install
2022-01-03 19:07:41 -08:00
Ryan Houdek 6d5322b349 Readme: Updates with links to the wiki's quick start guide and python command
Quick start guide on the wiki is currently not populated but will be
soon.
2022-01-03 16:05:05 -08:00
Ryan Houdek f836ff0897 Scripts: Adds a python script that can hand hold a user through FEX install
This is definitely a bit divisive but overall this is a win.
This is a user pattern that is emerging in a bunch of projects.
Allow an officially sourced script that lets you pipe a script directly
in to python/bash and setup the environment entirely.

This only supports Ubuntu {20.04, 21.04, 21.10, 22.04} which matches
exactly what we expose in the PPA.

Once this is in the repo, and our PPA is updated to the latest release
tag you can run this script like:
`curl --silent <Direct Github raw link> | python3`

Once the PPA is updated, the README and Wiki will be updated with Quick
Start guides to use this path.

The script steps
1) Checks if ARMv8, or fail
2) Checks if supported Ubuntu, or fail
3) Checks if PPA installed
3a) Install PPA if not, or fail
4) Check if packages are installed
4a) Install non-installed packages, or fail
5) Check if RootFS is configured/exists
5a) Run through FEXRootFSFetcher to get/setup rootfs, or fail
6) Attempt to run emulated uname -a through FEX, or fail
7) Provide some examples for how to use FEX
8) Exit with success!
2022-01-03 14:20:09 -08:00
Mai M 0bf577654d Merge pull request #1491 from Sonicadvance1/fix_signal_pause
FEXCore: Don't use Run() inside RunUntilExit()
2022-01-03 06:56:43 -05:00
Mai M ac63a2bac4 Merge pull request #1493 from Sonicadvance1/fix_pidfd_send_signal32
Linux: Fixes pidfd_send_signal for 32-bit
2022-01-03 06:37:13 -05:00
Ryan Houdek 6b8e6c18ee SignalDelegator: Increase signal stack size
Just to help in the case that a program will dequeue a bucketload of
signals in one go. Basic signal stack size wasn't enough.
2022-01-03 03:29:35 -08:00
Ryan Houdek 93dccffc81 unittests: Updates posix tests with new passing tests 2022-01-03 03:29:35 -08:00
Ryan Houdek 22810208b4 FEXCore: Don't use Run() inside RunUntilExit()
Run() is expected to be used after a pause at this point.
Make RunUntilExit() notify using the StartRunning variable entirely.

This is because if we used Run() then the initial thread hasn't yet set
up its core or signal handlers. So it'll ignore SIG63. Then its reason
for pause is still set to Resume.
Once an application tries sending SIG63, we try restoring the context
from a paused state that does exist and crash
2022-01-03 03:29:35 -08:00
Mai M 154468c172 Merge pull request #1492 from Sonicadvance1/disable_unsync_log_message
Dispatcher: Disables unsync context message
2022-01-03 06:26:08 -05:00
Mai M 0e45b98fb2 Merge pull request #1490 from Sonicadvance1/fix_cpuid_memcpy
CPUID: Fixes another memory overflow issue
2022-01-03 06:24:47 -05:00
Ryan Houdek 2d573f7bf9 Linux: Fixes pidfd_send_signal for 32-bit
Due to the way this syscall handles siginfo_t, we only need to shift
data through the padding.

This is due to the syscall not allowing you to send siginfo_t with
si_code being positive. Which is what the kernel uses to interpret the
data if it is reinterpreting it.
2022-01-03 03:19:12 -08:00
Ryan Houdek 01587270d9 Dispatcher: Disables unsync context message
This message can actually cause a perfectly running application to crash
due to signaling at the wrong time.

Instead constexpr disable it so if you want to see it. You can still
enable it
2022-01-03 03:02:33 -08:00
Ryan Houdek 0a63a15287 CPUID: Fixes another memory overflow issue
Just like the previous one, but this time when we are NOT on a release
2022-01-03 02:56:48 -08:00
Mai M a65e7a09e7 Merge pull request #1489 from Sonicadvance1/release_process_documentation
Docs: Adds documentation about the FEX monthly release process
2022-01-03 01:01:38 -05:00
Ryan Houdek 7e5c561849 Docs: Adds documentation about the FEX monthly release process
Walks through all the steps required for a monthly release.
2022-01-02 21:13:54 -08:00
Mai M a45bc7598c Merge pull request #1487 from Sonicadvance1/rootfs_stop_trying_so_hard
RootFS: Stop trying to retry rootfs after five times
2022-01-01 03:40:40 -05:00
Ryan Houdek 9f8071ce8f RootFS: Stop trying to retry rootfs after five times
In most cases this will be fixed on the first retry, but give it a
chance at five.
Otherwise you have a chance to infinite loop.
2021-12-31 20:04:57 -08:00
Mai M 0808599813 Merge pull request #1486 from Sonicadvance1/add_version_override
CMake: Adds an OVERRIDE_VERSION option
2021-12-31 02:51:15 -05:00
Mai M 340699ca33 Merge pull request #1485 from Sonicadvance1/fix_fortify
FEXCore: Fixes a memcpy overflow in processor brand
2021-12-31 02:51:06 -05:00
Ryan Houdek 00a572c00c CMake: Adds an OVERRIDE_VERSION option
This is intended to be used by a package maintainer to override the
version when the git repo or git executable isn't available.

Not expected to be used by normal users. Instead by our automated ppa
tooling.
2021-12-30 22:55:46 -08:00
Ryan Houdek d300181ab9 FEXCore: Fixes a memcpy overflow in processor brand
When our FEX version string has less than 16 characters this would
result in size_t underflow, resulting in a crash.

This only ever occurs on a release tag so it went uncaught at first.
The PPA builder managed to uncover this problem since it only deals with
release tags.
2021-12-30 22:54:44 -08:00
Ryan Houdek 5146651908 Merge pull request #1482 from Sonicadvance1/override_tune_cpu
CMake: Adds a TUNE_CPU option
2021-12-29 20:36:04 -08:00
Ryan Houdek 875bae41a3 CMake: Adds a TUNE_CPU option
By default we tune to the native CPU, using some heuristics to determine
the true native CPU since Apple doesn't expose an MIDR.

Adds an option that allows the user to pass in the CPU to tune for.
This will be useful for debugging and also for package building.
2021-12-29 17:27:19 -08:00
Mai M 95bd309db9 Merge pull request #1481 from Sonicadvance1/fix_config_dependency
FEXCore: Fixes FEXCore_Base config dependency
2021-12-29 20:07:22 -05:00
Ryan Houdek 6f1e7e7a4c FEXCore: Fixes FEXCore_Base config dependency
This depends on Config. Then everything else links to this, pulling in
the dependency.
2021-12-29 16:52:26 -08:00
Stefanos Kornilios Mitsis Poiitidis b078f51199 Merge pull request #1480 from Sonicadvance1/fix_asmtest_dependencies
unittests: Fixes a missing dependency on ASM tests
2021-12-29 15:37:32 +02:00
Ryan Houdek d03856f0b7 unittests: Fixes a missing dependency on ASM tests
The output asm folder needs to be created before the config file can be
generated. Otherwise the python script will fail.
2021-12-29 00:27:13 -08:00
Ryan Houdek 98210fdf47 Merge pull request #1479 from Sonicadvance1/fix_msqid_ds
Linux: Finish 32-bit msqid_ds usage
2021-12-28 23:37:20 -08:00
Ryan Houdek 3d1cef4afe Linux: Documents msgctl syscall only supporting IPC_64
I thought this was incorrect but since it only supports IPC_64 it is
actually correct.
If a 32-bit application wants to use the old encoding then it needs to
use the ipc syscall instead.
Fixes #1253
2021-12-28 23:04:53 -08:00
Ryan Houdek 060ab98378 Linux: Implements 32-bit msgctl IPC_SET
Needed struct version conversion support.
Removes the remaining UNHANDLED defines since we support them all now.
2021-12-28 23:02:18 -08:00
Ryan Houdek 1d5bd1520a Merge pull request #1478 from Sonicadvance1/stop_hardcode
Stop hardcoding /usr/bin paths
2021-12-28 22:35:15 -08:00
Ryan Houdek 9eca823e84 Stop hardcoding /usr/bin paths
Instead of using execve, use execvpe.
This glibc helper will search PATH or `confstr(_CS_PATH)`

Fixes #1475
2021-12-28 21:40:22 -08:00
Ryan Houdek cd4269f4e8 Merge pull request #1477 from Sonicadvance1/implement_into
FEXCore: Implements recoverable INTO instruction
2021-12-28 20:38:32 -08:00
Ryan Houdek 2f1d44f838 unittests: Basic INTO test
This only tests for the non-faulting INTO instruction.

This is because our ASM tests can not test for signals. Nor can it
recover.
2021-12-28 19:58:35 -08:00
Ryan Houdek 95eb456065 FEXCore: Implements recoverable INTO instruction
This instruction is only available on 32-bit x86. On overflow flag set,
this instruction will raise a SIGSEGV with overflow exception set in
siginfo_t's si_code.

This implements the everything required to raise the signal and modify
our signal results that we are giving to the guest.

This is effectively the first step in generating signal frames with
custom data in it but only enough for INTO today.
2021-12-28 19:58:35 -08:00
Ryan Houdek d4b31cd4c6 Merge pull request #1476 from Sonicadvance1/squashfs_robust_lock
RootFS: Be more robust against stale lock files
2021-12-28 17:26:54 -08:00
Ryan Houdek c611eb228f RootFS: Be more robust against stale lock files
In the case of having a stale lock file on the system, usually due to an
unclean shutdown, have FEX retry if the socket for that lock file is
also not available.

This resolves the exact case of:
```
FEXBash xeyes &
kill -9 `pidof FEXMountDaemon`
sudo shutdown -r now
FEXBash xeyes & <--- This command would now fail
```

The first command would do:
 - Check lock file and socket
 - Spin up FEXMountDaemon
 - Create Lock file and Socket file
 - Start executing XEyes in FEXInterpreter

The second command would do:
 - Kills the FEXMountDaemon process
 - This leaves squashfuse running
 - The lock file and socket file are now unmanaged and just files on the filesystem

The third command would do:
 - unmounts the squashfuse mount
 - Leaving a dangling folder as a mount point

The fourth command would do:
 - Check for lock file and socket
 - Attempt to use lock file, socket, and mount that is dangling
 - Fail with error

Now instead fourth command will do:
 - Check for lock file and socket
 - Check if filesystem is valid
 - Knows that lock file exists but socket isn't active at all
 - Deletes lock file
 - Spin up FEXMountDaemon
 - Create Lock File and Socket File
 - Start executing XEyes in FEXInterpreter
2021-12-28 16:48:46 -08:00
Ryan Houdek bc120b9f97 Merge pull request #1471 from Sonicadvance1/remaining_syscalls
Linux: Implements the remaining 32-bit syscalls
2021-12-27 19:25:35 -08:00
Ryan Houdek 995672921f Linux: Implements the remaining 32-bit syscalls
The main one that we can't implement is readdir. Falls in to the same
problem space as getdents.

With this, we have the full entry tables filled out for both 64-bit and
32-bit. With some holes in the implementation, we have almost all
coverage now.
2021-12-27 02:19:30 -08:00
Ryan Houdek 514b6a822f Linux: Oops, forgot to pass usig to io_pggetevents
Easy enough to fix
2021-12-27 02:17:52 -08:00
Ryan Houdek 0945c728e9 Linux: Oops, missed fsmount syscall implementation
Simple fix
2021-12-27 02:17:20 -08:00
Ryan Houdek 3a3e2776ba Adds Game bug issue template 2021-12-25 12:29:59 -08:00
Ryan Houdek b72245a2d5 Merge pull request #1469 from Sonicadvance1/remove_0x
FEXRootFSFetcher: Remove 0x prefix on hash
2021-12-24 20:46:43 -08:00
Ryan Houdek 9c49bd3c8e FEXRootFSFetcher: Remove 0x prefix on hash
So as to not confuse users who just copy and paste the hash without
thinking
2021-12-24 20:33:23 -08:00
Ryan Houdek c691d70919 Merge pull request #1467 from Sonicadvance1/fix_really_old_ubuntu
Improves compile ability for older libraries
2021-12-24 18:50:39 -08:00
Ryan Houdek f3a27a57f1 Merge pull request #1468 from Sonicadvance1/experimental_libcxx
CMake: Adds experimental libc++ option
2021-12-24 18:50:31 -08:00
Ryan Houdek edce981824 Merge pull request #1463 from Sonicadvance1/enable_threads_default
Config: Enables all host threads by default
2021-12-24 18:50:20 -08:00
Ryan Houdek 9bffaeea40 Merge pull request #1462 from Sonicadvance1/sanitize_core_option
Config: Sanitize Core option
2021-12-24 18:50:11 -08:00
Ryan Houdek 88ce9b5cd2 Merge pull request #1461 from Sonicadvance1/fix_gdb_map
GDBServer: Fixes long string packet encodings
2021-12-24 18:50:01 -08:00
Ryan Houdek 72e8a997f6 Merge pull request #1460 from Sonicadvance1/FEXRootFSFetcher
FEXRootFSFetcher: Adds a new tool to help set up a new RootFS
2021-12-24 18:49:50 -08:00
Ryan Houdek d708cbad5e Config: Always fixup Threads option
In the case of nothing being set then with the default being zero now we
will not have fixed it up to calculate the number of threads based on
host core count.
2021-12-24 16:32:58 -08:00
Ryan Houdek 152eaff00b CMake: Adds experimental libc++ option
FEX currently doesn't compile with it mainly because libc++ doesn't
support c++ pmr at all.
2021-12-24 15:46:58 -08:00
Ryan Houdek 2079f6b3c7 Improves compile ability for older libraries
Adds a header only include utility folder that can be included from
everywhere.

Contains syscall helpers for older glibc and defines for older Linux
uapi headers missing some defines.
2021-12-24 15:43:25 -08:00
Ryan Houdek 2c31080fd4 FEXRootFSFetcher: Adds a new tool to help set up a new RootFS
This tool supports both a zenity and tty interface.
TTY will be presented if available while Zenity will be used otherwise.

This tool pulls a rootfs list from https://rootfs.fex-emu.org/
It then allows you to select a rootfs from the list, download it, place
it in to the correct working folder, extract it if desired, and set it
as the current default RootFS.

This requires curl and potentially zenity installed to use.
Maybe also unsquashfs if the user chooses to extract the image.

This is an all in one tool to quickly get a new user up and running.

Additionally if you pass in a file path in to the tool, it will generate
an xxhash of the file and exit out. This is the hash used to ensure the
files are valid.
2021-12-24 13:38:20 -08:00
Ryan Houdek ef7b77dff7 Config: Enables all host threads by default
We've solved the few threading problems that we had before. So now allow
FEX to query the host for the number of threads.
2021-12-24 13:22:01 -08:00
Ryan Houdek 777aadb73e Config: Sanitize Core option
If set to an invalid Core option from the json file then sanitize it
back to JIT.

Otherwise FEX has a chance of just crashing.
2021-12-24 13:21:47 -08:00
Ryan Houdek ccd06e2097 GDBServer: Fixes long string packet encodings
Fixes thread, memory-map, and OS data packet types.
These were attempting to substr when the encode function already handles
that.
Was making it so gdb was only ever receiving the first 1000 bytes of the
data and then decoding incorrectly.
2021-12-24 13:21:35 -08:00
Ryan Houdek 09a5f8c6b5 FEX: Move common config loading to a helper 2021-12-24 13:21:24 -08:00
Ryan Houdek 8f835678f5 Merge pull request #1466 from Sonicadvance1/fix_old_ubuntu3
Linux: Fixes build on Ubuntu 20.04 take 3
2021-12-24 13:20:45 -08:00
Ryan Houdek 8f2cd39802 Linux: Fixes build on Ubuntu 20.04 take 3
Fixes #1457
Header is necessary for x64 types as well for struct verifier
2021-12-24 13:05:32 -08:00
Ryan Houdek 059e8a9099 Merge pull request #1465 from Sonicadvance1/fix_old_ubuntu2
Linux: Fixes build on Ubuntu 20.04 take 2
2021-12-24 12:40:34 -08:00
Ryan Houdek 5fae07f6ee Linux: Fixes build on Ubuntu 20.04 take 2
Fixes #1457

Header include order matters here. Rude.
2021-12-24 12:23:52 -08:00
Ryan Houdek 02005e71ca Merge pull request #1464 from Sonicadvance1/fix_old_ubuntu
Linux: Fixes build on Ubuntu 20.04
2021-12-24 12:11:49 -08:00
Ryan Houdek 59cc5228b8 Linux: Fixes build on Ubuntu 20.04
Fixes #1457
2021-12-24 02:46:29 -08:00
Ryan Houdek 117cbde226 Merge pull request #1455 from Sonicadvance1/detect_fpcr_exceptions
HostFeatures: Detect if the host CPU suports float exceptions
2021-12-23 13:42:08 -08:00
Ryan Houdek 565d1e27d7 Merge pull request #1456 from Sonicadvance1/fix_sized_constant
IR: Fixes sized constant mask
2021-12-18 23:31:33 -08:00
Ryan Houdek 22ecff05c9 IR: Fixes sized constant mask
This mask was subtracting backwards, which made anything smaller than 8
byte constants not mask correctly
2021-12-18 23:14:55 -08:00
Ryan Houdek 17480c0e2d HostFeatures: Detect if the host CPU suports float exceptions
On x86 this is always supported.
On ARM this is only supported if FPCR writes actually enable the things.
Also detects the AFP feature for flushing input denormals to zero.

These are all part of the x86 MXCSR.
No Cortex supports FPCR exceptions, while Apple M1 CPUs support
Exceptions but not the true "AFP" extension
Apple instead supports some additional flags in their
`SYS_APL_AFPCR_EL0` register for enabling this.
2021-12-17 16:12:31 -08:00
Ryan Houdek 0608d95322 Updates externals vixl 2021-12-17 16:11:40 -08:00
Ryan Houdek 47bd47950b Merge pull request #1454 from Sonicadvance1/fix_aarch64_ubuntu_20.04
Linux: Fixes Ubuntu 20.04 compilation on AArch64
2021-12-16 22:41:05 -08:00
Ryan Houdek 9aab7e07f4 Linux: Fixes Ubuntu 20.04 compilation on AArch64
This header was missing
2021-12-16 22:23:32 -08:00
Mai M f2fd9d9e3f Merge pull request #1451 from Sonicadvance1/fix_semid_ds
Linux: Fixes semid_ds_64 definition
2021-12-17 01:19:41 -05:00
Mai M 45430e9569 Merge pull request #1452 from Sonicadvance1/fix_32bit_shmctl
Linux: Fixes 32-bit shmctl
2021-12-17 01:19:30 -05:00
Mai M be3e3a351a Merge pull request #1453 from Sonicadvance1/fix_cpuid_crash
CPUID: Fixes crash on unknown CPU
2021-12-17 01:19:19 -05:00
Ryan Houdek 2cee9e5d4b CPUID: Fixes crash on unknown CPU
If the CPU is unknown inside of the ARM CPU detection then the
MIDROption selected could have fallen down a path where it is set to
nullptr.

Resolve this crash by doing a nullptr check.
2021-12-16 21:50:42 -08:00
Ryan Houdek 03cc35f341 Linux: Fixes 32-bit shmctl
We were dereferencing the shmun ptr instead of using it directly.
Resulting in an almost immediate crash

Additionally IPC_SET doesn't write back to the shmid_ds provided.

Additional still SHM_STATE/SHM_STAT_ANY/IPC_STAT wasn't writing its
result back to the guest. Resulting in invalid stat information for the
guest.

Additionally SHM_INFO doesn't follow IPC64 behaviour since there is no
shm_info 64-bit type for a 32-bit OS.
The kernel just truncates the results in this case. Could be an
oversight on the kernel dev's side?

Fixes #1252
2021-12-15 23:34:15 -08:00
Ryan Houdek 3aead5ef65 Linux: Move some 32-bit ipc types to Types.h
This way we can run the struct verifier over it
2021-12-15 23:29:08 -08:00
Ryan Houdek b9abd091d5 Linux: Fixes semid_ds_64 definition
Wasn't quite the expected definition.

Fixes #1450
2021-12-15 22:02:37 -08:00
Ryan Houdek 43dc232e5a Merge pull request #1442 from Sonicadvance1/more_aot_changes
More AOT code movement
2021-12-15 03:27:01 -08:00
Ryan Houdek ea73c9d7ea Merge pull request #1443 from Sonicadvance1/squashfs_fixes
FEXMountDaemon squashfs fixes.
2021-12-15 03:04:15 -08:00
Ryan Houdek fa6f1b1d90 Merge pull request #1446 from Sonicadvance1/more_fault_reconstruction
Dispatcher: Adds more state reconstruction to state restore
2021-12-15 03:04:09 -08:00
Ryan Houdek e24eb7a72d Merge pull request #1447 from Sonicadvance1/consolidate_host_features
HostFeatures: Consolidates HostFeatures flags
2021-12-15 02:55:42 -08:00
Ryan Houdek e05d116b05 Merge pull request #1448 from Sonicadvance1/fix_proton_experimental
Linux: Adds the safe syscall unimplemented gap for x86-64
2021-12-15 02:55:35 -08:00
Ryan Houdek ee821b9cbf Merge pull request #1449 from Sonicadvance1/implement_rdtscp
Implements RDTSCP
2021-12-15 02:52:47 -08:00
Ryan Houdek c29e563836 unittests: Adds tests for RDTSCP 2021-12-15 01:21:16 -08:00
Ryan Houdek fd59fb1a7a OpcodeDispatcher: Implements support for RDTSCP
Have fun
2021-12-15 01:18:34 -08:00
Ryan Houdek 04deeb3911 IR: Adds Processor ID IR op
For the x86-64 JIT this is implemented with pulling rdtscp's result for
this value.
For Interpreter and AArch64 JIT this is implemented with the getcpu
syscall.

Theoretically AArch64 could implement this with MPIDR_EL1 but because
SoC vendors hecked this up, we can't. Thanks.
Kernel just returns zero + reserved bits if you try reading it.
2021-12-15 01:17:24 -08:00
Ryan Houdek 4a7c65b20d Linux: Adds the safe syscall unimplemented gap for x86-64
Syscalls on x86-64 have a gap in the range of [335, 424) where these are
defined as returning ENOSYS.
This was a decision that the Linux developers decided on so x86-64 can
realign its syscall numbers to a common infrastructure.

Turns out that Proton Experimental was using a 32-bit only syscall to
check if the feature existed and was listening for ENOSYS.
They did this unconditionally on both x86 and x86-64 and it hits this
gap section.

This was introduced in Proton Experimental with commit
563fb0fbe2a49734aede87bedccb20daefc553b0
Message: "ntdll: Use clock_gettime64 if supported."

This is technically valid since it falls within this gap section but
leaves a bad taste.

We can safely ignore this section so set it up with the
UnimplementedSyscallSafe handler and also assign it to Syscall MAX so it
can go down our fast handler.

Not that this is likely to be a performance critical path.

Fixes Proton Experimental crashing when run under FEX.
2021-12-14 03:45:55 -08:00
Ryan Houdek 23cb0dea00 HostFeatures: Consolidates HostFeatures flags
Some of these were in the Emitter class and some were in the
HostFeatures.

Merge these together since in the future I'm going to be using all of
this data as a key for our AOT code cache.
2021-12-14 02:41:42 -08:00
Ryan Houdek bcad7a9eea Dispatcher: Adds more state reconstruction to state restore
Still only setting the new state if RIP is affected for now.
Noticed a bug where we weren't setting our frame RIP to the new RIP on
32-bit.
Decided to walk through more of the state setting while fixing that.

This gets #1214 further but then it eventually crashes with a read to
0x11.
2021-12-14 02:36:24 -08:00
Ryan Houdek 0a3a270a63 Merge pull request #1445 from lioncash/pdep
OpcodeDispatcher: Handle PDEP
2021-12-13 16:43:35 -08:00
lioncash 06e4a5a5b7 CPUID: Signify full support for BMI2
With PDEP support dropped in, we now support all of BMI2, so we can
signify that we support it in our emulated CPUID.
2021-12-13 14:07:41 -05:00
lioncash 6ff80670b3 OpcodeDispatcher: Handle PDEP
Now all of BMI2 is handled.
2021-12-13 14:06:55 -05:00
lioncash 38eea80b8d IR: Add PDep IR opcode 2021-12-13 13:54:15 -05:00
Ryan Houdek c43af0e10f Merge pull request #1444 from neobrain/fix_jemalloc_overrides
Update jemalloc for fixed libc overrides
2021-12-13 04:48:05 -08:00
Tony Wasserka 2fa8cd7e97 Update jemalloc for fixed libc overrides 2021-12-13 13:38:39 +01:00
Ryan Houdek 5d0734a7f2 FEXMountDaemon: Attempt to increase FD max when close to limit
On FEXMountDaemon startup, fetch the FD limit and current number of open
files. This allows us to track the number of open FDs we have.

Once we get close the the safe threshold we then try increasing the soft
limit towards the hard limit. If we bump up to the max (Say someone
setting the hard limit to 256) then print some fairly big warning
messages.

With the previously fixed FD leak, we could very quickly hit the pipe
limit because Steam is constantly opening new processes all the time.
With the leak fixed, running Steam and idling, FEXMountDaemon only has
around 16 pipes being watched.
2021-12-12 22:51:53 -08:00
Ryan Houdek 3e0e922fef FEXMountDaemon: Fixes EPoll thread not shutting down on SIGTERM
This is another edge case where if the process was sent a SIGTERM then
it can get in to a weird state where pipe tracking is broken.

We already check for zero pipes being tracked at the end of this loop,
just check for the shutdown variable to be set instead
2021-12-12 22:51:15 -08:00
Ryan Houdek e2004f4999 FEXMountDaemon: Fix leaked pipe FDs
Even though we were removing the pipe from the epoll interest list,
we were failing to close the pipe fd. This was leaking the FD and
running in to the Linux process FD limitation over time.

This fixes a class of issues where if you were running FEXMountDaemon
that hit the max, then there were FEX processes not being watched. Which
means that FEXMountDaemon could exit while FEX was still trying to run.

Usually a non-issue since squashfuse would keep running in the
background since open files in the mount were still in use.

Could hit fun race conditions because of it though.
2021-12-12 22:45:40 -08:00
Ryan Houdek 13f37efe12 SquashFS: Work around a race condition on FEXMountDaemon ack
On an overburdened system then FEXMountDaemon may notrespond within the two
seconds of FEX waiting for an ack.
In this case then we hit a bad edge case where we overwrite the lock
file with a new FEXMountDaemon instance then spin up a new
FEXMountDaemon.
This new FEXMountDaemon will recognize that one is still running and
early exit, but now that we've overwritten the lock file with a new
location, all future FEX instances won't be able to find the squashfs
rootfs.

Instead, ignore that the ack wasn't received and attempt running
instead.

FEXMountDaemon should still pick up the pipe message and monitor FEX
regardless.

This fixes the `/usr not found` messages and crashes resulting from it.
2021-12-12 22:44:48 -08:00
Ryan Houdek 80e66a36eb FEXConfig: Adds AOT options to the UI 2021-12-12 18:22:02 -08:00
Ryan Houdek 4998d35ec5 FEXLoader: Moves AOT handling to its own independent files
This is only moving it outside of FEXLoader to make it easier to work
on.

I'll be hammering on this more soon.
2021-12-12 18:19:12 -08:00
Ryan Houdek f0655874e3 ELFCodeLoader: Fixes bug in ELF section tracking
Each section map invocation was clearing the sections pointer, meaning
we only ever had one section in the vector.

We want to keep all of the sections around from the ELFCodeLoader.
Fixes a bug where we were effectively never catching any code segments
from the executable passed in.
2021-12-12 18:15:06 -08:00
Ryan Houdek 9394e49c95 Merge pull request #1431 from Sonicadvance1/expose_arm_names
CPUID Expose Hybrid flag and CPU names
2021-12-12 18:07:53 -08:00
Ryan Houdek 91984003b6 Merge pull request #1440 from lioncash/pext
OpcodeDispatcher: Handle PEXT
2021-12-10 16:27:49 -08:00
lioncash dbf571fdfb OpcodeDispatcher: Handle PEXT 2021-12-10 19:03:58 -05:00
lioncash b6abcc5e3c IR: Add PExt opcode 2021-12-10 19:03:54 -05:00
Ryan Houdek e99a23cfc6 Merge pull request #1439 from neobrain/opt_x86tables_init
Speed up initialization of X86Tables
2021-12-10 09:11:47 -08:00
Ryan Houdek 425d9323d5 Merge pull request #1295 from neobrain/feature_new_thunk_gen
Implement thunk library generation using libclang
2021-12-10 09:10:00 -08:00
Tony Wasserka a7c0997daf OpcodeDispatcher: Mark const-initialized instruction tables as constexpr
This allows the compiler back these tables into the executable, which
reduces the amount of work the function has to do at runtime.

Reduces the runtime of this function by 30% relative to the previous commit.
2021-12-10 13:25:24 +01:00
Tony Wasserka 6eba3f331e OpcodeDispatcher: Store instruction tables as arrays instead of std::vector
Reduces the runtime of this function by 90%.
2021-12-10 13:25:23 +01:00
Tony Wasserka bcc75e3312 X86Tables: Use stack-allocated arrays instead of std::vector 2021-12-10 13:25:23 +01:00
Tony Wasserka 30672e1517 X86Tables: Remove unneeded initialization code from the Debug mode path
The tables have recently been changed to be zeroed out as a whole on startup.

Speeds up InstallDebugInfo by about two orders of magnitude and reduces Debug
executable size by 1.8%.
2021-12-10 13:25:23 +01:00
Tony Wasserka b6e46fd44d Thunks: Update documentation to reflect generator changes 2021-12-10 11:25:01 +01:00
Tony Wasserka 32a0e37569 Thunks: Remove now unneeded generator files 2021-12-10 11:25:01 +01:00
Tony Wasserka 5b6175b702 Remove now unused Vulkan-Docs submodule 2021-12-10 11:25:01 +01:00
Tony Wasserka 0285e35c87 Thunks: Use libclang-based code generation for libvulkan 2021-12-10 11:25:01 +01:00
Tony Wasserka bbfb8713a7 Thunks: Drop struct verifier tests for Vulkan
This is currently not supported with the new generator.
2021-12-10 11:25:00 +01:00
Tony Wasserka c1d6967cb6 Thunks/gen: Add support for expanding symbol lists using a macro 2021-12-10 11:25:00 +01:00
Tony Wasserka 20e24eae9c Thunks: Use libclang-based code generation for libdrm 2021-12-10 11:25:00 +01:00
Tony Wasserka b8d9027680 Thunks: Use libclang-based code generation for libxshmfence 2021-12-10 11:25:00 +01:00
Tony Wasserka 688ef9f5af Thunks: Use libclang-based code generation for libxcb-xfixes 2021-12-10 11:25:00 +01:00
Tony Wasserka 181c9074ab Thunks: Use libclang-based code generation for libxcb-sync 2021-12-10 11:25:00 +01:00
Tony Wasserka 6f85f64bfd Thunks: Use libclang-based code generation for libxcb-shm 2021-12-10 11:25:00 +01:00
Tony Wasserka ba6ee61db2 Thunks: Use libclang-based code generation for libxcb-randr 2021-12-10 11:25:00 +01:00
Tony Wasserka e735281b2c Thunks: Use libclang-based code generation for libxcb-present 2021-12-10 11:24:59 +01:00
Tony Wasserka c8b2d714a8 Thunks: Use libclang-based code generation for libxcb-glx 2021-12-10 11:24:59 +01:00
Tony Wasserka bba0625a57 Thunks: Use libclang-based code generation for libxcb-dri3 2021-12-10 11:24:59 +01:00
Tony Wasserka c9a7b36210 Thunks: Use libclang-based code generation for libxcb-dri2 2021-12-10 11:24:59 +01:00
Tony Wasserka 99bf02f27f Thunks/gen: Add support for libraries with "-" in the filename 2021-12-10 11:24:59 +01:00
Tony Wasserka a2c02ae51e Thunks: Use libclang-based code generation for libxcb
Notably, xcb_take_socket uses the callback_guest annotation, since the
given callback is never called on the host (instead it's manually forwarded
back to a helper guest thread for calling).
2021-12-10 11:24:59 +01:00
Tony Wasserka c7b59143c6 Thunks/gen: Add custom guest entrypoint annotation 2021-12-10 11:24:59 +01:00
Tony Wasserka 1e23e61572 Thunks: Use libclang-based code generation for libEGL 2021-12-10 11:24:59 +01:00
Tony Wasserka f4dd1a895e Thunks: Use libclang-based code generation for libGL 2021-12-10 11:24:58 +01:00
Tony Wasserka 3a2f9a4d46 Thunks/gen: Support guest-side symbol tables and custom host symbol loaders 2021-12-10 11:24:58 +01:00
Tony Wasserka 4a848d7202 Thunks: Use libclang-based code generation for libX11 2021-12-10 11:24:58 +01:00
Tony Wasserka 1b58ed9f57 Thunks/X11: Fix incorrect library version 2021-12-10 11:24:58 +01:00
Tony Wasserka 1c86f7ed36 Thunks/gen: Support annotating guest-exclusive function pointers
Function pointer arguments given to the guest thunk library aren't
callable on the host, so they need special handling on a case-by-case
basis. Similar problems arise when the guest-side tries to consume a
pointer returned from a native host library. To automate some common
scenarios, this change adds two new annotations:

"callback_guest" indicates the callback parameter is never called on the
host and hence can be marshalled like any other argument. Accidental host
calls to the function pointer are prevented by wrapping it in an opaque
type alias.

"returns_guest_pointer" indicates that the host function returns a pointer
usable in the guest context. This applies e.g. to functions that derive
the returned pointer from input guest pointer arguments.
2021-12-10 11:24:58 +01:00
Tony Wasserka 2733b2ee1e Thunks/gen: Add annotation for custom host implementations 2021-12-10 11:24:58 +01:00
Tony Wasserka 4ceb2dfdf2 Thunks: Use libclang-based code generation for libXext 2021-12-10 11:24:58 +01:00
Tony Wasserka bd380e0f15 Thunks: Use libclang-based code generation for libXrender 2021-12-10 11:24:58 +01:00
Tony Wasserka 73ec786c60 Thunks: Use libclang-based code generation for libasound 2021-12-10 11:24:57 +01:00
Tony Wasserka e3a2c8dc80 Thunks: Use libclang-based code generation for libXfixes 2021-12-10 11:24:57 +01:00
Tony Wasserka d3b14df840 Thunks/gen: Add versioning support 2021-12-10 11:24:57 +01:00
Tony Wasserka a12ab8f98a Thunks/gen: Add support for "callback_stub" annotations
Some applications set callbacks that never get called in practice (such as
error handlers). It's sensible to just not implement these instead of
cluttering the code with effectively unused callback wrappers.
2021-12-10 11:24:57 +01:00
Tony Wasserka 966b9a69d8 Thunks/gen: Add support for function pointer parameters ("callbacks") 2021-12-10 11:24:57 +01:00
Tony Wasserka 927d3d00e2 Thunks/gen: Add support for thunking variadic functions 2021-12-10 11:24:57 +01:00
Tony Wasserka bfb9cabeb8 Thunks/gen: Add support for generation of host library functions 2021-12-10 11:24:57 +01:00
Tony Wasserka 22466a973c Thunks/gen: Add unit tests 2021-12-10 11:24:57 +01:00
Tony Wasserka c05e1c9797 Thunks: Add a new code generator based on libclang 2021-12-10 11:24:56 +01:00
Ryan Houdek c65be9f55d Merge pull request #1438 from Sonicadvance1/remove_jemalloc_from_fexconfig
FEXConfig: Removes jemalloc usage
2021-12-09 18:10:32 -08:00
Ryan Houdek edc31fe0a3 FEXConfig: Removes jemalloc usage
Splits out the few required dependencies to a FEXCore_Base static
library.

FEXCore then links to this directly.

Then make it so the FEX Common code links to FEXCore_Base so the
jemalloc dependency doesn't get pulled in.
2021-12-09 17:54:53 -08:00
Ryan Houdek d4655fbb17 Merge pull request #1424 from Sonicadvance1/aot_code_movement
FEXCore: Reorganizes some AOT related code
2021-12-09 13:55:50 -08:00
Ryan Houdek 5f0dfcd715 Merge pull request #1422 from Sonicadvance1/implement_clzero
Implements CLZero instruction
2021-12-09 13:55:31 -08:00
Ryan Houdek e62cb417c2 Merge pull request #1437 from Sonicadvance1/disable_flake_gvisor
gvisor: Disables flaky test
2021-12-09 12:58:29 -08:00
Ryan Houdek df486c0786 gvisor: Disables flaky test 2021-12-09 12:37:58 -08:00
Ryan Houdek 9d43904792 Merge pull request #1436 from lioncash/context-const
Context: Take some arguments as pointer-to-const
2021-12-09 12:23:34 -08:00
Ryan Houdek 5654f9a030 Merge pull request #1435 from Sonicadvance1/fix_vixl_assertions
Arm64: Fixes vixl assertions around ubfm usage
2021-12-09 11:54:41 -08:00
Ryan Houdek d74cf6d8d8 CPUID Expose Hybrid flag and CPU names
Had some idle time so I implemented this logic.

We do some tricky logic to have a big.little configuration even with
unknown CPU core types. Promoting or demoting a single MIDR depending on
if we have a mixed configuration or not.

In a non-hybrid design we only claim product names inside the CPUID
product string.

This will appear if you `/proc/cpuinfo` or read the CPUID registers
directly

eg on Snapdragon 888:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 1
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 2
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 3
model name      : FEX-2112-1-g13b14b85            Cortex-A55
processor       : 4
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 5
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 6
model name      : FEX-2112-1-g13b14b85            Cortex-A78
processor       : 7
model name      : FEX-2112-1-g13b14b85            Cortex-X1

eg on Macbook Pro VM which can't see the CPU type:
processor       : 0
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 1
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 2
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 3
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 4
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 5
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 6
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
processor       : 7
model name      : FEX-2112-1-g13b14b85            Unknown ARM CPU
2021-12-09 11:39:52 -08:00
Ryan Houdek 84905d2856 unittests: Adds 32-bit inline syscall test
Triggers the ubfm bug on AArch64.
Uses a fairly benign syscall that gets inlined
2021-12-09 11:34:35 -08:00
Ryan Houdek 92cb9d477e Arm64: Fixes vixl assertions around ubfm usage
vixl has an assert check to ensure the register sizes are the same for
ubfm.

32-bit Inline syscalls hit this for both arguments and return value.
VCastFromGPR would have hit this but there aren't any x86 instructions
that move 8-bit and 16-bit values in to a vector register.
2021-12-09 11:34:20 -08:00
lioncash 7ed6007252 Context: std::move functions in signal registration functions
Prevents potential reallocations, given we're using std::function here.
2021-12-09 14:31:47 -05:00
Ryan Houdek b6499ac724 Merge pull request #1434 from lioncash/syscallthread
{x32, x64}/Thread: Make use of .data() instead of .at(0)
2021-12-09 11:28:39 -08:00
lioncash c73991f467 Context: Take some arguments as pointer-to-const
Several API functions act as state querying functions. These can take
some parameters by const to communicate that we don't intend to modify
the respective passed in instance.
2021-12-09 14:26:35 -05:00
lioncash 0bf0fe4779 {x32, x64}/Thread: Make use of .data() instead of .at(0)
Follows the previous PRs of using .data() instead of .at(0) to retrieve
the buffer pointers.
2021-12-09 14:14:40 -05:00
Ryan Houdek 097f48f3ff Merge pull request #1433 from lioncash/syscallmsg
x32/Socket: Resolve sign comparison mismatch cases
2021-12-09 11:09:07 -08:00
Ryan Houdek 75e5df545a Merge pull request #1432 from lioncash/syscall-warn
LinuxSyscalls: Enable warnings
2021-12-09 10:51:21 -08:00
lioncash 171d5f7263 x32/Socket: Use .data() instead of .at(0) for pointer retrieval
.at(0) can blow up in the case we receive a message from user code with msg_iovlen=0
.data() will simply return a pointer to the empty buffer and we can
proceed with forwarding the call to the syscall.
2021-12-09 13:49:15 -05:00
lioncash 970067d19b x32/Socket: Resolve sign comparison mismatch cases 2021-12-09 13:43:26 -05:00
lioncash c26ff60949 LinuxSyscalls: Enable warnings
Enables some basic warnings that we have enabled for FEXCore in order to
prevent some small things from slipping through.
2021-12-09 12:58:54 -05:00
lioncash 04d830ed39 LinuxSyscalls: Minor tidy up of CMakeLists 2021-12-09 12:46:39 -05:00
Ryan Houdek b9c49027c7 Merge pull request #1409 from Sonicadvance1/reentrantmutex
Adds a new ReentrantMutex to use for FEXCore
2021-12-08 14:00:12 -08:00
Ryan Houdek 9860e8b71b Merge pull request #1430 from lioncash/net
NetStream: Move NetBuf definition into cpp file
2021-12-08 13:10:06 -08:00
lioncash a1e94a9863 NetStream: Move NetBuf definition into cpp file
Keeps the NetBuf class completely internal and also lessens the header
dependencies for NetStream.
2021-12-08 10:41:55 -05:00
lioncash 2945c13dcb NetStream: Mark virtual functions as override
Makes it explicit that we're overriding the interface of the class being
derived from.
2021-12-08 10:33:18 -05:00
lioncash 4f68821aef NetStream: Mark constructors as explicit
Prevents implicit conversion of ints to NetStreams.
2021-12-08 10:30:38 -05:00
Ryan Houdek 13087f8425 Merge pull request #1426 from Sonicadvance1/fcmp_unittests
Arm64: Fixes MapSelectCC FGT flag
2021-12-08 05:26:15 -08:00
Ryan Houdek fd0424768b Merge pull request #1427 from Sonicadvance1/enable_unaligned_atomic_tests
unittests: Enables remaining unaligned atomic tests on ARMv8.0
2021-12-08 05:26:03 -08:00
Ryan Houdek 5c4112f103 Merge pull request #1428 from Sonicadvance1/fix_known_failures
unittests: Fixes known failures
2021-12-08 05:25:50 -08:00
Ryan Houdek e8670bbab2 Merge pull request #1425 from Sonicadvance1/x86tables_unknown
X86Tables: Build Unknown op definition tables at compile time
2021-12-08 00:22:57 -08:00
Ryan Houdek 13b14b857b unittests: Fixes known failures
These were just incorrect results and I had failed to fix them.
2021-12-07 22:06:36 -08:00
Ryan Houdek 48c7ff2a23 unittests: Enables remaining unaligned atomic tests on ARMv8.0 2021-12-07 21:54:03 -08:00
Ryan Houdek 77a032a286 unittests: Adds tests for handling cmp merging
As a continuation to #1404 these unit tests were needing to be added.
This tests both select and conditional branch merging for integer ops
and float ops.

This is what discovered the issue fixed in f2de640395
2021-12-07 18:45:54 -08:00
Ryan Houdek f2de640395 Arm64: Fixes MapSelectCC FGT flag
As a continuation to #1404, this was mapped to the incorrect Arm64 flags
2021-12-07 18:43:50 -08:00
Ryan Houdek 9a642158e0 Workaround fmt not handling nullptr strings
Simple ternary check each time a name is used
2021-12-07 12:23:56 -08:00
Ryan Houdek dc44caa178 Github: Adds the new APITests to the github workflow 2021-12-07 11:55:26 -08:00
Ryan Houdek 081a003c6c unittests: Adds a new catch2 test for testing the InterruptableConditionVariable 2021-12-07 11:55:26 -08:00
Ryan Houdek 4c5fc6e813 FEXCore: Uses the InterruptableConditionVariable for StartRunning event
This can be used with the gdbstub and on pause could cause longjumps
out of the event.
Resolves the hang on InternalThreadState destruction.
2021-12-07 11:55:26 -08:00
Ryan Houdek 253cdb552d FEXCore/Utils: Adds InterruptableConditionVariable class
Mutex destruction is not safe in the face of longjump and signals.
Adds a new class that is very simple and will still work in this
instance.
2021-12-07 11:55:26 -08:00
Ryan Houdek 6404aba6e2 X86Tables: Build Unknown op definition tables at compile time
It doesn't make any sense anymore to have specific instruction names
set for UND versus a nullptr string anymore.

Was useful when we could use it to determine the difference between
undefined from the start versus set in the tables but with unknown
decoding. Which is an edge case.

Now instead just zero initialize the data, which means it is an unknown
type and nullptr name. Which works for use.

Improves initialization time of the InitializeInfoTables function from
423 microseconds to 37 microseconds.
2021-12-06 18:57:02 -08:00
Ryan Houdek 45f919683a FEXCore: Reorganizes some AOT related code
Specifically this tries to avoid changing much behaviour and keeping the
code the same. So most of it is a direct transplant without any
modifications. This is step one of the process so I can start logically
separating the code and making sense of it.

This mostly moves the AOT IR handling to its own independent file for
separation. Cleaning up the Core.cpp file quite heavily.

Two minor behaviour changes that got mixed up with this change.

The first one is an ASAN fix.
This is the FEX_PACKED on the RegisterAllocationData class.
I didn't want to change too heavily how this serialization works but I
wanted to resolve the ASAN error. This may change in the coming work.
Problem was the padding betwene the uint32_t and the PhysicalRegister
wasn't initialized but was being read.
Since it is all uint8_t types afterwards there isn't a perf issue here.

Second fix was a crash that occurs if you're attempting to both capture
and load IR on the same run. This is a quirk where we mmap the original
IR file. Then on shutdown the IR file is getting saved.
At which point we open the IR file again, truncate it, and start
serializing all of the IR data.
The truncation makes it so our mmap of the file is no longer resident,
resulting in a crash when reading our IR cache from the mmap region.
Now open a temporary file and rename it after storing.
Resolves the crash but still doesn't really solve the issue of multiple
processes overwriting the same IR files.
2021-12-06 18:42:21 -08:00
Ryan Houdek c3a6890a40 CMake: Adds Catch build when tests are enabled 2021-12-06 18:27:12 -08:00
Tony Wasserka 07be5a0bae Add Catch2 submodule 2021-12-06 18:27:11 -08:00
Ryan Houdek d0943322b8 Docs: Update for release FEX-2112 2021-12-05 17:44:04 -08:00
Ryan Houdek d6e4da7e77 CPUID: Exposes support for CLZero
Only exposed if the host if the HostFeatures for ARM claim to support
it.
2021-12-03 21:20:40 -08:00
Ryan Houdek ed5d7f6a62 unittests: Adds CLZero unit tests
First test just ensures that an aligned cacheline clear will clear out
the data.
Second test ensures an unaligned address will still only clear the
cacheline the address lies in without touching the other surrounding
cachelines.

Disabled on the host runner since the CI host runner doesn't support the
instruction
2021-12-03 21:19:14 -08:00
Ryan Houdek 18b223811a OpcodeDispatcher: Implements CLZero instruction
This instruction zeroes a cacheline in memory that is weakly ordered and
non-temporal.

It uses the RAX register for where in memory to clear and aligns the
address on cacheline regardless of actual alignment.
2021-12-03 21:15:51 -08:00
Ryan Houdek e1d21f0bff IR: Implements new CacheLineZero op
This zeroes out an emulated 64byte cacheline. Writing zeros to memory.
This very specifically is only 64bytes to match x86 behaviour.
Also specifically non-temporal and weakly ordered. Which matches x86
CLZero behaviour.
2021-12-03 21:14:02 -08:00
Ryan Houdek d2783f2edd HostFeatures: Adds new host feature flag for CLZero
If the DCZID block size matches the CLZero cacheline size. Then claim we
support CLZero here.
Otherwise we don't want to expose support for it.
2021-12-03 21:12:52 -08:00
Ryan Houdek 3377e5a50e HarnessHelpers: Fixes accidental fmt::format
Was meant to be fmt::print. Happened with 62dfccc989

oops
2021-12-03 21:01:59 -08:00
Ryan Houdek db0740becc Merge pull request #1421 from lioncash/symbols
JitSymbols: Make use of fmt
2021-12-03 00:41:05 -08:00
lioncash ec8f077e60 JitSymbols: Take std::string_view instead of std::string
Now that we're using fmt, we can make the API itself non-allocating and
allow passing any kind of character buffer to it.
2021-12-03 02:49:30 -05:00
lioncash 2b190f1713 JitSymbols: Make HostAddr const
We can seamlessly format const pointers, so we can allow this in the
interface.
2021-12-03 02:49:30 -05:00
lioncash 69413d51bf JitSymbols: Make use of fmt
Makes the constructed strings a little quicker to read.
2021-12-03 02:49:26 -05:00
Ryan Houdek 8c38f936da Merge pull request #1420 from lioncash/const
IR: Make some interface functions accept pointers to const
2021-12-02 21:14:05 -08:00
lioncash e3fbe48f7c JitSymbols: Make use of unique_ptr
Ensures that the FILE pointer will always be handled.
2021-12-02 23:58:28 -05:00
lioncash c1acd9ad59 IR: Mark relevant API functions as [[nodiscard]] 2021-12-02 23:48:53 -05:00
lioncash 381d07707c IntrusiveIRList: Mark APIs [[nodiscard]] where applicable
Allows us to be more aggressively warned by the compiler in cases it's
definitely a bug to ignore the returned value.
2021-12-02 23:38:34 -05:00
lioncash a07e77b028 IntrusiveIRList: Remove some redundant reinterpret_casts
GetListData() and GetData() already return uintptr_t values, so we don't
need to cast these anymore.
2021-12-02 23:32:40 -05:00
lioncash 6b5ccb8244 IR: Make some interface functions accept pointers to const
Some interface functions are just querying state, so we can allow them
being used with const qualified data.
2021-12-02 23:32:36 -05:00
Ryan Houdek 597e4e1d57 Merge pull request #1419 from lioncash/ra-cast
RegisterAllocationPass: Resolve sign comparison mismatch in CalculateNodeInterference()
2021-12-02 20:22:08 -08:00
lioncash cdab71e612 RegisterAllocationPass: Resolve sign comparison mismatch in CalculateNodeInterference()
GetSSACount() returns a size_t, rather than a signed value.
2021-12-02 18:13:01 -05:00
Ryan Houdek 8ef0278338 Merge pull request #1418 from lioncash/branch
JIT: Eliminate redundant jump target map lookups
2021-12-02 14:45:44 -08:00
lioncash 98f54eab9c Arm64/BranchOps: Mark MapBranchCC as static
This isn't used outside of this translation unit, so we can make that
explicit.
2021-12-02 17:18:44 -05:00
Ryan Houdek 919e2bc3bd Merge pull request #1408 from Sonicadvance1/add_timetrace
CMake: Adds option to enable time-trace compile option
2021-12-02 14:14:38 -08:00
Ryan Houdek ffee22d3fd Merge pull request #1417 from lioncash/cpuid
CPUID: Call handler functions directly
2021-12-02 14:14:20 -08:00
lioncash e648e90c2e JIT: Eliminate redundant jump target map lookups
try_emplace() will return the existing entry in the map if it
exists already, so we don't need to perform a find() and then
try_emplace(). We can just use try_emplace() by itself.
2021-12-02 17:03:04 -05:00
Ryan Houdek b93871ff55 Merge pull request #1416 from lioncash/strong
IR: Convert NodeID into a strong type
2021-12-02 13:39:24 -08:00
lioncash 352f0a1133 CPUID: Call handler functions directly
Given the calls are all internal, we can store the function pointers
directly and call them with the this pointer instead of indirecting
through std::bind and std::function.
2021-12-02 16:34:43 -05:00
Ryan Houdek 6c93b6f775 Merge pull request #1412 from Sonicadvance1/socket_logger
Adds Socket logger and tool
2021-12-02 13:32:04 -08:00
lioncash f806577c46 RegisterAllocationPass: Use std::optional instead of magic value
Instead of using UINT32_MAX to communicate not found, we can return an
empty optional instance.
2021-12-02 15:17:19 -05:00
lioncash fac4022302 IR: Convert NodeID into a strong type
Converts the NodeID alias into a strongly-typed structure, which will
allow catching attempts to pass NodeIDs into incorrect APIs at
compile-time.
2021-12-02 15:04:57 -05:00
lioncash b5b3cb252a BucketList: Pass T to Next explicitly
Without this, the internal BucketList type will always be allocated with
T as a uint32_t, due to the default type for T, even if T is specified
differently in other code.
2021-12-02 14:55:44 -05:00
Ryan Houdek d5bf6414f3 Merge pull request #1413 from Sonicadvance1/fix_sched_getaffinity
Linux: Fixes sched_getaffinity
2021-11-30 18:06:04 -08:00
Ryan Houdek fc858ca0a5 Adds FEXLogServer
This tool works in both gui-less and gui modes.
FEXLogServer executed alone will open a socket on `localhost:8087` and
pass log output to stderr.

If passed the `-g` argument then it will open the IMGui UI which has
some more options.
It's fairly basic right now but eventually should allow filtering by
PIDs, TIDs, and log level type
2021-11-30 14:58:06 -08:00
Ryan Houdek 8ddf1ccbfe Linux: Fixes sched_getaffinity
Noticed an issue where some applications were thinking that applications
were running on a system with 512 CPU cores.
Specifically freedreno/turnip was trying to create 512 worker threads
which would cause my ARM devices to run out of memory.

Looks like sysconf behaviour has changed around this where it is
actually expecting the uint64_t alignment that the Linux kernel does.
So follow that behaviour as well and resolve the issue.
2021-11-30 14:53:10 -08:00
Ryan Houdek f2292c08e9 FEXLoader: Adds Socket logging path
Adds a new OutputSocket config option that when set will force all logs
to go through a socket.

This allows us to avoid polluting the guest application's output through
a socket instead of a file. Alleviating the issue of a file output
overwriting when multiple applications are ran.
2021-11-30 12:05:45 -08:00
Ryan Houdek b1b616087b Merge pull request #1380 from neobrain/fix_thunk_fails
Thunks: Fail loudly if thunking is enabled for a library that isn't installed
2021-11-30 11:45:11 -08:00
Tony Wasserka 5542360948 Thunks: Fail if thunking is enabled for a library that isn't installed
Previously, the host thunk would seemingly succeed to load even when
the corresponding host library could not be loaded, leading to obscure
crashes later throughout execution.
2021-11-30 17:39:01 +01:00
Tony Wasserka 7490d15369 Thunks: Clean up helper macro formatting 2021-11-30 17:39:01 +01:00
Tony Wasserka 68587cd02c unittests/ASM: Drop OP_THUNK test
This test doesn't increase coverage significantly, since OP_THUNK is called
with an invalid library name. An ideal test should verify that thunk symbols
are loaded properly, whereas currently it's only ensured the opcode 0xF 0x3F
is recognized by FEX at all. That's better than nothing, but a regression
here would likely show up in other tests anyway.
2021-11-30 17:39:01 +01:00
Ryan Houdek a9e1d5f2a2 FEX/Common: Adds a SocketLogging namespace
This will be used for sending messages over a socket
2021-11-29 13:14:33 -08:00
Ryan Houdek a8b9b351cb FEXConfig: Splits GUI setup in to a header
This will be used in another tool soon
2021-11-29 13:14:33 -08:00
Ryan Houdek 559f7128c3 Netstream: Use MSG_NOSIGNAL to avoid signals on socket
If the socket was closed then Linux would raise a EPIPE signal.
We don't want this, so avoid it with this flag
2021-11-29 13:14:33 -08:00
Ryan Houdek 5c1f14a155 FEXCore: Moves NetStream to Utils
The Frontend will want to use this
2021-11-29 13:14:32 -08:00
Ryan Houdek 82ebdf52be Merge pull request #1411 from Sonicadvance1/non_fata_32bit_mapping
FEXLoader: Change 32-bit memory check error to warning
2021-11-27 00:02:12 -08:00
Ryan Houdek 86684f9033 FEXLoader: Change 32-bit memory check error to warning
Instead of early exiting, allow the application to continue running but
throw error messages anyway. Should allow some users to still run FEX
even if an application steals a page in the lower 32-bits
2021-11-26 23:52:38 -08:00
Ryan Houdek e6d285a716 Merge pull request #1410 from Sonicadvance1/print_memory_map
FEXLoader: Print memory map when 32-bit intersect happens
2021-11-26 22:56:05 -08:00
Ryan Houdek 5c98c782f6 FEXLoader: Print memory map when 32-bit intersect happens
To get more information as to what intersected
2021-11-26 22:46:46 -08:00
Ryan Houdek 9b43f94f51 Merge pull request #1407 from lioncash/irgen
IR.json/json_ir_generator: Minor touchups to generated IR utilities
2021-11-26 15:01:00 -08:00
Ryan Houdek 19a0651944 Merge pull request #1406 from lioncash/regnode
RegisterAllocationPass: Reduce usages of NodeIDs where applicable
2021-11-26 14:57:08 -08:00
Ryan Houdek 436d20d660 CMake: Adds option to enable time-trace compile option
Useful for running something like aras-p/ClangBuildAnalyzer on the
source and finding compile bottlenecks
2021-11-26 14:43:39 -08:00
Ryan Houdek e027b83a06 Merge pull request #1399 from Sonicadvance1/gdbstub_improvements
GDBStub: Fixes a few hangs and crashes
2021-11-26 13:52:47 -08:00
lioncash d0aa785e6f json_ir_generator: Mark generated IR functions [[nodiscard]] where applicable
Marks generated functions as [[nodiscard]] where it would otherwise be a
logic bug to ignore their return values.
2021-11-26 14:09:17 -05:00
lioncash b2f44d607f json_ir_generator: Add helper functions for accessing header SSA args
In several parts of the IR we have cases where code accesses header SSA
arguments via:

Op->Header.Args[index]

which can be quite noisy when repeated quite a bit.

This adds member functions to the IR ops that have SSA arguments, that
allow it to simply be:

Op->Args(index)
2021-11-26 13:59:04 -05:00
lioncash 0aa53f726a json_ir_generator: Remove extra tab for enum entries
Makes the generated indentation consistent with the rest of the file.
2021-11-26 13:28:25 -05:00
lioncash b7e2dcc6c1 json_ir_generator: Make some utilities take a pointer to const
A few of these functions are just querying state or returning it, so we
can allow pointers to const to make the functions a little more
flexible.
2021-11-26 13:24:14 -05:00
lioncash 90e9010b17 json_ir_generator: Use _v equivalent of type traits
Same behavior, less writing.
2021-11-26 13:09:34 -05:00
lioncash c62945e615 IR.json: Remove semicolons from MEM_OFFSET constants
Also avoids some -Wextra-semi warnings, given the generation script will
add semicolons for these lines.
2021-11-26 12:57:17 -05:00
lioncash c78976adf1 json_ir_generator: Don't print out semicolons for empty entries
These can cause -Wextra-semi warnings on higher warning levels, so we
can just print out a newline and move on.
2021-11-26 12:55:34 -05:00
Ryan Houdek 13ad5b66bb Merge pull request #1405 from lioncash/interparr
InterpreterOps: Mark F80CMP array as static constexpr
2021-11-26 09:38:59 -08:00
lioncash 25b9bcc652 IR.json: Mark TypeDefinition instances as constexpr
Create() is constexpr, so we can also mark the constants as such to
avoid potentially having static constructors.
2021-11-26 12:23:38 -05:00
lioncash e2e2687a85 IR.json: Remove static from constants
Namespace-scope variables declared as const or constexpr have internal
linkage by default.

Also makes the output a little more consistent in terms of const/static
ordering.
2021-11-26 12:20:53 -05:00
lioncash 2d3aca2398 RegisterAllocationPass: Eliminate trivial casts
We can just specify LiveRange** as the return type of the lambdas to
allow the compiler to assume all returns are of this type.
2021-11-26 12:06:54 -05:00
lioncash a1d1674033 RegisterAllocatorPass: Convert macro functions into concrete ones
Gets rid of some preprocessor leakage across the file and makes the
functions more statically typed.
2021-11-26 12:06:54 -05:00
lioncash a281b6b38b RegisterAllocationPass: Reduce usages of NodeIDs where applicable
Deduplicates some code that makes use of NodeIDs to make migration to a
strongly-typed NodeID a little more straightforward.
2021-11-26 12:06:50 -05:00
lioncash 9e0c652a9a RegisterAllocationPass: Extract remat cost calc to its own function
Keeps it nicely separated from the live-range calculation code.
2021-11-26 10:52:55 -05:00
lioncash fdc0ce7101 InterpreterOps: Mark GetFallbackInfo as static
This is only used within this translation unit.
2021-11-26 10:31:41 -05:00
lioncash 600b1ad88a InterpreterOps: Mark F80CMP array as static constexpr
We don't need to construct this every time this code is executed.
2021-11-26 10:26:37 -05:00
Ryan Houdek b45603d174 Merge pull request #1404 from FEX-Emu/skmp/fix-steamwebhelper
jit/arm64: COND_FGT should map to gt, not hi, as it is not FGTU
2021-11-25 12:41:35 -08:00
Stefanos Kornilios Mitsis Poiitidis 28cb1240b4 jit/arm64: COND_FGT should map to gt, not hi, as it is not FGTU 2021-11-25 22:32:02 +02:00
Ryan Houdek 63f9e0b410 Merge pull request #1403 from lioncash/interp
Interpreter: Build op handler table at compile-time
2021-11-24 22:33:31 -08:00
lioncash e4c1a285ce Interpreter: Build op handler table at compile-time
We can create the table at compile-time, since we already have the
function labels available ahead of time.

This also centralizes the op table in one read-only location
2021-11-25 00:18:04 -05:00
Ryan Houdek 1f306d666c Merge pull request #1401 from lioncash/nodeid
IR: Add type alias for Node IDs
2021-11-24 17:12:58 -08:00
Ryan Houdek baadf0b98e Merge pull request #1402 from lioncash/bucket
BucketList: Minor API touchups
2021-11-24 17:07:44 -08:00
lioncash 0c697af5fb Passes: Make use of NodeID alias where applicable
Unifies the passes so that they're using the NodeID alias as well.
2021-11-24 17:15:09 -05:00
lioncash e6ad608226 IR: Add alias for Node IDs
In quite a few places we have a raw primitive to represent an IR node's
ID. This can make reading some bits of the API (or the passes) a little
confusing to take in, since there's no meaningful type name for some
data structure members.

We can provide an alias that communicates this directly to the reader.

This also has the nice benefit of providing a single point of definition
for node IDs which can allow for easier changes in the future (e.g.
making Node IDs strongly-typed etc).
2021-11-24 17:15:06 -05:00
lioncash 1cf0c335a6 BucketList: Compare against T{} instead of zero directly
Allows BucketList to work with types that aren't a direct numeric
primitive, so long as the object has equality operators defined and are
default constructible
2021-11-24 17:09:20 -05:00
lioncash 52ed4a97cf BucketList: Generify Append() and Erase()
BucketList allows choosing an arbitrary type, but the interface was
assuming uint32_t was only desirable
2021-11-24 17:01:52 -05:00
lioncash db6cbbe7d4 BucketList: Remove dereference to static member
We can reference this directly, since it'll be the same for the lifetime
of the class. Other member functions already do this as well.
2021-11-24 16:57:00 -05:00
lioncash 70a66cdff2 BucketList: Resolve signed/unsigned mismatches in API
The size of the bucket list is expressed as an unsigned value, but all
indexing was taking place with a signed value.
2021-11-24 16:51:12 -05:00
Ryan Houdek ef772c3ef4 Merge pull request #1400 from lioncash/service
CompileService: Store WorkItem instances as unique_ptr
2021-11-24 10:41:05 -08:00
Ryan Houdek 8ab6830fef Merge pull request #1388 from Sonicadvance1/reduce_signal_stalls
Arm64: Reduce the chance of hanging on reentrant allocations
2021-11-24 10:36:07 -08:00
lioncash 5d62a8cedc CompileService: Store WorkItem instances as unique_ptr
Lets us simplify our GC management a little bit and also makes it so we
won't potentially leak memory if an exception occurs anywhere.
2021-11-24 11:48:46 -05:00
Ryan Houdek 262c1c55d2 GDBStub: Fixes a few hangs and crashes
Enough to get a backtrace sometimes but otherwise still pretty finicky.
2021-11-23 17:45:23 -08:00
Ryan Houdek 6913a8b5ec Dispatcher: Initialize Context member states in StoreThreadState
In the case of storing a thread state without a guest signal then this
would be filled with garbage data.
This would then cause a crash on thread restart.

Only happens using gdbstub
2021-11-23 17:42:32 -08:00
Ryan Houdek 5345f7fa37 Merge pull request #1392 from Sonicadvance1/race_condition_compileservice
FEXCore: Fixes CompileService race condition on thread creation
2021-11-23 13:05:34 -08:00
Ryan Houdek 966272ceaf Merge pull request #1398 from lioncash/math
FEXCore: Centralize alignment utility functions in one header
2021-11-23 11:33:48 -08:00
lioncash bb881c5c9b FEXCore: Centralize alignment utility functions
Previously, these alignment functions were in four separate places. We
can centralize these in one predictable spot to remove a little
duplication.
2021-11-23 14:21:48 -05:00
Ryan Houdek 92ae6ae2e4 Merge pull request #1393 from Sonicadvance1/fix_map32bit_error
Linux: Fixes MAP_32BIT error case
2021-11-23 10:42:16 -08:00
Ryan Houdek 7bb789121b Merge pull request #1387 from Sonicadvance1/syscall_fix_sigprocask
Linux: Fixes sigprocmask with new and old set being the same address
2021-11-23 10:39:40 -08:00
Ryan Houdek 97e3a36427 Merge pull request #1386 from Sonicadvance1/ensure_compileservice_signals
JIT: Ensures signals in compileservice JIT space is handled
2021-11-23 10:39:20 -08:00
Ryan Houdek 0560e9be4e Merge pull request #1397 from lioncash/alloc
Syscalls: Move construction of syscall name map into GetSyscallName()
2021-11-23 10:38:46 -08:00
Ryan Houdek 513b02fda3 Merge pull request #1396 from lioncash/fmtstr
General: Migrate logging over to fmt where possible
2021-11-23 10:20:44 -08:00
lioncash 0511764df3 Syscalls: Move construction of syscall name map into GetSyscallName()
These maps are primarily used for debugging, so we don't need to
construct and allocate the data until GetSyscallName() is called.

Removes a bit of memory usage and a static constructor for release
builds.
2021-11-23 13:17:24 -05:00
Lioncash 38b9e85f4f LogManager: Remove now unused portions of the printf-style logger
Now with many of the facilities from the printf logger removed, we can
narrow the exposed functions and in other cases, remove them completely.
2021-11-23 12:51:59 -05:00
Lioncash 75b2f226f6 General: Migrate over to fmt where possible
Migrates lingering instances of the old logger over to fmt where
applicable. This allows removing some of the old defines and functions.

The only remaining usages of the printf-based variant of the logger is
in Tests/LinuxSyscalls/Syscalls.cpp for the strace handling.
2021-11-23 12:51:57 -05:00
Ryan Houdek c4d05fb8e3 Merge pull request #1395 from Sonicadvance1/update_linux_5.16
Updates syscalls to 5.16 syscalls
2021-11-23 09:16:08 -08:00
Ryan Houdek e0d40dd403 Merge pull request #1394 from Sonicadvance1/update_thunksdb
ThunksDB: Updates file to include local libs
2021-11-23 09:16:00 -08:00
Ryan Houdek f97fdd4593 Linux: Fixes MAP_32BIT error case
If MAP_32BIT was failing then we weren't checking error result
correctly.
This was causing us to crash.

Now add a helper function for checking syscall results since it is easy
to mess up.

Replaces the few cases where we were manually checking with the helper.

Fixes a Dota: Underlords crash
2021-11-22 14:15:51 -08:00
Ryan Houdek 6b6b5e880c Linux: Bumps emulated Linux kernel version to 5.16
Only if available as usual.

Fixes #1159
2021-11-22 14:03:28 -08:00
Ryan Houdek 83dc458019 Linux: Adds syscalls for 5.16
Only futex_waitv
2021-11-22 14:03:28 -08:00
Ryan Houdek 4623e4ca21 Linux: Adds syscalls for 5.15
Only process_mrelease
2021-11-22 14:03:28 -08:00
Ryan Houdek bc442871b4 Linux: Adds syscalls for 5.14
memfd_secret and quotactl_fd
2021-11-22 14:03:28 -08:00
Ryan Houdek abaddcccf4 Linux: Adds syscalls for 5.13
Only adds the landlock syscalls
2021-11-22 14:03:28 -08:00
Ryan Houdek ee912c1bb8 Linux: Switches syscall definition usage to FEX's definitions
This works around the problem where some syscall defines may not exist
depending on your compilation environment.
No change in behaviour.
2021-11-22 14:03:28 -08:00
Ryan Houdek f69013fd44 Linux: Updates Arm64 syscall definitions
Fixes the missing definitions from the previous script change
2021-11-22 14:03:28 -08:00
Ryan Houdek 33be4fa98d Scripts: Updates syscall definition extractor
We were missing one of the definition types which means we missed
definitions
2021-11-22 14:03:28 -08:00
Ryan Houdek 99f760dcab ThunksDB: Updates file to include local libs
We were missing installed things in /usr/local, Update the DB to include
those.

Fixes the Ender Lilies vulkan thunk hang that I was encountering
2021-11-22 09:13:15 -08:00
Ryan Houdek 16f1ad4692 Merge pull request #1385 from Sonicadvance1/pressure_vessel_workaround
RootFS: Keep rootfs around while in container
2021-11-22 07:28:07 -08:00
Ryan Houdek e90602f79c Linux: Fixes sigprocmask with new and old set being the same address
If set and oldset point to the same location. We need to be careful to
store the old mask before returning it. Otherwise we will will store the
current mask over the new mask and not change anything.
2021-11-22 06:36:32 -08:00
Ryan Houdek f2d06ecf67 RootFS: Keep rootfs around while in container
For some reason the full container path isn't actually setup right now.
Until this problem is resolved, keep the FEX rootfs configuration
around.

This lets applications still execute albeit still wrapped with the FEX
rootfs rather than the container's
2021-11-22 06:32:56 -08:00
Ryan Houdek 23e583e131 Merge pull request #1390 from Sonicadvance1/packaging_improvements
Various packaging improvements
2021-11-21 18:02:51 -08:00
Ryan Houdek d883b4bf1c Merge pull request #1391 from Sonicadvance1/oops_message
OpcodeDispatcher: Remove debug message
2021-11-21 18:02:37 -08:00
Ryan Houdek f584f16ca4 FEXCore: Fixes CompileService race condition on thread creation
The CompileService was spinning up with the incoming thread mask and
then setting the mask once running.

Instead set the mask, which the thread inherits, then set it back once
it is created.
2021-11-21 11:08:02 -08:00
Ryan Houdek fe3164924c OpcodeDispatcher: Remove debug message 2021-11-21 10:51:41 -08:00
Ryan Houdek 826818192d CPack: Adds more library dependencies
FEXConfig mostly needs these

squashfuse is for FEXMountDaemon
2021-11-21 10:41:52 -08:00
Ryan Houdek 46e30495dc CPack: Add an ldconfig trigger
lintian was complaining about this
2021-11-21 10:38:56 -08:00
Ryan Houdek 702c9e97e4 CPack: Update conflicts and change package name when built statically
Will allow users to choose between a static package or a non-static
package

These packages will conflict with each other, so you can only choose one
or the other.

fex-emu-static: Necessary for Chroots, Can't use thunking.
fex-emu: Necessary for thunking, Can't as easily be used for chroots
2021-11-21 10:36:02 -08:00
Ryan Houdek 2bdd100d2c CPack: Update package contact
lintian was complaining about short description
2021-11-21 10:35:18 -08:00
Ryan Houdek 97ba05ca5b CPack: Update package name
Didn't expose a real version before
2021-11-21 10:34:44 -08:00
Ryan Houdek 508d72d6cc CPack: Add description file
lintian complains about this
2021-11-21 10:32:36 -08:00
Ryan Houdek 3183cf79e5 FEXCore: Adds library soversion
lintian complains about this
2021-11-21 10:31:34 -08:00
Ryan Houdek 5d77a64c8f FEXCore: Compress man page
lintiant complains about this
2021-11-21 10:31:05 -08:00
Ryan Houdek 005c3ce4b3 Tools: Strip binaries on release
lintian complains about this
2021-11-21 10:30:26 -08:00
Ryan Houdek 41adfebf9d JIT: Ensures signals in compileservice JIT space is handled
If a SIGBUS is received in compile service code then we weren't handling
it correctly. Instead we would fail the JIT space check and hand it off
to the guest.
2021-11-20 13:49:54 -08:00
Ryan Houdek e1604fb32f Arm64: Reduce the chance of hanging on reentrant allocations
When compiling code and linking blocks, disable all signals then
reenable once back in the dispatcher.

This adds 15-40 microseconds on to each of these steps and some
additional branching overhead.

This is a problem where our allocator is not reentrant safe and when
receiving a signal we need to jump to a new code region and compile.
Depending on the mutexes that are currently live, we will just stall out
forever.

This is /very/ commonly picked up when attempting to run pressure-vessel
2021-11-20 13:47:30 -08:00
Ryan Houdek 3730a4284c Merge pull request #1382 from neobrain/refactor_thunk_cleanup
Various thunk library cleanups
2021-11-20 09:49:33 -08:00
Ryan Houdek fbd14b65f7 Merge pull request #1383 from Sonicadvance1/fix_error_and_die
FEXCore: Fixes ERROR_AND_DIE
2021-11-19 20:03:00 -08:00
Ryan Houdek 6e1ea92c09 Merge pull request #1384 from Sonicadvance1/support_guest_sigill
FEXCore: Supports guest SIGILL
2021-11-19 20:02:53 -08:00
Ryan Houdek e68d4c52f7 unittests: Updates IR tests for new Break argument
Textual rather than numerical
2021-11-19 12:49:36 -08:00
Ryan Houdek 5759b0d503 FEXCore: Supports guest SIGILL
Fixes #1217

Instead of throwing an error and closing down FEX. Instead pass the
SIGILL to the guest application.

On unhandled instruction implementation the instruction, we instead emit
a _Break IR op at that location.

A _Break IR op will ensure the context state is synchronized at the
point of of the fault and has fairly low overhead. We branch to the
dispatcher which does the SRA spilling.

Tested this with an application that attempts an AVX512 instruction,
catches the fault, and continues onward.
With #1383 in place, we also won't pass spurious ERROR_AND_DIE to the guest anymore.
2021-11-19 12:49:23 -08:00
Ryan Houdek b10ee67525 RCLSE: Don't optimize through a BREAK irop
The end of this optimization pass is getting to a bit long.
It might be worth adding a new flag to IR ops soon for ops we can't
optimize context loadstores through.
2021-11-19 12:19:37 -08:00
Ryan Houdek 47d04ae807 FEXCore: Fixes ERROR_AND_DIE
ERROR_AND_DIE was using __builtin_trap which would send our application
either a SIGILL or SIGTRAP depending on architecture.
This would then be captured by our faulting system and passed over to
the guest application.

If the guest application happened to have a signal handler installed for
these then it would pick up this fault and potentially continue
unsafely.

Now we can remove this usage of __builtin_trap and switch over to our
own handler.

Our frontend will check to see if the fault came from our handler and
uninstall the host signal handlers in this case. Which is what we want
for "ERROR_AND_DIE"
2021-11-19 12:05:30 -08:00
Ryan Houdek 3dc938b8c5 Merge pull request #1377 from Sonicadvance1/inline_32bit_syscalls
Linux: Passthrough 32-bit syscalls that can be
2021-11-18 20:23:59 -08:00
Ryan Houdek d81d5e21ba Merge pull request #1378 from Sonicadvance1/remove_jit_signalframe
Dispatcher: Removes usage of SignalFrame stack on JITs
2021-11-18 20:23:34 -08:00
Ryan Houdek a92233c7f0 Merge pull request #1379 from Sonicadvance1/update_drm
Updates 32-bit DRM emulation
2021-11-18 19:42:00 -08:00
Ryan Houdek 61ccec3726 Merge pull request #1381 from neobrain/fix_slow_stack
Dispatcher: Use std::vector as the underlying container for std::stack
2021-11-18 13:21:36 -08:00
Tony Wasserka 08bfc5e0c0 Thunks: Catch unresolved function references at build time
This could also prove useful for guest libraries, however the implicit
dependency of libEGL on libGL (via glXGetProcAddress) cannot be resolved
easily with the current build setup.
2021-11-18 19:27:09 +01:00
Tony Wasserka f62ec61e4f Thunks: Add missing includes 2021-11-18 19:27:09 +01:00
Tony Wasserka 06881af363 Thunks/xcb: Remove unused function
The callback mechanism is not used for this API since the function pointer
argument is called on the guest itself (if through some indirection).
2021-11-18 19:27:09 +01:00
Tony Wasserka 598533555d Thunks/vulkan: Small cleanup for pointer lookup logic 2021-11-18 19:27:09 +01:00
Tony Wasserka 85651ad090 Thunks/vulkan: Simplify API entrypoint lookup
This avoids the need for the heap-allocated string->symbol map that had
to be kept around previously.
2021-11-18 19:27:09 +01:00
Tony Wasserka 25545dcd66 Thunks/vulkan: Use more efficient unordered_map initialization
Dynamically adding each entry involves unneeded map rehashing that can be
avoided by constructing the container using an initializer_list.
2021-11-18 19:27:09 +01:00
Tony Wasserka a5ac66c1ce Thunks/vulkan: Use std::string_view to look up API entrypoints
This eliminates unneeded memory allocations on library initialization
without change of behavior.
2021-11-18 19:27:09 +01:00
Tony Wasserka b6effa7a9b Thunks/vulkan: Ignore null-instances when looking up symbols and remove obsolete code 2021-11-18 19:27:08 +01:00
Tony Wasserka 730ba42cf3 Dispatcher: Use std::vector as the underlying container for std::stack
By default std::stack uses the heavyweight std::deque container for storage,
but std::vector is perfectly suitable for our purposes and it has better
performance characteristics.
2021-11-18 19:15:11 +01:00
Ryan Houdek 0165329285 Ioctl: Update i915 drm 2021-11-18 03:27:21 -08:00
Ryan Houdek f12b4aa3d1 Ioctl: Update v3d drm 2021-11-18 03:22:48 -08:00
Ryan Houdek a9852d31e7 Ioctl: Update virtio drm 2021-11-18 03:13:59 -08:00
Ryan Houdek 265e8b4d39 Updates External drm-headers repo 2021-11-18 03:13:42 -08:00
Ryan Houdek ef3338ec0d Dispatcher: Removes usage of SignalFrame stack on JITs
Only the interpreter requires this currently since the stack register
doesn't match.
The JITs correctly set up their stack register and matches what the API
wants.

This is guaranteed to work. signal + rt_sigreturn is a handshake with
the kernel to return the data that the kernel wants.
Only time this doesn't align is when the guest signal handler longjumps
out of the signal handler. Which leaves a dangling entry in the
SignalFrame stack on the interpreter.
2021-11-18 03:10:14 -08:00
Ryan Houdek d47182b631 Merge pull request #1376 from lioncash/vex
Frontend: Handle VEX RXB bits
2021-11-18 01:48:24 -08:00
Ryan Houdek 9ac9b89ab3 Linux: Passthrough 32-bit syscalls that can be
Most of these were already marked for passthrough, just needed to be
enabled.

Some syscalls can be passed through but need to be renamed for 32-bit.
This adds another define which does that for us.
2021-11-18 00:37:28 -08:00
Ryan Houdek 50c8d9edab Merge pull request #1375 from lioncash/bzhi
OpcodeDispatcher: Implement BZHI
2021-11-18 00:31:15 -08:00
lioncash 77c969b424 Frontend: Handle VEX RXB bits
These are equivalent to the REX prefix's RXB bits, except that they're
in 1's complement form.

These are trivial to handle and fix usages of the upper range of
registers for ModRM encoded fields.

To ensure that we have coverage for this, I've altered the BEXTR test to
make use of R14 and R15.
2021-11-17 14:50:21 -05:00
lioncash 95aabc0947 OpcodeDispatcher: Implement BZHI 2021-11-17 13:44:57 -05:00
Ryan Houdek a04cc4dc96 Merge pull request #1373 from lioncash/unused
InterpreterCore/Dispatcher: Resolve unused variable warnings
2021-11-17 07:49:11 -08:00
Ryan Houdek 2a894bc111 Merge pull request #1372 from lioncash/copy
Tests: Minor cleanup
2021-11-17 07:49:02 -08:00
Ryan Houdek f746870356 Merge pull request #1371 from Sonicadvance1/inline_syscall
Arm64: Adds an inline syscall optimization
2021-11-17 07:48:51 -08:00
lioncash 70f8779793 Dispatcher: Mark guest_siginfo as [[maybe_unused]] 2021-11-17 08:43:44 -05:00
lioncash 1615ed9ed7 InterpreterCore: Remove unused variable
This became unused once SIGBUS handling was centralized in one location.
2021-11-17 08:38:48 -05:00
lioncash e3a7c14740 HarnessHelpers: Construct fstream directly in ReadFile()
We can also make use of data() here to avoid undefined behavior if any
inputs are 0 for whatever reason (unlikely, but still nice to have the
reassurance).
2021-11-17 07:42:20 -05:00
lioncash 66a5057914 HarnessHelpers: Mark lookup tables as static
We don't need to keep pushing these onto the stack over and over
(especially given how many tests we have).
2021-11-17 07:42:20 -05:00
lioncash 62dfccc989 HarnessHelpers: Migrate to fmt
May as well, given we're in the area
2021-11-17 07:42:17 -05:00
lioncash 51d8bb9020 ELFCodeLoader2: Migrate to fmt
May as well, given we're in the area.
2021-11-17 07:06:27 -05:00
lioncash 65e28ba6cf IRLoader: Migrate logging to fmt
Given we're in proximity of FEXLoader, we may as well.
2021-11-17 06:58:57 -05:00
lioncash c181033263 FEXLoader: Take string as const in RanAsInterpreter()
This doesn't modify the input string.
2021-11-17 06:55:43 -05:00
lioncash f371b04d7c FEXLoader: Construct fstreams directly
We can use the constructor to directly open the fstream
2021-11-17 06:55:43 -05:00
lioncash 98f42aafa9 FEXLoader: Migrate logging to fmt
While we're in the area, lets move logging over to fmt to get it out of
the way.

This also gives us an opportunity to merge some CMakeLists things
together to organize it.
2021-11-17 06:55:40 -05:00
Lioncash 69437eaf11 FEXLoader: Take sections by const reference in loop
Given LoadedSection instances are 72 bytes, this avoids a minor bit of
copy churn
2021-11-17 06:01:35 -05:00
Ryan Houdek c63c5fb664 Merge pull request #1369 from lioncash/cast
OpcodeDispatcher: Amend return types of BitSize helpers
2021-11-17 02:00:50 -08:00
Ryan Houdek ec8e24e84a IR: Adds documentation description for syscall and inlinesyscall 2021-11-16 22:53:44 -08:00
Ryan Houdek c6b26739db Interpreter: Implements the InlineSyscall for the interpreter
This isn't optimal, just here to work
2021-11-16 22:53:44 -08:00
Ryan Houdek f9e22432d2 Arm64: Adds an inline syscall optimization
In a syscall microbench this improves performance by ~19% on my
Snapdragon 888.
Going from ~9.6 million syscalls per second to ~11.5 million.

Macbook Pro is less effective here due to high syscall overhead due to
VM. Going form  7.2M/s to 7.6M/s, ~6% improvement

We can also inline some 32-bit syscalls but that will need some more
work which isn't done yet. Even though the op in the JIT supports it.
2021-11-16 22:53:44 -08:00
Ryan Houdek d54b9cd272 IREmitter: ReplaceAllWith can't remove sideeffect nodes
Ran in to this when removing syscall nodes. If an IR op has side-effects
then this generic helper can not remove them.
First time this was encountered and it was confusing
2021-11-15 14:37:25 -08:00
Ryan Houdek 92b29138ba Linux: Describes syscalls that can be passed through without change
~0 is used as an invalid syscall indicator to signify that it can be
passed through
2021-11-15 14:36:33 -08:00
Ryan Houdek bab4f53b59 Linux: Updates syscalls enums
Adds the Arm64 one as well so we can map syscall names directly with
renaming here

Adds some comments about which syscalls have no implementation.
Automatically generated.
2021-11-15 14:30:57 -08:00
Ryan Houdek aae1dd4d81 Adds new script to generate syscall enums
Lets us define our own enums for syscall mapping.
Necessary to get all the syscall IDs in a sane way
2021-11-15 14:29:49 -08:00
Ryan Houdek 977d92b7bc Merge pull request #1370 from lioncash/enum-shift
EnumUtils: Remove shift operators from helper macro
2021-11-15 14:16:01 -08:00
lioncash e1cf60c2ab EnumUtils: Remove shift operators from helper macro
These aren't strictly necessary for enum flags and through discussion in
\#1363, would lead to an awkward to use overload.

If these are ever needed, they can be added back at a later date.
2021-11-15 12:03:43 -05:00
Ryan Houdek cabb58ad29 Merge pull request #1368 from Sonicadvance1/32bit_host_runner
unittests: Enables 32-bit host runner
2021-11-15 09:00:17 -08:00
lioncash 44f27c81d3 OpcodeDispatcher: Mark size retrieval functions as nodiscard
Allows the compiler to warn about obvious bug cases.
2021-11-15 11:56:21 -05:00
lioncash c9c8c38d86 OpcodeDispatcher: Increase BitSize helper return vals to uint32_t
Pointed out by neobrain in #1362 that there are some cases where this
will overflow and not fit into 8 bits.
2021-11-15 11:51:09 -05:00
lioncash 5c8c36c5d2 OpcodeDispatcher: Amend cast in GPROffset
There's no bug here, but it definitely looks weird to not be aligned
with the return type.
2021-11-15 11:48:18 -05:00
Ryan Houdek d413cf8d62 unittests: Enables 32-bit host runner
A few tests messing with segments can't be run on the host. This is
because they don't exactly match expected Linux LDT/GDT setup.
2021-11-13 01:37:09 -08:00
Ryan Houdek 18623d07aa TestHarnessRunner: Adds support for 32-bit host runner
Little bit of work to set up the local descriptor table entries for
32-bit execution.
Also setting up the far call state.
2021-11-13 01:37:09 -08:00
Ryan Houdek e63c832a5a unittests: Fixes 32-bit unit test missing flag
This was being run as 64-bit and happening to work
2021-11-13 01:36:09 -08:00
Ryan Houdek 42161b20a4 Updates xbyak external 2021-11-13 01:36:09 -08:00
Ryan Houdek 081d0169d8 Merge pull request #1360 from Sonicadvance1/map_32bit_emulation
Linux: Emulate MAP_32BIT on mmap
2021-11-12 21:49:35 -08:00
Ryan Houdek 77093a7f4d Merge pull request #1358 from Sonicadvance1/signal_rip_adjust
Dispatcher: Partial support for RIP adjust in signal handler
2021-11-12 21:49:25 -08:00
Ryan Houdek bb7f542813 Linux: Emulate MAP_32BIT on mmap
Previous we were just using an address hint to emulate MAP_32BIT.
Seemingly this behaviour has changed on AArch64 where it now isn't
guaranteed to scan up from the hint provided if exact allocation fails.

Now we pull in the full 32-bit allocator and add support for MAP_32BIT
in it. This limits the allocations there in to the first 2GB which Linux
expects.

Necessary for Mono's trampolines to work since it requires code to be in
the first 2GB on x86-64.
2021-11-12 21:25:15 -08:00
Ryan Houdek 2f28548ce4 Dispatcher: Partial support for RIP adjust in signal handler
When the context structure is adjusted in the signal handler then this
updates the state of the CPU on sigreturn.

We check to see if the RIP was adjusted and in this case we will update
the JIT in a potentially unsafe fashion.
This works well enough with signal handlers that effectively just long
jump and not much else.

Fixes wine initial prefix setup where rundl32 expects to do a try-catch
fault capture but infinite loops without this.
2021-11-12 21:22:07 -08:00
Ryan Houdek 498a845c0b Merge pull request #1365 from lioncash/mulx
OpcodeDispatcher: Implement MULX
2021-11-12 21:17:25 -08:00
Ryan Houdek 59ed91c50a Merge pull request #1366 from lioncash/enumreg
OpcodeDecoder: Shorten up GPR and MM base offset retrieval
2021-11-12 20:56:01 -08:00
Ryan Houdek ad864e0e52 Merge pull request #1367 from lioncash/flag
OpcodeDispatcher: Resolve sign mismatches in loops
2021-11-12 17:35:16 -08:00
lioncash 9849aef5bc Frontend: Handle ignoring of widening modes if not 64-bit mode 2021-11-12 17:00:43 -05:00
lioncash 11c3ae19b4 OpcodeDispatcher: Add helper for retrieving MM base offset
Shortens up some lines to make them easier to read.
2021-11-12 16:52:20 -05:00
lioncash 4351ec9f4e OpcodeDispatcher: Add helper for retrieving GPR offsets
Greatly shortens the length of multiple lines. This makes it nicer to
see which register is being loaded or stored.
2021-11-12 16:52:13 -05:00
lioncash 59469705e4 OpcodeDispatcher: Implement MULX 2021-11-12 16:46:38 -05:00
lioncash 763e7c5e08 OpcodeDispatcher: Resolve sign mismatches in loops
Silences compiler warnings
2021-11-12 13:59:46 -05:00
lioncash 2edb5d3004 X86Enums: Convert constants into enums
Allows for parameters and functions to make use of the enum type for
enforcing type checking.
2021-11-12 13:34:55 -05:00
Ryan Houdek c6662a46b5 Merge pull request #1364 from lioncash/fptr
Frontend: Mark modrm function LUT as static
2021-11-12 06:02:08 -08:00
Ryan Houdek d1fa8bcfba Merge pull request #1363 from lioncash/enum
OpcodeDecoder: Convert enums into enum classes where applicable
2021-11-12 05:56:23 -08:00
Ryan Houdek 3a3cbf2b5f Merge pull request #1362 from lioncash/bit
OpcodeDispatcher: Add helper for getting bit sizes
2021-11-12 05:53:05 -08:00
lioncash a0eb2eb2e2 Frontend: Mark modrm function LUT as static
This doesn't need to be constructed on a by-instance basis.
2021-11-12 08:48:49 -05:00
lioncash d58952d1ef OpcodeDispatcher: Make selection flag enum an enum class
Makes the enum and member variable strongly typed, so that it's harder
to implicitly use incorrect values.
2021-11-12 08:41:07 -05:00
lioncash 05ce65322f OpcodeDispatcher: Make X87Tag an enum class
Makes the tag parameter strongly-typed and prevents implicitly passing
it an incorrect value.
2021-11-12 08:31:50 -05:00
lioncash 3e5037042f OpcodeDispatcher: Make Segment enum an enum class
Avoids polluting the surrounding scope.
2021-11-12 08:24:21 -05:00
lioncash 6d82410f44 OpcodeDispatcher: Convert OpType enum into enum class
Prevents pollution of the surrounding scope a little.
2021-11-12 08:19:14 -05:00
lioncash 49966f2954 Utils: Add header for enum-based utilities
Adds a header with a few utilities that make working with strongly typed
enums a little more convenient, especially when working with enum
classes.
2021-11-12 08:15:08 -05:00
Ryan Houdek 51d21d861d Merge pull request #1361 from lioncash/rorx
OpcodeDispatcher: Handle RORX
2021-11-12 05:06:01 -08:00
lioncash 6a79b5cf56 OpcodeDispatcher: Add helper for getting bit sizes
Allows expressing what we're directly getting instead of needing to
repeat it in several places, given how common this is.
2021-11-12 08:01:43 -05:00
lioncash 1d88b499e0 OpcodeDispatcher: Handle RORX 2021-11-12 07:28:28 -05:00
Ryan Houdek b327daf5f5 Merge pull request #1359 from Sonicadvance1/fix_brk_allocate
ELFCodeLoader: Fixes BRK allocation on some ELF loading
2021-11-12 02:40:11 -08:00
Ryan Houdek 90511ddd19 Merge pull request #1357 from Sonicadvance1/fix_sigsuspend_signalmask
Linux: SignalDelegator update signal mask on return from sigsuspend
2021-11-12 02:39:57 -08:00
Ryan Houdek 4be8e23299 ELFCodeLoader: Fixes BRK allocation on some ELF loading
With a DYN ELF in most cases the first program header will be at offset
0.

This isn't always true though, in the case of `micro` the first program
header is at a 4MB offset.
This was causing us to miscalculate the size of ELF in memory, which was
causing BRK to intersect with the ELF on allocation.
This would cause ELF loading to fail and then us to close in some cases.

Also makes sure to pass -1 as FD in mmap with anonymous to match.
2021-11-11 22:13:59 -08:00
Ryan Houdek d4a7afbf5c Linux: SignalDelegator update signal mask on return from sigsuspend
We were failing to update the emulated signal mask on return from
sigsuspend.
This was breaking Unity titles which was using sigsuspend and then
modifying the signal mask with sigprocaddr
2021-11-11 22:06:51 -08:00
Ryan Houdek d3cd354c48 Merge pull request #1356 from lioncash/adx
OpcodeDecoder: Add support for ADX
2021-11-10 10:06:19 -08:00
lioncash 5d707cf831 OpcodeDecoder: Add support for ADX
Another CPU extension we can cross off the list.

Addresses #1355
2021-11-10 12:11:01 -05:00
Ryan Houdek e37a0bbd34 Merge pull request #1354 from lioncash/bmishift
OpcodeDispatcher: Implement BMI2 SARX/SHLX/SHRX
2021-11-07 00:42:41 -07:00
lioncash 9df8245ef2 OpcodeDispatcher: Implement BMI2 SARX/SHLX/SHRX
Gets the basic BMI2 shifts out of the way.
2021-11-07 02:12:06 -05:00
538 changed files with 41151 additions and 28536 deletions

No files matched your search

@@ -0,0 +1,45 @@
---
name: Potential Game Bug
about: A bug in FEX-Emu that causes a problem in a game
title: "[Game]: [Short Problem Description]"
labels: Game related
assignees: ''
---
**What Game**
The game name.
A link to the storefront where to get the game. GOG, Steam, Itch.io, etc
**Describe the bug**
A clear and concise description of what the bug is.
**To Reproduce**
Steps to reproduce the behavior:
1. Go to '...'
2. Click on '....'
3. Scroll down to '....'
4. See error
**Expected behavior**
A clear and concise description of what you expected to happen.
**Screenshots and Video**
If applicable, add screenshots and video to help explain your problem.
**System information:**
- OS: [eg: Ubuntu 21.10]
- CPU/SoC: [eg: Snapdragon 888, Intel Core i8-12900k]
- Video driver version: [eg: OpenGL ES 3.2 Mesa 22.0.0-devel (git-9ff086052a)]
- RootFS used: [eg: Ubuntu 21.10 Official Rootfs]
- FEX version: (FEXGetConfig --version) [eg: FEX-2112-155-gc691d709]
- Thunks Enabled: [Yes/No]
**Additional context**
- Is this an x86 or x86-64 game: [x86/x86-64/Both]
- Does this reproduce on x86-64 host with FEX: [Yes/No/Untested]
- Does this reproduce on AArch64 with Radeon/Intel/Nvidia: [Yes/No/Untested]
- Is this a Vulkan game: [Yes/No/Unknown]
- If Yes, What is your Vulkan driver:
Add any other context about the problem here.
+15 -2
View File
@@ -31,7 +31,9 @@ jobs:
- name : submodule checkout
# Need to update submodules
run: git submodule update --init --depth 1
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
@@ -49,7 +51,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -140,6 +142,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_StructVerifier.log || true
- name: APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target api_tests
- name: APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+8 -5
View File
@@ -1,7 +1,7 @@
[submodule "External/vixl"]
shallow = true
path = External/vixl
url = https://github.com/Sonicadvance1/vixl.git
url = https://github.com/FEX-Emu/vixl.git
[submodule "External/cpp-optparse"]
path = External/cpp-optparse
url = https://github.com/Sonicadvance1/cpp-optparse
@@ -42,7 +42,10 @@
[submodule "External/xxhash"]
path = External/xxhash
url = https://github.com/FEX-Emu/xxHash.git
[submodule "External/Vulkan-Docs"]
shallow = true
path = External/Vulkan-Docs
url = https://github.com/KhronosGroup/Vulkan-Docs.git
[submodule "External/Catch2"]
path = External/Catch2
url = https://github.com/catchorg/Catch2.git
[submodule "External/robin-map"]
shallow = true
path = External/robin-map
url = https://github.com/Tessil/robin-map.git
+157 -90
View File
@@ -17,11 +17,21 @@ option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_STATIC_PIE "Enables static-pie build" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
set (X86_C_COMPILER "x86_64-linux-gnu-gcc" CACHE STRING "c compiler for compiling x86 guest libs")
set (X86_CXX_COMPILER "x86_64-linux-gnu-g++" CACHE STRING "c++ compiler for compiling x86 guest libs")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
set (OVERRIDE_VERSION "detect" CACHE STRING "Override the FEX version in the format of <MMYY>{.<REV>}")
string(TOUPPER "${CMAKE_BUILD_TYPE}" CMAKE_BUILD_TYPE)
if (CMAKE_BUILD_TYPE MATCHES "DEBUG")
set(ENABLE_ASSERTIONS TRUE)
@@ -32,6 +42,11 @@ if (ENABLE_ASSERTIONS)
add_definitions(-DASSERTIONS_ENABLED=1)
endif()
if (ENABLE_INTERPRETER)
message(STATUS "Interpreter enabled")
add_definitions(-DINTERPRETER_ENABLED=1)
endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
set(CMAKE_RUNTIME_OUTPUT_DIRECTORY ${CMAKE_BINARY_DIR}/Bin)
@@ -66,10 +81,12 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
add_definitions(-D_M_ARM_64=1)
endif()
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
message(STATUS "CCache enabled")
set_property(GLOBAL PROPERTY RULE_LAUNCH_COMPILE "${CCACHE_PROGRAM}")
endif()
endif()
if (ENABLE_XRAY)
@@ -77,17 +94,35 @@ if (ENABLE_XRAY)
link_libraries(-fxray-instrument)
endif()
if (ENABLE_COMPILE_TIME_TRACE)
add_compile_options(-ftime-trace)
link_libraries(-ftime-trace)
endif()
set (PTHREAD_LIB pthread)
if (ENABLE_LLD)
set (LD_OVERRIDE "-fuse-ld=lld")
link_libraries(${LD_OVERRIDE})
endif()
if (ENABLE_LIBCXX)
message(WARNING "This is an unsupported configuration and should only be used for testing")
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -std=c++11 -stdlib=libc++")
set(CMAKE_EXE_LINKER_FLAGS "${CMAKE_EXE_LINKER_FLAGS} -stdlib=libc++ -lc++abi")
endif()
if (NOT ENABLE_OFFLINE_TELEMETRY)
# Disable FEX offline telemetry entirely if asked
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
add_definitions(-DTERMUX_BUILD=1)
set(TERMUX_BUILD 1)
# Termux doesn't support Jemalloc due to bad interactions between emutls, jemalloc, and scudo
set(ENABLE_JEMALLOC FALSE)
endif()
if (ENABLE_STATIC_PIE)
if (_M_ARM_64 AND ENABLE_LLD)
message (FATAL_ERROR "Static linking does not currently work with AArch64+LLD. Use GNU ld for now.")
@@ -241,6 +276,8 @@ set (CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fn
set (CMAKE_CXX_FLAGS_RELEASE "${CMAKE_CXX_FLAGS_RELEASE} -fomit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-pointer")
include_directories(External/robin-map/include/)
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
@@ -251,17 +288,23 @@ endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
pkg_check_modules(XXHASH libxxhash>=0.8.0 QUIET)
if (NOT XXHASH_FOUND)
message(STATUS "xxHash not found. Using Externals")
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
endif()
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
option(CATCH_BUILD_STATIC_LIBRARY "" ON)
set(CATCH_BUILD_STATIC_LIBRARY ON)
add_subdirectory(External/Catch2/)
# Pull in catch_discover_tests definition
list(APPEND CMAKE_MODULE_PATH "${CMAKE_CURRENT_SOURCE_DIR}/External/Catch2/contrib/")
include(Catch)
endif()
add_subdirectory(External/cpp-optparse/)
include_directories(External/cpp-optparse/)
@@ -308,33 +351,51 @@ if(ENABLE_WERROR OR ENABLE_STRICT_WERROR)
endif()
endif()
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=native")
if (NOT TUNE_ARCH STREQUAL "generic")
check_cxx_compiler_flag("-march=${TUNE_ARCH}" COMPILER_SUPPORTS_ARCH_TYPE)
if(COMPILER_SUPPORTS_ARCH_TYPE)
add_compile_options("-march=${TUNE_ARCH}")
else()
message(FATAL_ERROR "Trying to compile arch type '${TUNE_ARCH}' but the compiler doesn't support this")
endif()
endif()
if (TUNE_CPU STREQUAL "native")
if(_M_ARM_64)
if (CMAKE_CXX_COMPILER_VERSION VERSION_GREATER_EQUAL 999999.0)
# Clang 12.0 fixed the -mcpu=native bug with mixed big.little implementers
# Clang can not currently check for native Apple M1 type in hypervisor. Currently disabled
check_cxx_compiler_flag("-mcpu=native" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=native")
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${AARCH64_CPU}")
endif()
endif()
else()
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo" "${CMAKE_CXX_COMPILER_VERSION}"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
add_compile_options("-march=native")
endif()
endif()
else()
check_cxx_compiler_flag("-march=native" COMPILER_SUPPORTS_MARCH_NATIVE)
if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
add_compile_options("-mcpu=${TUNE_CPU}")
else()
message(FATAL_ERROR "Trying to compile cpu type '${TUNE_CPU}' but the compiler doesn't support this")
endif()
endif()
@@ -408,6 +469,10 @@ if (BUILD_TESTS)
enable_testing()
message(STATUS "Unit tests are enabled")
endif()
add_subdirectory(FEXHeaderUtils/)
include_directories(FEXHeaderUtils/)
add_subdirectory(External/FEXCore)
# Binfmt_misc files must be installed prior to Source/ installs
@@ -426,29 +491,16 @@ if (BUILD_TESTS)
endif()
if (BUILD_THUNKS)
add_subdirectory(ThunkLibs/Generator)
# Thunk targets for both host libraries and IDE integration
add_subdirectory(ThunkLibs/HostLibs)
# Thunk targets for IDE integration of guest code, only
add_subdirectory(ThunkLibs/GuestLibs)
# Thunk targets for guest libraries
include(ExternalProject)
ExternalProject_Add(host-libs
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
CMAKE_ARGS
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
install(
CODE "MESSAGE(\"-- Installing: host-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkHostsInstall
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Host
)"
DEPENDS host-libs
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
@@ -459,15 +511,16 @@ if (BUILD_THUNKS)
"-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DVULKAN_XML=${CMAKE_SOURCE_DIR}/External/Vulkan-Docs/xml/vk.xml"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
install(
CODE "MESSAGE(\"-- Installing: guest-libs\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target ThunkGuestsInstall
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
DEPENDS guest-libs
@@ -478,59 +531,73 @@ set(FEX_VERSION_MAJOR "0")
set(FEX_VERSION_MINOR "0")
set(FEX_VERSION_PATCH "0")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=0
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
RESULT_VARIABLE GIT_ERROR
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (OVERRIDE_VERSION STREQUAL "detect")
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe --abbrev=0
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
RESULT_VARIABLE GIT_ERROR
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (NOT ${GIT_ERROR} EQUAL 0)
# Likely built in a way that doesn't have tags
# Setup a version tag that is unknown
set(GIT_DESCRIBE_STRING "FEX-0000")
if (NOT ${GIT_ERROR} EQUAL 0)
# Likely built in a way that doesn't have tags
# Setup a version tag that is unknown
set(GIT_DESCRIBE_STRING "FEX-0000")
endif()
endif()
else()
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
# Change something like `FEX-2106.1-76-<hash>` in to a list
string(REPLACE "-" ";" DESCRIBE_LIST ${GIT_DESCRIBE_STRING})
# Parse the version here
# Change something like `FEX-2106.1-76-<hash>` in to a list
string(REPLACE "-" ";" DESCRIBE_LIST ${GIT_DESCRIBE_STRING})
# Extract the `2106.1` element
list(GET DESCRIBE_LIST 1 DESCRIBE_LIST)
# Extract the `2106.1` element
list(GET DESCRIBE_LIST 1 DESCRIBE_LIST)
# Change `2106.1` in to a list
string(REPLACE "." ";" DESCRIBE_LIST ${DESCRIBE_LIST})
# Change `2106.1` in to a list
string(REPLACE "." ";" DESCRIBE_LIST ${DESCRIBE_LIST})
# Calculate list size
list(LENGTH DESCRIBE_LIST LIST_SIZE)
# Calculate list size
list(LENGTH DESCRIBE_LIST LIST_SIZE)
# Pull out the major version
list(GET DESCRIBE_LIST 0 FEX_VERSION_MAJOR)
# Pull out the major version
list(GET DESCRIBE_LIST 0 FEX_VERSION_MAJOR)
# Minor version only exists if there is a .1 at the end
# eg: 2106 versus 2106.1
if (LIST_SIZE GREATER 1)
list(GET DESCRIBE_LIST 1 FEX_VERSION_MINOR)
endif()
# Minor version only exists if there is a .1 at the end
# eg: 2106 versus 2106.1
if (LIST_SIZE GREATER 1)
list(GET DESCRIBE_LIST 1 FEX_VERSION_MINOR)
endif()
# Package creation
set (CPACK_GENERATOR "DEB")
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_PACKAGE_CONTACT "team@fex-emu.org")
if (ENABLE_STATIC_PIE)
set (CPACK_PACKAGE_NAME fex-emu-static)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu")
else()
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu-static")
endif()
set (CPACK_PACKAGE_FILE_NAME "${CPACK_PACKAGE_NAME}-${GIT_DESCRIBE_STRING}_${CMAKE_SYSTEM_PROCESSOR}")
set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.org>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libstdc++6")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA "${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm")
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "qemu-user-static")
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "${CPACK_DEBIAN_PACKAGE_CONFLICTS}, qemu-user-static")
endif()
include (CPack)
+3
View File
@@ -0,0 +1,3 @@
x86 and x86-64 Linux emulator
FEX is very much work in progress, so expect things to change.
+1
View File
@@ -0,0 +1 @@
activate-noawait ldconfig
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"StallProcess": "1"
}
}
+62
View File
@@ -10,6 +10,10 @@
"/usr/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.2.0",
"/usr/local/lib/x86_64-linux-gnu/libGL.so.1.7.0",
"/lib/x86_64-linux-gnu/libGL.so",
"/lib/x86_64-linux-gnu/libGL.so.1",
"/lib/x86_64-linux-gnu/libGL.so.1.2.0",
@@ -25,6 +29,9 @@
"/usr/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/usr/local/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0",
"/lib/x86_64-linux-gnu/libGLESv2.so",
"/lib/x86_64-linux-gnu/libGLESv2.so.2",
"/lib/x86_64-linux-gnu/libGLESv2.so.2.0.0"
@@ -36,6 +43,9 @@
"/usr/lib/x86_64-linux-gnu/libX11.so",
"/usr/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libX11.so",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6",
"/usr/local/lib/x86_64-linux-gnu/libX11.so.6.4.0",
"/lib/x86_64-linux-gnu/libX11.so",
"/lib/x86_64-linux-gnu/libX11.so.6",
"/lib/x86_64-linux-gnu/libX11.so.6.4.0"
@@ -48,6 +58,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_radeon.so",
"/lib/x86_64-linux-gnu/libvulkan_radeon.so"
],
"Comment": [
@@ -61,6 +72,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_lvp.so",
"/lib/x86_64-linux-gnu/libvulkan_lvp.so"
]
},
@@ -71,6 +83,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_freedreno.so",
"/lib/x86_64-linux-gnu/libvulkan_freedreno.so"
]
},
@@ -81,6 +94,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_intel.so",
"/lib/x86_64-linux-gnu/libvulkan_intel.so"
]
},
@@ -91,6 +105,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_panfrost.so",
"/lib/x86_64-linux-gnu/libvulkan_panfrost.so"
]
},
@@ -101,6 +116,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/usr/local/lib/x86_64-linux-gnu/libGLX_nvidia.so.0",
"/lib/x86_64-linux-gnu/libGLX_nvidia.so.0"
],
"Comment": [
@@ -114,6 +130,7 @@
],
"Overlay": [
"/usr/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/usr/local/lib/x86_64-linux-gnu/libvulkan_virtio.so",
"/lib/x86_64-linux-gnu/libvulkan_virtio.so"
]
},
@@ -123,6 +140,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb.so",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb.so.1.1.0",
"/lib/x86_64-linux-gnu/libxcb.so",
"/lib/x86_64-linux-gnu/libxcb.so.1",
"/lib/x86_64-linux-gnu/libxcb.so.1.1.0"
@@ -134,6 +154,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri2.so.0.0.0"
@@ -145,6 +168,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0",
"/lib/x86_64-linux-gnu/libxcb-dri3.so.0.0.0"
@@ -156,6 +182,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0",
"/lib/x86_64-linux-gnu/libxcb-xfixes.so.0.0.0"
@@ -167,6 +196,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0",
"/lib/x86_64-linux-gnu/libxcb-shm.so.0.0.0"
@@ -178,6 +210,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0",
"/lib/x86_64-linux-gnu/libxcb-sync.so",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1",
"/lib/x86_64-linux-gnu/libxcb-sync.so.1.0.0"
@@ -189,6 +224,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0",
"/lib/x86_64-linux-gnu/libxcb-randr.so.0.1.0"
@@ -200,6 +238,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-present.so",
"/lib/x86_64-linux-gnu/libxcb-present.so.0",
"/lib/x86_64-linux-gnu/libxcb-present.so.0.0.0"
@@ -211,6 +252,9 @@
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/usr/local/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0",
"/lib/x86_64-linux-gnu/libxcb-glx.so.0.0.0"
@@ -222,6 +266,9 @@
"/usr/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/usr/local/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0",
"/lib/x86_64-linux-gnu/libxshmfence.so",
"/lib/x86_64-linux-gnu/libxshmfence.so.1",
"/lib/x86_64-linux-gnu/libxshmfence.so.1.0.0"
@@ -233,6 +280,9 @@
"/usr/lib/x86_64-linux-gnu/libdrm.so",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2",
"/usr/local/lib/x86_64-linux-gnu/libdrm.so.2.4.0",
"/lib/x86_64-linux-gnu/libdrm.so",
"/lib/x86_64-linux-gnu/libdrm.so.2",
"/lib/x86_64-linux-gnu/libdrm.so.2.4.0"
@@ -244,6 +294,9 @@
"/usr/lib/x86_64-linux-gnu/libasound.so",
"/usr/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/usr/local/lib/x86_64-linux-gnu/libasound.so",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2",
"/usr/local/lib/x86_64-linux-gnu/libasound.so.2.0.0",
"/lib/x86_64-linux-gnu/libasound.so",
"/lib/x86_64-linux-gnu/libasound.so.2",
"/lib/x86_64-linux-gnu/libasound.so.2.0.0"
@@ -255,6 +308,9 @@
"/usr/lib/x86_64-linux-gnu/libXrender.so",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1",
"/usr/local/lib/x86_64-linux-gnu/libXrender.so.1.3.0",
"/lib/x86_64-linux-gnu/libXrender.so",
"/lib/x86_64-linux-gnu/libXrender.so.1",
"/lib/x86_64-linux-gnu/libXrender.so.1.3.0"
@@ -266,6 +322,9 @@
"/usr/lib/x86_64-linux-gnu/libXext.so",
"/usr/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/usr/local/lib/x86_64-linux-gnu/libXext.so",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6",
"/usr/local/lib/x86_64-linux-gnu/libXext.so.6.4.0",
"/lib/x86_64-linux-gnu/libXext.so",
"/lib/x86_64-linux-gnu/libXext.so.6",
"/lib/x86_64-linux-gnu/libXext.so.6.4.0"
@@ -277,6 +336,9 @@
"/usr/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3",
"/usr/local/lib/x86_64-linux-gnu/libXfixes.so.3.1.0",
"/lib/x86_64-linux-gnu/libXfixes.so",
"/lib/x86_64-linux-gnu/libXfixes.so.3",
"/lib/x86_64-linux-gnu/libXfixes.so.3.1.0"
Vendored Submodule
+1
Submodule External/Catch2 added at c4e3767e26.
+23 -18
View File
@@ -37,27 +37,32 @@ endif()
set(CMAKE_CXX_STANDARD 20)
set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
# Find our git hash
find_package(Git)
set(GIT_SHORT_HASH "Unknown")
set(GIT_DESCRIBE_STRING "FEX-Unknown")
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
if (OVERRIDE_VERSION STREQUAL "detect")
# Find our git hash
find_package(Git)
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
endif()
else()
set(GIT_SHORT_HASH "${OVERRIDE_VERSION}")
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
configure_file(
+2 -12
View File
@@ -18,14 +18,8 @@ This project aims to provide a fast and functional x86-64 emulation library that
* Portable library implementation in order to support easy integration in to applications
### Target Host Architecture
The target host architecture for this library is AArch64. Specifically the ARMv8.1 version or newer.
The CPU IR is designed with AArch64 in mind but there is a desire to run the recompiled code on other architectures as well.
Multiple architecture support is desired for easier bringup and debugging, performance isn't as much of a priority there (ex. x86-64(guest) translated to x86-64(host))
### Not currently goals but will be in the future
* 32bit x86 support
* This will be a desire in the future, but to lower the amount of work required, decided to push this off for now.
* Integration in to WINE
* Later generation of x86-64 instruction sets
* Including AVX, F16C, XOP, FMA, AVX2, etc
The CPU IR is designed with AArch64 in mind but should allow for other architectures as well.
x86-64 host support is available for ease of development, but is not a priority.
### Not desired
* Kernel space emulation
* CPL0-2 emulation
@@ -33,7 +27,3 @@ Multiple architecture support is desired for easier bringup and debugging, perfo
* IRQs
* SVM
* "Cycle Accurate" emulation
### Dependencies
* clang-tidy if you want to ensure the code stays tidy
* cmake
* A C++17 compliant compiler (There are assumptions made about using Clang and LTO)
+6 -43
View File
@@ -7,17 +7,12 @@ OpClasses = collections.OrderedDict()
def get_ir_classes(ops, defines):
global OpClasses
for op_key, op_vals in ops.items():
if not ("Last" in op_vals):
OpClass = "#Unknown"
for op_class, opslist in ops.items():
if not (op_class in OpClasses):
OpClasses[op_class] = []
if ("OpClass" in op_vals):
OpClass = op_vals["OpClass"]
if not (OpClass in OpClasses):
OpClasses[OpClass] = []
OpClasses[OpClass].append([op_key, op_vals])
for op, op_val in opslist.items():
OpClasses[op_class].append([op, op_val])
# Sort the dictionary after we are done parsing it
OpClasses = collections.OrderedDict(sorted(OpClasses.items()))
@@ -38,41 +33,9 @@ def print_ir_ops():
op_key = op[0]
op_vals = op[1]
output_file.write("## %s\n" % (op_key))
HasDest = ("HasDest" in op_vals and op_vals["HasDest"] == True)
HasSSAArgs = ("SSAArgs" in op_vals and len(op_vals["SSAArgs"]) > 0)
HasSSAArgNames = "SSANames" in op_vals
HasArgs = "Args" in op_vals
SSAArgsCount = 0
ArgCount = 0
if (HasSSAArgs):
SSAArgsCount = int(op_vals["SSAArgs"])
if (HasArgs):
ArgCount = len(op_vals["Args"])
TotalArgsCount = SSAArgsCount + (ArgCount / 2)
output_file.write(">")
if (HasDest):
output_file.write("%dest = ")
output_file.write("%s " % op_key)
ArgComma = (", ", "")
if (HasSSAArgs):
for i in range(0, SSAArgsCount):
FinalArg = (i + 1) == TotalArgsCount
if (HasSSAArgNames):
output_file.write("%%%s%s" % (op_vals["SSANames"][i], ArgComma[FinalArg]))
else:
output_file.write("%%ssa%d%s" % (i, ArgComma[FinalArg]))
if (HasArgs):
Args = op_vals["Args"]
for i in range(0, ArgCount, 2):
FinalArg = ((i / 2) + SSAArgsCount + 1) == TotalArgsCount
data_type = Args[i]
data_name = Args[i + 1]
output_file.write("\<%s %s\>%s" % (data_type, data_name, ArgComma[FinalArg]))
output_file.write(op_key)
output_file.write("\n\n")
Vendored Regular → Executable
+493 -392
View File
File diff suppressed because it is too large. Load diff
+67 -43
View File
@@ -1,9 +1,15 @@
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (SRCS
set (FEXCORE_BASE_SRCS
Common/Paths.cpp
Interface/Config/Config.cpp
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
)
set (SRCS
Common/JitSymbols.cpp
Common/NetStream.cpp
Common/SoftFloat-3e/extF80_add.c
Common/SoftFloat-3e/extF80_div.c
Common/SoftFloat-3e/extF80_sub.c
@@ -71,7 +77,6 @@ set (SRCS
Common/SoftFloat-3e/f32_to_extF80.c
Common/SoftFloat-3e/s_normSubnormalF32Sig.c
Common/SoftFloat-3e/s_f32UIToCommonNaN.c
Interface/Config/Config.cpp
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
@@ -95,19 +100,7 @@ set (SRCS
Interface/Core/Dispatcher/Dispatcher.cpp
Interface/Core/Dispatcher/X86Dispatcher.cpp
Interface/Core/Dispatcher/Arm64Dispatcher.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp
Interface/Core/Interpreter/InterpreterFallbacks.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -121,6 +114,7 @@ set (SRCS
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/HLE/Thunks/Thunks.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
Interface/IR/IREmitter.cpp
@@ -141,12 +135,28 @@ set (SRCS
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
Utils/Allocator/64BitAllocator.cpp
Utils/FileLoading.cpp
Utils/LogManager.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
)
if (ENABLE_INTERPRETER)
list(APPEND SRCS
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/Interpreter/ALUOps.cpp
Interface/Core/Interpreter/AtomicOps.cpp
Interface/Core/Interpreter/BranchOps.cpp
Interface/Core/Interpreter/ConversionOps.cpp
Interface/Core/Interpreter/EncryptionOps.cpp
Interface/Core/Interpreter/F80Ops.cpp
Interface/Core/Interpreter/FlagOps.cpp
Interface/Core/Interpreter/MemoryOps.cpp
Interface/Core/Interpreter/MiscOps.cpp
Interface/Core/Interpreter/MoveOps.cpp
Interface/Core/Interpreter/VectorOps.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp)
@@ -194,7 +204,7 @@ if (ENABLE_JIT_ARM64)
Interface/Core/JIT/Arm64/VectorOps.cpp)
endif()
set (LIBS vixl dl fmt::fmt xxhash tiny-json)
set (LIBS vixl dl xxhash tiny-json)
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
endif()
@@ -248,6 +258,7 @@ set(OUTPUT_CONFIG_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigValues.inl")
set(OUTPUT_CONFIG_OPTION_NAME "${OUTPUT_CONFIG_FOLDER}/ConfigOptions.inl")
set(INPUT_CONFIG_NAME "${CMAKE_BINARY_DIR}/generated/Config/Config.json")
set(OUTPUT_MAN_NAME "${CMAKE_BINARY_DIR}/generated/FEX.1")
set(OUTPUT_MAN_NAME_COMPRESS "${CMAKE_BINARY_DIR}/generated/FEX.1.gz")
add_custom_target(CREATE_CONFIG_FOLDER ALL
COMMAND ${CMAKE_COMMAND} -E make_directory "${OUTPUT_CONFIG_FOLDER}")
@@ -263,6 +274,12 @@ add_custom_command(
"${OUTPUT_CONFIG_OPTION_NAME}"
)
add_custom_command(
OUTPUT "${OUTPUT_MAN_NAME_COMPRESS}"
DEPENDS "${OUTPUT_MAN_NAME}"
COMMAND "gzip" "-kf9n" "${OUTPUT_MAN_NAME}"
)
set_source_files_properties(${OUTPUT_CONFIG_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
@@ -270,32 +287,28 @@ set_source_files_properties(${OUTPUT_CONFIG_OPTION_NAME} PROPERTIES
set_source_files_properties(${OUTPUT_MAN_NAME} PROPERTIES
GENERATED TRUE)
set_source_files_properties(${OUTPUT_MAN_NAME_COMPRESS} PROPERTIES
GENERATED TRUE)
# Create the target
add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_CONFIG_NAME}"
DEPENDS "${OUTPUT_CONFIG_OPTION_NAME}"
DEPENDS "${OUTPUT_MAN_NAME}")
DEPENDS "${OUTPUT_MAN_NAME}"
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the man page
install(FILES ${OUTPUT_MAN_NAME} DESTINATION ${MAN_DIR}/man1)
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
check_cxx_compiler_flag(-fdiagnostics-color=always GCC_COLOR)
check_cxx_compiler_flag(-fcolor-diagnostics CLANG_COLOR)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
add_dependencies(${Name} CONFIG_INC)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
function(AddDefaultOptionsToTarget Name)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PRIVATE IncludePrivate/)
@@ -305,6 +318,7 @@ function(AddObject Name Type)
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
target_compile_definitions(${Name} PRIVATE ${DEFINES})
add_dependencies(${Name} CONFIG_INC)
target_compile_options(${Name}
PRIVATE
@@ -327,19 +341,6 @@ function(AddObject Name Type)
PRIVATE
"-fcolor-diagnostics")
endif()
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
set_target_properties(${Name} PROPERTIES C_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES CXX_VISIBILITY_PRESET hidden)
set_target_properties(${Name} PROPERTIES VISIBILITY_INLINES_HIDDEN TRUE)
target_include_directories(${Name} PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
target_include_directories(${Name} PUBLIC "${PROJECT_SOURCE_DIR}/include/")
target_include_directories(${Name} PUBLIC "${CMAKE_BINARY_DIR}/include/")
if (CMAKE_BUILD_TYPE MATCHES "RELEASE")
target_link_options(${Name}
@@ -351,6 +352,29 @@ function(AddLibrary Name Type)
endif()
endfunction()
# Build FEXCore_Config static library
add_library(FEXCore_Base STATIC ${FEXCORE_BASE_SRCS})
target_link_libraries(FEXCore_Base fmt::fmt tiny-json)
AddDefaultOptionsToTarget(FEXCore_Base)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
AddDefaultOptionsToTarget(${Name})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
AddDefaultOptionsToTarget(${Name})
endfunction()
AddObject(${PROJECT_NAME}_object OBJECT)
AddLibrary(${PROJECT_NAME} STATIC)
AddLibrary(${PROJECT_NAME}_shared SHARED)
+14 -10
View File
@@ -1,13 +1,16 @@
#pragma once
#include "Common/MathUtils.h"
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include <stdint.h>
#include <stdlib.h>
#include <type_traits>
namespace FEXCore {
template<typename T>
struct BitSet final {
using ElementType = T;
@@ -17,12 +20,12 @@ struct BitSet final {
ElementType *Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
void Free() {
@@ -61,8 +64,8 @@ struct BitSetView final {
ElementType *Memory;
void GetView(BitSet<T> &Set, uint64_t ElementOffset) {
LOGMAN_THROW_A((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
@@ -87,11 +90,12 @@ struct BitSetView final {
bool operator[](T Element) {
return Get(Element);
}
};
static_assert(sizeof(BitSet<uint32_t>) == sizeof(uintptr_t), "Needs to just be a pointer");
static_assert(std::is_trivially_copyable<BitSet<uint32_t>>::value, "Needs to trivially copyable");
static_assert(std::is_trivially_copyable_v<BitSet<uint32_t>>, "Needs to trivially copyable");
static_assert(sizeof(BitSetView<uint32_t>) == sizeof(uintptr_t), "Needs to just be a pointer");
static_assert(std::is_trivially_copyable<BitSetView<uint32_t>>::value, "Needs to trivially copyable");
static_assert(std::is_trivially_copyable_v<BitSetView<uint32_t>>, "Needs to trivially copyable");
} // namespace FEXCore
+17 -30
View File
@@ -1,66 +1,53 @@
#include "Common/JitSymbols.h"
#include <string>
#include <sstream>
#include <unistd.h>
namespace FEXCore {
JITSymbols::JITSymbols() {
std::stringstream PerfMap;
PerfMap << "/tmp/perf-" << getpid() << ".map";
#include <fmt/format.h>
fp = fopen(PerfMap.str().c_str(), "wb");
namespace FEXCore {
JITSymbols::JITSymbols() : fp{nullptr, std::fclose} {
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
fp.reset(fopen(PerfMap.c_str(), "wb"));
if (fp) {
// Disable buffering on this file
setvbuf(fp, nullptr, _IONBF, 0);
setvbuf(fp.get(), nullptr, _IONBF, 0);
}
}
JITSymbols::~JITSymbols() {
if (fp) {
fclose(fp);
}
}
JITSymbols::~JITSymbols() = default;
void JITSymbols::Register(void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << "JIT_0x" << GuestAddr << "_" << HostAddr << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
}
void JITSymbols::Register(void *HostAddr, uint32_t CodeSize, std::string const &Name) {
void JITSymbols::Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << Name << "_" << HostAddr << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{} {:x} {}_{}\n", HostAddr, CodeSize, Name, HostAddr);
}
void JITSymbols::RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name) {
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " " << Name << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{} {:x} {}\n", HostAddr, CodeSize, Name);
}
void JITSymbols::RegisterJITSpace(void *HostAddr, uint32_t CodeSize) {
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
if (!fp) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
std::stringstream String;
String << std::hex << HostAddr << " " << CodeSize << " FEXJIT" << std::endl;
fwrite(String.str().c_str(), 1, String.str().size(), fp);
fmt::print(fp.get(), "{} {:x} FEXJIT\n", HostAddr, CodeSize);
}
}
} // namespace FEXCore
+11 -6
View File
@@ -1,19 +1,24 @@
#pragma once
#include <cstdint>
#include <cstdio>
#include <string>
#include <memory>
#include <string_view>
namespace FEXCore {
class JITSymbols final {
public:
JITSymbols();
~JITSymbols();
void Register(void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterNamedRegion(void *HostAddr, uint32_t CodeSize, std::string const &Name);
void RegisterJITSpace(void *HostAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
private:
FILE* fp{};
using FILEPtr = std::unique_ptr<FILE, decltype(&std::fclose)>;
FILEPtr fp;
};
}
-13
View File
@@ -1,13 +0,0 @@
#pragma once
#include <stdint.h>
static inline uint64_t AlignUp(uint64_t value, uint64_t size) {
return value + (size - value % size) % size;
};
static inline uint64_t AlignDown(uint64_t value, uint64_t size) {
return value - value % size;
};
-42
View File
@@ -1,42 +0,0 @@
#pragma once
#include <array>
#include <iostream>
#include <iterator>
#include <string.h>
class NetStream : public std::iostream {
public:
NetStream(int socketfd) : std::iostream(new NetBuf(socketfd)) {}
virtual ~NetStream();
private:
class NetBuf : public std::streambuf {
public:
NetBuf(int socketfd) {
socket = socketfd;
reset_output_buffer();
}
virtual ~NetBuf();
protected:
virtual std::streamsize xsputn(const char* buffer, std::streamsize size);
virtual std::streambuf::int_type underflow();
virtual std::streambuf::int_type overflow(std::streambuf::int_type ch);
virtual int sync();
private:
void reset_output_buffer() {
// we always leave room for one extra char
setp(std::begin(output_buffer), std::end(output_buffer) -1);
}
int flushBuffer(const char *buffer, size_t size);
int socket;
std::array<char, 1400> output_buffer;
std::array<char, 1500> input_buffer; // enough for a typical packet
};
};
+1 -1
View File
@@ -72,7 +72,7 @@ namespace FEXCore::Paths {
// Ensure the folder structure is created for our Data
if (!std::filesystem::exists(*EntryCache, ec) &&
!std::filesystem::create_directories(*EntryCache, ec)) {
LogMan::Msg::D("Couldn't create EntryCache directory: '%s'", EntryCache->c_str());
LogMan::Msg::DFmt("Couldn't create EntryCache directory: '{}'", *EntryCache);
}
}
+276 -16
View File
@@ -57,51 +57,183 @@ struct X80SoftFloat {
// Ops
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
faddp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_add(lhs, rhs);
#endif
}
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fsubp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_sub(lhs, rhs);
#endif
}
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fmulp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_mul(lhs, rhs);
#endif
}
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fdivp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_div(lhs, rhs);
#endif
}
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
X80SoftFloat Rem = extF80_rem(lhs, rhs);
if (SignBit(Rem)) {
Rem = extF80_add(Rem, rhs);
}
else {
Rem.Sign = SignBit(lhs);
}
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Rem;
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fprem1;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
return extF80_rem(lhs, rhs);
#endif
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Tmp = lhs;
Tmp.Exponent = 0x3FFF;
Tmp.Sign = lhs.Sign;
return Tmp;
#endif
}
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fxtract;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
return Result;
#else
int32_t TrueExp = lhs.Exponent - ExponentBias;
return i32_to_extF80(TrueExp);
#endif
}
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
@@ -111,62 +243,190 @@ struct X80SoftFloat {
}
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FSCALE which may have accuracy problems");
X80SoftFloat Int = FRNDINT(rhs);
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fscale; # st0 = st0 * 2^(rdint(st1))
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
BIGFLOAT Src2_d = Int;
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
return Result;
#endif
}
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used F2XM1 which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
f2xm1; # st0 = 2^st(0) - 1
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
#endif
}
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FYL2X which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[rhs]; # st(1)
fldt %[lhs]; # st(0)
fyl2x; # st(1) * log2l(st(0))
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = Src2_d * log2l(Src1_d);
return Tmp;
#endif
}
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
WARN_ONCE("x87: Application used FATAN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs];
fldt %[rhs];
fpatan;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
return Result;
#else
BIGFLOAT Src1_d = lhs;
BIGFLOAT Src2_d = rhs;
BIGFLOAT Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
#endif
}
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FTAN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fptan;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = tanl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FSIN which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsin;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = sinl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
WARN_ONCE("x87: Application used FCOS which may have accuracy problems");
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fcos;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
BIGFLOAT Src_d = lhs;
Src_d = cosl(Src_d);
return Src_d;
#endif
}
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
fninit;
fldt %[lhs]; # st0
fsqrt;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
return Result;
#else
return extF80_sqrt(lhs);
#endif
}
operator float() const {
+29
View File
@@ -0,0 +1,29 @@
#pragma once
#include <string>
namespace FEXCore::StringUtils {
// Trim the left side of the string of whitespace and new lines
[[maybe_unused]] static std::string LeftTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(TrimTokens)) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
// Trim the right side of the string of whitespace and new lines
[[maybe_unused]] static std::string RightTrim(std::string String, std::string TrimTokens = " \t\n\r") {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(TrimTokens)) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
// Trim both the left and right of the string of whitespace and new lines
[[maybe_unused]] static std::string Trim(std::string String, std::string TrimTokens = " \t\n\r") {
return RightTrim(LeftTrim(String, TrimTokens), TrimTokens);
}
}
+29 -30
View File
@@ -1,4 +1,5 @@
#include "Common/StringConv.h"
#include "Common/StringUtils.h"
#include "Common/Paths.h"
#include "Utils/FileLoading.h"
@@ -74,14 +75,14 @@ namespace JSON {
json_t const *json = json_createWithPool(&Data.at(0), &Pool.PoolObject);
if (!json) {
LogMan::Msg::E("Couldn't create json");
LogMan::Msg::EFmt("Couldn't create json");
return;
}
json_t const* ConfigList = json_getProperty(json, "Config");
if (!ConfigList) {
LogMan::Msg::E("Couldn't get config list");
LogMan::Msg::EFmt("Couldn't get config list");
return;
}
@@ -92,12 +93,12 @@ namespace JSON {
const char* ConfigString = json_getValue(ConfigItem);
if (!ConfigName) {
LogMan::Msg::E("Couldn't get config name");
LogMan::Msg::EFmt("Couldn't get config name");
return;
}
if (!ConfigString) {
LogMan::Msg::E("Couldn't get ConfigString for '%s'", ConfigName);
LogMan::Msg::EFmt("Couldn't get ConfigString for '{}'", ConfigName);
return;
}
@@ -173,7 +174,7 @@ namespace JSON {
if (!Global &&
!std::filesystem::exists(ConfigFile, ec) &&
!std::filesystem::create_directories(ConfigFile, ec)) {
LogMan::Msg::D("Couldn't create config directory: '%s'", ConfigFile.c_str());
LogMan::Msg::DFmt("Couldn't create config directory: '{}'", ConfigFile);
// Let's go local in this case
return "./" + Filename + ".json";
}
@@ -370,29 +371,6 @@ namespace JSON {
return {};
}
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
const static std::string ContainerManager = "/run/host/container-manager";
@@ -401,7 +379,7 @@ namespace JSON {
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
std::string ManagerStr = Manager.data();
ManagerStr = trim(ManagerStr);
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
@@ -416,7 +394,9 @@ namespace JSON {
Meta->Load();
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THREADS)) {
{
// Always fix up the number of threads and create the configuration
// Otherwise the application could receive zero as the number of threads
FEX_CONFIG_OPT(Cores, THREADS);
if (Cores == 0) {
// When the number of emulated CPU cores is zero then auto detect
@@ -424,6 +404,25 @@ namespace JSON {
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CORE)) {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
#if (_M_X86_64)
constexpr uint32_t MaxCoreNumber = 2;
#else
constexpr uint32_t MaxCoreNumber = 1;
#endif
#ifdef INTERPRETER_ENABLED
constexpr uint32_t MinCoreNumber = 0;
#else
constexpr uint32_t MinCoreNumber = 1;
#endif
if (Core > MaxCoreNumber || Core < MinCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, std::to_string(FEXCore::Config::CONFIG_IRJIT));
}
}
std::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, std::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
+10 -1
View File
@@ -32,7 +32,7 @@
},
"Threads": {
"Type": "uint32",
"Default": "1",
"Default": "0",
"ShortArg": "T",
"Desc": [
"Number of physical hardware threads to tell the process we have.",
@@ -201,6 +201,15 @@
"Disables logging"
]
},
"OutputSocket": {
"Type": "str",
"Default": "",
"Desc": [
"Socket to connect to",
"eg: localhost:8087",
"If set will override the OutputLog location"
]
},
"OutputLog": {
"Type": "str",
"Default": "stderr",
+20 -12
View File
@@ -51,7 +51,7 @@ namespace FEXCore::Context {
CTX->CustomExitHandler = std::move(handler);
}
ExitHandler GetExitHandler(FEXCore::Context::Context *CTX) {
ExitHandler GetExitHandler(const FEXCore::Context::Context *CTX) {
return CTX->CustomExitHandler;
}
@@ -71,23 +71,23 @@ namespace FEXCore::Context {
return CTX->RunUntilExit();
}
int GetProgramStatus(FEXCore::Context::Context *CTX) {
int GetProgramStatus(const FEXCore::Context::Context *CTX) {
return CTX->GetProgramStatus();
}
FEXCore::Context::ExitReason GetExitReason(FEXCore::Context::Context *CTX) {
FEXCore::Context::ExitReason GetExitReason(const FEXCore::Context::Context *CTX) {
return CTX->ParentThread->ExitReason;
}
bool IsDone(FEXCore::Context::Context *CTX) {
bool IsDone(const FEXCore::Context::Context *CTX) {
return CTX->IsPaused();
}
void GetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
void GetCPUState(const FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
memcpy(State, CTX->ParentThread->CurrentFrame, sizeof(FEXCore::Core::CPUState));
}
void SetCPUState(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *State) {
void SetCPUState(FEXCore::Context::Context *CTX, const FEXCore::Core::CPUState *State) {
memcpy(CTX->ParentThread->CurrentFrame, State, sizeof(FEXCore::Core::CPUState));
}
@@ -115,11 +115,11 @@ namespace FEXCore::Context {
}
void RegisterHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterHostSignalHandler(Signal, Func, Required);
CTX->RegisterHostSignalHandler(Signal, std::move(Func), Required);
}
void RegisterFrontendHostSignalHandler(FEXCore::Context::Context *CTX, int Signal, HostSignalDelegatorFunction Func, bool Required) {
CTX->RegisterFrontendHostSignalHandler(Signal, Func, Required);
CTX->RegisterFrontendHostSignalHandler(Signal, std::move(Func), Required);
}
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Context::Context *CTX, FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) {
@@ -162,12 +162,20 @@ namespace FEXCore::Context {
return CTX->CPUID.RunFunction(Function, Leaf);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->AOTIRLoader = CacheReader;
FEX_DEFAULT_VISIBILITY FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CTX->CPUID.RunFunctionName(Function, Leaf, CPU);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter) {
CTX->AOTIRWriter = CacheWriter;
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<int(const std::string&)> CacheReader) {
CTX->SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
CTX->SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(FEXCore::Context::Context *CTX, std::function<void(const std::string&)> CacheRenamer) {
CTX->SetAOTIRRenamer(CacheRenamer);
}
void FinalizeAOTIRCache(FEXCore::Context::Context *CTX) {
+22 -68
View File
@@ -4,6 +4,7 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
@@ -57,38 +58,6 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct AOTIRInlineEntry {
uint64_t GuestHash;
uint64_t GuestLength;
/* RAData followed by IRData */
uint8_t InlineData[0];
IR::RegisterAllocationData *GetRAData();
IR::IRListView *GetIRData();
};
struct AOTIRInlineIndexEntry {
uint64_t GuestStart;
uint64_t DataOffset;
};
struct AOTIRInlineIndex {
uint64_t Count;
uint64_t DataBase;
AOTIRInlineIndexEntry Entries[0];
AOTIRInlineEntry *Find(uint64_t GuestStart);
AOTIRInlineEntry *GetInlineEntry(uint64_t DataOffset);
};
struct AOTIRCaptureCacheEntry {
std::unique_ptr<std::ostream> Stream;
std::map<uint64_t, uint64_t> Index;
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData);
};
struct Context {
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
@@ -156,32 +125,6 @@ namespace FEXCore::Context {
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
struct AOTIRCacheEntry {
AOTIRInlineIndex *Array;
void *mapping;
size_t size;
};
std::unordered_map<std::string, AOTIRCacheEntry> AOTIRCache;
std::function<int(const std::string&)> AOTIRLoader;
std::function<std::unique_ptr<std::ostream>(const std::string&)> AOTIRWriter;
std::unordered_map<std::string, AOTIRCaptureCacheEntry> AOTIRCaptureCache;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
std::string filename;
void *CachedFileEntry;
bool ContainsCode;
};
using AddrToFileMapType = std::map<uint64_t, AddrToFileEntry>;
AddrToFileMapType AddrToFile;
std::map<std::string, std::string> FilesWithCode;
AddrToFileMapType::iterator FindAddrForFile(uint64_t Entry, uint64_t Length);
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
@@ -253,9 +196,6 @@ namespace FEXCore::Context {
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
bool LoadAOTIRCache(int streamfd);
void FinalizeAOTIRCache();
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer);
/**
* @brief Initializes the JIT compilers for the thread
*
@@ -337,6 +277,26 @@ namespace FEXCore::Context {
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void FinalizeAOTIRCache() {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void SetAOTIRLoader(std::function<int(const std::string&)> CacheReader) {
IRCaptureCache.SetAOTIRLoader(CacheReader);
}
void SetAOTIRWriter(std::function<std::unique_ptr<std::ofstream>(const std::string&)> CacheWriter) {
IRCaptureCache.SetAOTIRWriter(CacheWriter);
}
void SetAOTIRRenamer(std::function<void(const std::string&)> CacheRenamer) {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
@@ -362,13 +322,7 @@ namespace FEXCore::Context {
std::mutex ExitMutex;
std::unique_ptr<GdbServer> DebugServer;
std::shared_mutex AOTIRCacheLock;
std::shared_mutex AOTIRCaptureCacheWriteoutLock;
std::atomic<bool> AOTIRCaptureCacheWriteoutFlusing;
std::queue<std::function<void()>> AOTIRCaptureCacheWriteoutQueue;
void AOTIRCaptureCacheWriteoutQueue_Flush();
void AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn);
IR::AOTIRCaptureCache IRCaptureCache;
bool StartPaused = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
+11 -18
View File
@@ -535,7 +535,6 @@ uint64_t HandleCASPAL_ARMv8(void *_ucontext, void *_info, uint32_t Instr) {
}
bool HandleAtomicVectorStore(void *_ucontext, void *_info, uint32_t Instr) {
mcontext_t* mcontext = &reinterpret_cast<ucontext_t*>(_ucontext)->uc_mcontext;
siginfo_t* info = reinterpret_cast<siginfo_t*>(_info);
if (info->si_code != BUS_ADRALN) {
@@ -1437,9 +1436,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS16<true>(
@@ -1499,9 +1497,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS32<true>(
@@ -1561,9 +1558,8 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
DesiredFunction = SWAPDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", Op);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", Op);
return false;
break;
}
auto Res = DoCAS64<true>(
@@ -1744,8 +1740,8 @@ static uint64_t HandleCAS_NoAtomics(void *_ucontext, void *_info)
if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
const uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A_FMT(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
#endif
DesiredReg = GetRdReg(NextInstr);
}
@@ -1865,8 +1861,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
else if ((NextInstr & FEXCore::ArchHelpers::Arm64::STLXR_MASK) == FEXCore::ArchHelpers::Arm64::STLXR_INST) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Just double check that the memory destination matches
uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
const uint32_t StoreAddressReg = GetRnReg(NextInstr);
LOGMAN_THROW_A_FMT(StoreAddressReg == AddressReg, "StoreExclusive memory register didn't match the store exclusive register");
#endif
uint32_t StatusReg = GetRmReg(NextInstr);
uint32_t StoreResultReg = GetRdReg(NextInstr);
@@ -1885,7 +1881,7 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
break;
}
else {
LogMan::Msg::A("Unknown instruction 0x%08x", NextInstr);
LogMan::Msg::AFmt("Unknown instruction 0x{:08x}", NextInstr);
}
}
@@ -1951,9 +1947,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS16<DoRetry>(
@@ -2028,9 +2023,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS32<DoRetry>(
@@ -2105,9 +2099,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
return false;
break;
}
auto Res = DoCAS64<DoRetry>(
@@ -1,7 +1,9 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include "aarch64/cpu-aarch64.h"
#include "cpu-features.h"
@@ -14,53 +16,44 @@ namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size) : vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode) {
CPU.SetUp();
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
if (SupportsAtomics) {
if (ctx->HostFeatures.SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
SetCPUFeatures(Features);
if (!SupportsAtomics) {
WARN_ONCE("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
#endif
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
if (Is64Bit && ((~Constant)>> 16) == 0) {
movn(Reg, (~Constant) & 0xFFFF);
if (NOPPad) {
nop(); nop(); nop();
}
return;
}
int NumMoves = 1;
movz(Reg, (Constant) & 0xFFFF, 0);
for (int i = 1; i < Segments; ++i) {
uint16_t Part = (Constant >> (i * 16)) & 0xFFFF;
if (Part) {
movk(Reg, Part, i * 16);
++NumMoves;
}
}
if (NOPPad) {
for (int i = NumMoves; i < Segments; ++i) {
nop();
}
}
}
@@ -69,7 +62,7 @@ void Arm64Emitter::PushCalleeSavedRegisters() {
// We need to save pairs of registers
// We save r19-r30
MemOperand PairOffset(sp, -16, PreIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x19, x20},
{x21, x22},
{x23, x24},
@@ -132,7 +125,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
MemOperand PairOffset(sp, 16, PostIndex);
const std::array<std::pair<vixl::aarch64::XRegister, vixl::aarch64::XRegister>, 6> CalleeSaved = {{
const std::array<std::pair<vixl::aarch64::Register, vixl::aarch64::Register>, 6> CalleeSaved = {{
{x29, x30},
{x27, x28},
{x25, x26},
@@ -147,26 +140,76 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
void Arm64Emitter::SpillStaticRegs() {
void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
stp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & GPRSpillMask) &&
((1U << Reg2.GetCode()) & GPRSpillMask)) {
stp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & GPRSpillMask)) {
str(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & GPRSpillMask)) {
str(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
stp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (((1U << Reg1.GetCode()) & FPRSpillMask) &&
((1U << Reg2.GetCode()) & FPRSpillMask)) {
stp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRSpillMask)) {
str(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRSpillMask)) {
str(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
}
}
}
}
}
void Arm64Emitter::FillStaticRegs() {
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
if (StaticRegisterAllocation()) {
for (size_t i = 0; i < SRA64.size(); i+=2) {
ldp(SRA64[i], SRA64[i+1], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
auto Reg1 = SRA64[i];
auto Reg2 = SRA64[i+1];
if (((1U << Reg1.GetCode()) & GPRFillMask) &&
((1U << Reg2.GetCode()) & GPRFillMask)) {
ldp(Reg1, Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg1.GetCode()) & GPRFillMask)) {
ldr(Reg1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i])));
}
else if (((1U << Reg2.GetCode()) & GPRFillMask)) {
ldr(Reg2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[i+1])));
}
}
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
ldp(SRAFPR[i].Q(), SRAFPR[i+1].Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (((1U << Reg1.GetCode()) & FPRFillMask) &&
((1U << Reg2.GetCode()) & FPRFillMask)) {
ldp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRFillMask)) {
ldr(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRFillMask)) {
ldr(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
}
}
}
}
}
@@ -58,15 +58,18 @@ const std::array<aarch64::VRegister, 12> RAFPR = {
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(size_t size);
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
vixl::aarch64::CPU CPU;
bool SupportsAtomics{};
bool SupportsRCPC{};
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant);
void SpillStaticRegs();
void FillStaticRegs();
static constexpr uint32_t CALLER_GPR_MASK = 0b0011'1111'1111'1111'1111;
// This isn't technically true because the lower 64-bits of v8..v15 are callee saved
// We can't guarantee only the lower 64bits are used so flush everything
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
@@ -79,9 +82,6 @@ protected:
uint32_t SpillSlots{};
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
@@ -12,15 +12,15 @@ namespace FEXCore::ArchHelpers::Arm64 {
// Obvously such a configuration can't do the actual arm64-specific stuff
bool HandleCASPAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASPAL Not Implemented");
ERROR_AND_DIE_FMT("HandleCASPAL Not Implemented");
}
bool HandleCASAL(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleCASAL Not Implemented");
ERROR_AND_DIE_FMT("HandleCASAL Not Implemented");
}
bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE("HandleAtomicMemOp Not Implemented");
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
#endif
+31 -11
View File
@@ -13,17 +13,27 @@
namespace FEXCore::ArchHelpers::Context {
enum ContextFlags : uint32_t {
CONTEXT_FLAG_INJIT = (1U << 0),
CONTEXT_FLAG_32BIT = (1U << 1),
};
struct X86ContextBackup {
// Host State
// RIP and RSP is stored in GPRs here
uint64_t GPRs[23];
FEXCore::x86_64::_libc_fpstate FPRState;
uint64_t sa_mask;
bool FaultToTopAndGeneratedException;
// Guest state
int Signal;
uint32_t Flags;
uint64_t OriginalRIP;
uint64_t FPStateLocation;
uint64_t UContextLocation;
uint64_t SigInfoLocation;
FEXCore::Core::CPUState GuestState;
static constexpr int RedZoneSize = 128;
};
@@ -37,9 +47,15 @@ struct ArmContextBackup {
uint32_t FPCR;
__uint128_t FPRs[32];
uint64_t sa_mask;
bool FaultToTopAndGeneratedException;
// Guest state
int Signal;
uint32_t Flags;
uint64_t OriginalRIP;
uint64_t FPStateLocation;
uint64_t UContextLocation;
uint64_t SigInfoLocation;
FEXCore::Core::CPUState GuestState;
// Arm64 doesn't have a red zone
@@ -108,7 +124,7 @@ static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
auto MContext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&MContext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
return HostState->FPRs[id];
}
@@ -127,7 +143,7 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
@@ -135,7 +151,8 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -146,7 +163,7 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x%08x", HostState->Head.Magic);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Backup->FPRs[0], 32 * sizeof(__uint128_t));
HostState->FPCR = Backup->FPCR;
HostState->FPSR = Backup->FPSR;
@@ -160,7 +177,8 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -193,15 +211,15 @@ static inline void SetState(void* ucontext, uint64_t val) {
}
static inline uint64_t GetArmReg(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not impelented for x86 host");
ERROR_AND_DIE_FMT("Not impelented for x86 host");
}
static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
ERROR_AND_DIE("Not impelented for x86 host");
ERROR_AND_DIE_FMT("Not impelented for x86 host");
}
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
ERROR_AND_DIE("Not implemented for x86 host");
ERROR_AND_DIE_FMT("Not implemented for x86 host");
}
using ContextBackup = X86ContextBackup;
@@ -220,7 +238,8 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Save the signal mask so we can restore it
memcpy(&Backup->sa_mask, &_ucontext->uc_sigmask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -238,7 +257,8 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
// Restore the signal mask now
memcpy(&_ucontext->uc_sigmask, &Backup->sa_mask, sizeof(uint64_t));
} else {
ERROR_AND_DIE("Wrong context type"); // This must be a runtime error
// This must be a runtime error
ERROR_AND_DIE_FMT("Wrong context type");
}
}
@@ -27,7 +27,7 @@ namespace FEXCore {
<< std::endl;
}
Output.close();
LogMan::Msg::D("Dumped %d blocks of sampling data", SamplingMap.size());
LogMan::Msg::DFmt("Dumped {} blocks of sampling data", SamplingMap.size());
}
BlockSamplingData::BlockData *BlockSamplingData::GetBlockData(uint64_t RIP) {
+420 -91
View File
@@ -5,14 +5,16 @@ desc: Handles presented capability bits for guest cpu
$end_info$
*/
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include "Common/StringConv.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXHeaderUtils/Syscalls.h>
#include "git_version.h"
#include <cstring>
@@ -21,6 +23,66 @@ $end_info$
#endif
namespace FEXCore {
namespace ProductNames {
#ifdef _M_ARM_64
static const char ARM_UNKNOWN[] = "Unknown ARM CPU";
static const char ARM_A57[] = "Cortex-A57";
static const char ARM_A72[] = "Cortex-A72";
static const char ARM_A73[] = "Cortex-A73";
static const char ARM_A75[] = "Cortex-A75";
static const char ARM_A76[] = "Cortex-A76";
static const char ARM_A76AE[] = "Cortex-A76AE";
static const char ARM_V1[] = "Neoverse V1";
static const char ARM_A77[] = "Cortex-A77";
static const char ARM_A78[] = "Cortex-A78";
static const char ARM_A78AE[] = "Cortex-A78AE";
static const char ARM_A78C[] = "Cortex-A78C";
static const char ARM_A710[] = "Cortex-A710";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
static const char ARM_E1[] = "Neoverse E1";
static const char ARM_A35[] = "Cortex-A35";
static const char ARM_A53[] = "Cortex-A53";
static const char ARM_A55[] = "Cortex-A55";
static const char ARM_A65[] = "Cortex-A65";
static const char ARM_A510[] = "Cortex-A510";
static const char ARM_Kryo200[] = "Kryo 2xx";
static const char ARM_Kryo300[] = "Kryo 3xx";
static const char ARM_Kryo400[] = "Kryo 4xx/5xx";
static const char ARM_Kryo200S[] = "Kryo 2xx S";
static const char ARM_Kryo300S[] = "Kryo 3xx S";
static const char ARM_Kryo400S[] = "Kryo 4xx/5xx S";
static const char ARM_Denver[] = "Nvidia Denver";
static const char ARM_Carmel[] = "Nvidia Carmel";
static const char ARM_Firestorm[] = "Apple Firestorm";
static const char ARM_Icestorm[] = "Apple Icestorm";
#else
static const char UNKNOWN[] = "Unknown CPU";
#endif
}
static uint32_t GetCPUID() {
uint32_t CPU{};
FHU::Syscalls::getcpu(&CPU, nullptr);
return CPU;
}
static uint32_t CalculateNumberOfCPUs() {
size_t CPUs = 1;
while(std::filesystem::exists("/sys/devices/system/cpu/cpu" + std::to_string(CPUs))) {
CPUs++;
}
return CPUs;
}
constexpr uint32_t SUPPORTS_AVX = 0;
// #define CPUID_AMD
#ifdef CPUID_AMD
@@ -49,60 +111,242 @@ static uint32_t GetCycleCounterFrequency() {
return Result;
}
static bool GetHostHybridFlag() {
int MaxCPUs = 64;
size_t AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
cpu_set_t *Set = CPU_ALLOC(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
void CPUIDEmu::SetupHostHybridFlag() {
size_t CPUs = CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
int Result{};
for (;;) {
Result = sched_getaffinity(0, AllocSize, Set);
if (Result == 0 ||
(Result == -1 && errno != EINVAL)) {
break;
}
MaxCPUs <<= 1;
CPU_FREE(Set);
Set = CPU_ALLOC(MaxCPUs);
AllocSize = CPU_ALLOC_SIZE(MaxCPUs);
CPU_ZERO_S(AllocSize, Set);
}
if (Result != 0) {
return false;
}
int CPUs = CPU_COUNT_S(AllocSize, Set);
bool Hybrid = false;
uint64_t MIDR{};
for (int i = 0; i < CPUs; ++i) {
if (CPU_ISSET_S(i, AllocSize, Set)) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
if (FEXCore::StrConv::Conv(&Data.at(0), &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
break;
}
MIDR = NewMIDR;
for (size_t i = 0; i < CPUs; ++i) {
std::error_code ec{};
std::string MIDRPath = "/sys/devices/system/cpu/cpu" + std::to_string(i) + "/regs/identification/midr_el1";
if (std::filesystem::exists(MIDRPath, ec)) {
std::vector<char> Data{};
// Needs to be a fixed size since depending on kernel it will try to read a full page of data and fail
// Only read 18 bytes for a 64bit value prefixed with 0x
if (FEXCore::FileLoading::LoadFile(Data, MIDRPath, 18)) {
uint64_t NewMIDR{};
std::string_view MIDRView(&Data.at(0), 18);
if (FEXCore::StrConv::Conv(MIDRView, &NewMIDR)) {
if (MIDR != 0 && MIDR != NewMIDR) {
// CPU mismatch, claim hybrid
Hybrid = true;
}
// Truncate to 32-bits, top 32-bits are all reserved in MIDR
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
PerCPUData[i].MIDR = NewMIDR;
MIDR = NewMIDR;
}
}
}
}
CPU_FREE(Set);
return Hybrid;
struct CPUMIDR {
uint8_t Implementer;
uint16_t Part;
bool DefaultBig; // Defaults to a big core
const char *ProductName{};
};
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 35> CPUMIDRs = {{
// Typically big CPU cores
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd4b, 1, ProductNames::ARM_A78C}, // A78C
{0x41, 0xd4a, 1, ProductNames::ARM_E1}, // E1
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd48, 1, ProductNames::ARM_X2}, // X2
{0x41, 0xd47, 1, ProductNames::ARM_A710}, // A710
{0x41, 0xd44, 1, ProductNames::ARM_X1}, // X1
{0x41, 0xd42, 1, ProductNames::ARM_A78AE}, // A78AE
{0x41, 0xd41, 1, ProductNames::ARM_A78}, // A78
{0x41, 0xd40, 1, ProductNames::ARM_V1}, // V1
{0x41, 0xd0e, 1, ProductNames::ARM_A76AE}, // A76AE
{0x41, 0xd0d, 1, ProductNames::ARM_A77}, // A77
{0x41, 0xd0c, 1, ProductNames::ARM_N1}, // N1
{0x41, 0xd0b, 1, ProductNames::ARM_A76}, // A76
{0x51, 0x804, 1, ProductNames::ARM_Kryo400}, // Kryo 4xx Gold (A76 based)
{0x41, 0xd0a, 1, ProductNames::ARM_A75}, // A75
{0x51, 0x802, 1, ProductNames::ARM_Kryo300}, // Kryo 3xx Gold (A75 based)
{0x41, 0xd09, 1, ProductNames::ARM_A73}, // A73
{0x51, 0x800, 1, ProductNames::ARM_Kryo200}, // Kryo 2xx Gold (A73 based)
{0x41, 0xd08, 1, ProductNames::ARM_A72}, // A72
{0x4e, 0x004, 1, ProductNames::ARM_Carmel}, // Carmel
// Denver rated above A57 to match TX2 weirdness
{0x4e, 0x003, 1, ProductNames::ARM_Denver}, // Denver
{0x41, 0xd07, 1, ProductNames::ARM_A57}, // A57
// Typically Little CPU cores
{0x61, 0x022, 0, ProductNames::ARM_Icestorm}, // Apple M1 Icestorm
{0x41, 0xd46, 0, ProductNames::ARM_A510}, // A510
{0x41, 0xd06, 0, ProductNames::ARM_A65}, // A65
{0x41, 0xd05, 0, ProductNames::ARM_A55}, // A55
{0x51, 0x805, 0, ProductNames::ARM_Kryo400S}, // Kryo 4xx/5xx Silver (A55 based)
{0x51, 0x803, 0, ProductNames::ARM_Kryo300S}, // Kryo 3xx Silver (A55 based)
{0x41, 0xd03, 0, ProductNames::ARM_A53}, // A53
{0x51, 0x801, 0, ProductNames::ARM_Kryo200S}, // Kryo 2xx Silver (A53 based)
{0x41, 0xd04, 0, ProductNames::ARM_A35}, // A35
{0x41, 0, 0, ProductNames::ARM_UNKNOWN}, // Invalid CPU or Apple CPU inside Parallels VM
{0x0, 0, 0, ProductNames::ARM_UNKNOWN}, // Invalid starting point is lowest ranked
}};
auto FindDefinedMIDR = [](uint32_t MIDR) -> const CPUMIDR* {
uint8_t Implementer = MIDR >> 24;
uint16_t Part = (MIDR >> 4) & 0xFFF;
for (auto &MIDROption : CPUMIDRs) {
if (MIDROption.Implementer == Implementer &&
MIDROption.Part == Part) {
return &MIDROption;
}
}
return nullptr;
};
if (Hybrid) {
// Walk the MIDRs and calculate big little designs
std::vector<const CPUMIDR*> BigCores;
std::vector<const CPUMIDR*> LittleCores;
// Separate CPU cores out to big or little selected
for (size_t i = 0; i < CPUs; ++i) {
uint32_t MIDR = PerCPUData[i].MIDR;
auto MIDROption = FindDefinedMIDR(MIDR);
if (MIDROption) {
// Found one
if (MIDROption->DefaultBig) {
BigCores.emplace_back(MIDROption);
}
else {
LittleCores.emplace_back(MIDROption);
}
}
else {
// If we didn't insert this MIDR then claim it is a little core.
LittleCores.emplace_back(&CPUMIDRs.back());
}
}
if (LittleCores.empty()) {
// If we only ended up with big cores then we need to move some to be little cores
uint32_t LowestMIDR = ~0U;
uint32_t LowestMIDRIdx = 0;
// Walk all the big cores
for (size_t i = 0; i < BigCores.size(); ++i) {
uint8_t Implementer = BigCores[i]->Implementer;
uint16_t Part = BigCores[i]->Part;
// Walk our list of CPUMIDRs to find the most little core
for (size_t j = LowestMIDRIdx; j < CPUMIDRs.size(); ++j) {
auto &MIDROption = CPUMIDRs[i];
if ((MIDROption.Implementer == Implementer &&
MIDROption.Part == Part) ||
(MIDROption.Implementer == 0 &&
MIDROption.Part == 0)) {
LowestMIDRIdx = j;
LowestMIDR = MIDR;
break;
}
}
}
// Now we WILL have found a big core to demote to little status
// Demote them
std::erase_if(BigCores, [&LittleCores, LowestMIDR](auto *Entry) {
// Demote by erase copy to little array
uint8_t Implementer = LowestMIDR >> 24;
uint16_t Part = (LowestMIDR >> 4) & 0xFFF;
if (Entry->Implementer == Implementer &&
Entry->Part == Part) {
// Add it to the BigCore list
LittleCores.emplace_back(Entry);
return true;
}
return false;
});
}
if (BigCores.empty()) {
// We never found a CPU core we understand
// Grab the first core, consider it as little, move everything else to Big
uint32_t LittleMIDR = PerCPUData[0].MIDR;
// Now walk the little cores and move them to Big if they don't match
std::erase_if(LittleCores, [&BigCores, LittleMIDR](auto *Entry) {
// You're promoted now
uint8_t Implementer = LittleMIDR >> 24;
uint16_t Part = (LittleMIDR >> 4) & 0xFFF;
if (Entry->Implementer != Implementer ||
Entry->Part != Part) {
// Add it to the BigCore list
BigCores.emplace_back(Entry);
return true;
}
return false;
});
}
// Now walk the per CPU data one more time and set if it is big or little
for (auto &Data : PerCPUData) {
uint8_t Implementer = Data.MIDR >> 24;
uint16_t Part = (Data.MIDR >> 4) & 0xFFF;
bool FoundBig{};
const CPUMIDR *MIDR{};
for (auto Big : BigCores) {
if (Big->Implementer == Implementer &&
Big->Part == Part) {
FoundBig = true;
MIDR = Big;
break;
}
}
if (!FoundBig) {
for (auto Little : LittleCores) {
if (Little->Implementer == Implementer &&
Little->Part == Part) {
MIDR = Little;
break;
}
}
}
Data.IsBig = FoundBig;
if (MIDR) {
Data.ProductName = MIDR->ProductName ?: ProductNames::ARM_UNKNOWN;
}
else {
Data.ProductName = ProductNames::ARM_UNKNOWN;
}
}
}
else {
// If we aren't hybrid then just claim everything is big
for (size_t i = 0; i < CPUs; ++i) {
uint32_t MIDR = PerCPUData[i].MIDR;
auto MIDROption = FindDefinedMIDR(MIDR);
PerCPUData[i].IsBig = true;
if (MIDROption) {
PerCPUData[i].ProductName = MIDROption->ProductName ?: ProductNames::ARM_UNKNOWN;
}
else {
PerCPUData[i].ProductName = ProductNames::ARM_UNKNOWN;
}
}
}
}
#else
@@ -119,16 +363,21 @@ static uint32_t GetCycleCounterFrequency() {
return 0;
}
static bool GetHostHybridFlag() {
void CPUIDEmu::SetupHostHybridFlag() {
uint32_t eax, ebx, ecx, edx;
__cpuid(0, eax, ebx, ecx, edx);
if (eax >= 0x7) {
__cpuid(0x7, eax, ebx, ecx, edx);
// Bit 15 of edx claims hybrid CPU
return (edx & (1U << 15)) != 0;
Hybrid = (edx & (1U << 15)) != 0;
}
return false;
size_t CPUs = CalculateNumberOfCPUs();
PerCPUData.resize(CPUs);
for (size_t i = 0; i < CPUs; ++i) {
PerCPUData[i].IsBig = true;
PerCPUData[i].ProductName = ProductNames::UNKNOWN;
}
}
#endif
@@ -155,6 +404,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
uint32_t CoreCount = Cores();
// XXX: Enable once the rest of the SSE4.2 instructions are emulated
uint32_t SupportsSSE42 = CTX->HostFeatures.SupportsCRC && false ? 1 : 0;
Res.eax = FAMILY_IDENTIFIER;
@@ -184,7 +435,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(0 << 17) | // Process-context identifiers
(0 << 18) | // Prefetching from memory mapped device
(1 << 19) | // SSE4.1
(0 << 20) | // SSE4.2
(SupportsSSE42 << 20) | // SSE4.2
(0 << 21) | // X2APIC
(1 << 22) | // MOVBE
(1 << 23) | // POPCNT
@@ -194,7 +445,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
(0 << 27) | // OSXSAVE
(SUPPORTS_AVX << 28) | // AVX
(0 << 29) | // F16C
(0 << 30) | // RDRAND
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(0 << 31); // Hypervisor always returns zero
Res.edx =
@@ -386,7 +637,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(0 << 8) | // BMI2
(1 << 8) | // BMI2
(0 << 9) | // Enhanced REP MOVSB/STOSB
(1 << 10) | // INVPCID for system software control of process-context
(0 << 11) | // Restricted transactional memory
@@ -396,8 +647,8 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) {
(0 << 15) | // Intel Resource Directory Technology Allocation
(0 << 16) | // Reserved
(0 << 17) | // Reserved
(0 << 18) | // RDSEED
(0 << 19) | // ADCX and ADOX instructions
(CTX->HostFeatures.SupportsRAND << 18) | // RDSEED
(1 << 19) | // ADCX and ADOX instructions
(0 << 20) | // SMAP Supervisor mode access prevention and CLAC/STAC instructions
(0 << 21) | // Reserved
(0 << 22) | // Reserved
@@ -549,6 +800,59 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) {
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_1Ah(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Hybrid) {
uint32_t CPU = GetCPUID();
auto &Data = PerCPUData[CPU];
// 0x40 is a big CPU
// 0x20 is a little CPU
Res.eax |= (Data.IsBig ? 0x40 : 0x20) << 24;
}
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
// Maximum supported hypervisor leafs
// We only expose the information leaf
//
// Common courtesy to follow VMWare's "Hypervisor CPUID Interface proposal"
// 4000_0000h - Information leaf. Advertising to the software which hypervisor this is
// 4000_0001h - 4000_000Fh - Hypervisor specific leafs. FEX can use these for anything
// 4000_0010h - 4000_00FFh - "Generic Leafs" - Try not to overwrite, other hypervisors might expect information in these
//
// CPUID documentation information:
// 4000_0000h - 4FFF_FFFFh - No existing or future CPU will return information in this range
// Reserved entirely for VMs to do whatever they want.
Res.eax = 0x40000001;
// EBX, EDX, ECX become the hypervisor ID signature
constexpr static char HypervisorID[12] = "FEXIFEXIEMU";
memcpy(&Res.ebx, HypervisorID, sizeof(HypervisorID));
return Res;
}
// Hypervisor CPUID information leaf
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_4000_0001h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
if (Leaf == 0) {
// EAX[3:0] Is the host architecture that FEX is running under
#ifdef _M_X86_64
// EAX[3:0] = 1 = x86_64 host architecture
Res.eax |= 0b0001;
#elif defined(_M_ARM_64)
// EAX[3:0] = 2 = AArch64 host architecture
Res.eax |= 0b0010;
#else
// EAX[3:0] = 0 = Unknown architecture
#endif
}
return Res;
}
// Highest extended function implemented
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0000h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
@@ -636,35 +940,52 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) {
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(0 << 27) | // RDTSCP
(1 << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(0 << 30) | // 3DNow! Extensions
(0 << 31); // 3DNow!
(1 << 30) | // 3DNow! Extensions
(1 << 31); // 3DNow!
return Res;
}
constexpr char ProcessorBrand[48] = {
constexpr char ProcessorBrand[32] = {
GIT_DESCRIBE_STRING
"\0"
};
constexpr ssize_t DESCRIBE_STR_SIZE = std::char_traits<char>::length(GIT_DESCRIBE_STRING);
static_assert(DESCRIBE_STR_SIZE < 32);
//Processor brand string
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[0], sizeof(FEXCore::CPUID::FunctionResults));
return Res;
return Function_8000_0002h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[16], sizeof(FEXCore::CPUID::FunctionResults));
return Res;
return Function_8000_0003h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf) {
return Function_8000_0004h(Leaf, GetCPUID());
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0002h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memcpy(&Res, &ProcessorBrand[32], sizeof(FEXCore::CPUID::FunctionResults));
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[0], std::min(16L, DESCRIBE_STR_SIZE));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0003h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
memset(&Res, ' ', sizeof(FEXCore::CPUID::FunctionResults));
memcpy(&Res, &ProcessorBrand[16], std::max(0L, DESCRIBE_STR_SIZE - 16));
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0004h(uint32_t Leaf, uint32_t CPU) {
FEXCore::CPUID::FunctionResults Res{};
auto &Data = PerCPUData[CPU];
memcpy(&Res, Data.ProductName, std::min(strlen(Data.ProductName), sizeof(FEXCore::CPUID::FunctionResults)));
return Res;
}
@@ -757,7 +1078,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0008h(uint32_t Leaf) {
Res.ebx =
(0 << 2) | // XSaveErPtr: Saving and restoring error pointers
(0 << 1) | // IRPerf: Instructions retired count support
(0 << 0); // CLZERO support
(CTX->HostFeatures.SupportsCLZERO << 0); // CLZERO support
uint32_t CoreCount = Cores() - 1;
Res.ecx =
@@ -888,25 +1209,25 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_Reserved(uint32_t Leaf) {
void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
CTX = ctx;
using namespace std::placeholders;
RegisterFunction(0, std::bind(&CPUIDEmu::Function_0h, this, _1));
RegisterFunction(1, std::bind(&CPUIDEmu::Function_01h, this, _1));
RegisterFunction(2, std::bind(&CPUIDEmu::Function_02h, this, _1));
RegisterFunction(0, &CPUIDEmu::Function_0h);
RegisterFunction(1, &CPUIDEmu::Function_01h);
RegisterFunction(2, &CPUIDEmu::Function_02h);
// 3: Serial Number(previously), now reserved
#ifndef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x4, std::bind(&CPUIDEmu::Function_04h, this, _1));
RegisterFunction(0x4, &CPUIDEmu::Function_04h);
#endif
// 5: Monitor/mwait
// Thermal and power management
RegisterFunction(6, std::bind(&CPUIDEmu::Function_06h, this, _1));
RegisterFunction(6, &CPUIDEmu::Function_06h);
// Extended feature flags
RegisterFunction(7, std::bind(&CPUIDEmu::Function_07h, this, _1));
RegisterFunction(7, &CPUIDEmu::Function_07h);
// 9: Direct Cache Access information
// 0x0A: Architectural performance monitoring
// 0x0B: Extended topology enumeration
// 0x0D: Processor extended state enumeration
RegisterFunction(0x0D, std::bind(&CPUIDEmu::Function_0Dh, this, _1));
RegisterFunction(0x0D, &CPUIDEmu::Function_0Dh);
// 0x0F: Intel RDT monitoring
// 0x10: Intel RDT allocation enumeration
// 0x12: Intel SGX capability enumeration
@@ -915,38 +1236,46 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
#ifndef CPUID_AMD
// Timestamp counter information
// Doesn't exist on AMD hardware
RegisterFunction(0x15, std::bind(&CPUIDEmu::Function_15h, this, _1));
RegisterFunction(0x15, &CPUIDEmu::Function_15h);
#endif
// 0x16: Processor frequency information
// 0x17: SoC vendor attribute enumeration
// 0x1A: Hybrid Information Sub-leaf
#ifndef CPUID_AMD
RegisterFunction(0x1A, &CPUIDEmu::Function_1Ah);
#endif
// Hypervisor CPUID information leaf
RegisterFunction(0x4000'0000, &CPUIDEmu::Function_4000_0000h);
RegisterFunction(0x4000'0001, &CPUIDEmu::Function_4000_0001h);
// Largest extended function number
RegisterFunction(0x8000'0000, std::bind(&CPUIDEmu::Function_8000_0000h, this, _1));
RegisterFunction(0x8000'0000, &CPUIDEmu::Function_8000_0000h);
// Processor vendor
RegisterFunction(0x8000'0001, std::bind(&CPUIDEmu::Function_8000_0001h, this, _1));
RegisterFunction(0x8000'0001, &CPUIDEmu::Function_8000_0001h);
// Processor brand string
RegisterFunction(0x8000'0002, std::bind(&CPUIDEmu::Function_8000_0002h, this, _1));
RegisterFunction(0x8000'0002, &CPUIDEmu::Function_8000_0002h);
// Processor brand string continued
RegisterFunction(0x8000'0003, std::bind(&CPUIDEmu::Function_8000_0003h, this, _1));
RegisterFunction(0x8000'0003, &CPUIDEmu::Function_8000_0003h);
// Processor brand string continued
RegisterFunction(0x8000'0004, std::bind(&CPUIDEmu::Function_8000_0004h, this, _1));
RegisterFunction(0x8000'0004, &CPUIDEmu::Function_8000_0004h);
// 0x8000'0005: L1 Cache and TLB identifiers
#ifdef CPUID_AMD
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_8000_0005h, this, _1));
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_8000_0005h);
#else
// This is full reserved on Intel platforms
RegisterFunction(0x8000'0005, std::bind(&CPUIDEmu::Function_Reserved, this, _1));
RegisterFunction(0x8000'0005, &CPUIDEmu::Function_Reserved);
#endif
// 0x8000'0006: L2 Cache identifiers
RegisterFunction(0x8000'0006, std::bind(&CPUIDEmu::Function_8000_0006h, this, _1));
RegisterFunction(0x8000'0006, &CPUIDEmu::Function_8000_0006h);
// Advanced power management information
RegisterFunction(0x8000'0007, std::bind(&CPUIDEmu::Function_8000_0007h, this, _1));
RegisterFunction(0x8000'0007, &CPUIDEmu::Function_8000_0007h);
// Virtual and physical address sizes
RegisterFunction(0x8000'0008, std::bind(&CPUIDEmu::Function_8000_0008h, this, _1));
RegisterFunction(0x8000'0008, &CPUIDEmu::Function_8000_0008h);
// 0x8000'000A: SVM Revision
// TLB 1GB page identifiers
RegisterFunction(0x8000'0019, std::bind(&CPUIDEmu::Function_8000_0019h, this, _1));
RegisterFunction(0x8000'0019, &CPUIDEmu::Function_8000_0019h);
// 0x8000'001A: Performance optimization identifiers
// 0x8000'001B: Instruction based sampling identifiers
@@ -954,13 +1283,13 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
// 0x8000'001D: Cache properties
#ifdef CPUID_AMD
// Deterministic cache parameters for each level
RegisterFunction(0x8000'001D, std::bind(&CPUIDEmu::Function_8000_001Dh, this, _1));
RegisterFunction(0x8000'001D, &CPUIDEmu::Function_8000_001Dh);
#endif
// 0x8000'001E: Extended APIC ID
// 0x8000'001F: AMD Secure Encryption
// Setup some state tracking
Hybrid = GetHostHybridFlag();
SetupHostHybridFlag();
}
}
+41 -8
View File
@@ -1,13 +1,13 @@
#pragma once
#include <functional>
#include <cstdint>
#include <unordered_map>
#include <utility>
#include <vector>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Config/Config.h>
#include <cstdint>
#include <utility>
namespace FEXCore {
namespace Context {
struct Context;
@@ -24,28 +24,50 @@ private:
constexpr static uint32_t CPUID_VENDOR_AMD3 = 0x444D4163; // "cAMD"
public:
// X86 cacheline size effectively has to be hardcoded to 64
// if we report anything differently then applications are likely to break
constexpr static uint64_t CACHELINE_SIZE = 64;
void Init(FEXCore::Context::Context *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) {
auto Handler = FunctionHandlers.find(Function);
const auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end()) {
return Function_Reserved(Leaf);
}
return Handler->second(Leaf);
return (this->*Handler->second)(Leaf);
}
FEXCore::CPUID::FunctionResults RunFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
if (Function == 0x8000'0002U)
return Function_8000_0002h(Leaf, CPU % PerCPUData.size());
else if (Function == 0x8000'0003U)
return Function_8000_0003h(Leaf, CPU % PerCPUData.size());
else
return Function_8000_0004h(Leaf, CPU % PerCPUData.size());
}
private:
FEXCore::Context::Context *CTX;
bool Hybrid{};
FEX_CONFIG_OPT(Cores, THREADS);
using FunctionHandler = std::function<FEXCore::CPUID::FunctionResults(uint32_t Leaf)>;
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf);
void RegisterFunction(uint32_t Function, FunctionHandler Handler) {
FunctionHandlers[Function] = Handler;
FunctionHandlers.insert_or_assign(Function, Handler);
}
std::unordered_map<uint32_t, FunctionHandler> FunctionHandlers;
struct CPUData {
const char *ProductName{};
#ifdef _M_ARM_64
uint32_t MIDR{};
#endif
bool IsBig{};
};
std::vector<CPUData> PerCPUData{};
// Functions
FEXCore::CPUID::FunctionResults Function_0h(uint32_t Leaf);
@@ -56,11 +78,19 @@ private:
FEXCore::CPUID::FunctionResults Function_07h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_0Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_15h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_1Ah(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_4000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0000h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0001h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0003h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0004h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0002h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0003h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0004h(uint32_t Leaf, uint32_t CPU);
FEXCore::CPUID::FunctionResults Function_8000_0005h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0006h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_0007h(uint32_t Leaf);
@@ -69,5 +99,8 @@ private:
FEXCore::CPUID::FunctionResults Function_8000_0019h(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_8000_001Dh(uint32_t Leaf);
FEXCore::CPUID::FunctionResults Function_Reserved(uint32_t Leaf);
void SetupHostHybridFlag();
};
}
+27 -43
View File
@@ -35,7 +35,9 @@ namespace FEXCore {
CTX->InitializeCompiler(CompileThreadData.get(), true);
CompileThreadData->CPUBackend->CopyNecessaryDataForCompileThread(ParentThread->CPUBackend.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CompileService::Initialize() {
@@ -58,23 +60,15 @@ namespace FEXCore {
// Grab the work queue and clear it
// We don't need to grab the queue mutex since this thread will no longer receive any work events
// Threads are bounded 1:1
while (WorkQueue.size()) {
WorkItem *Item = WorkQueue.front();
while (!WorkQueue.empty()) {
WorkQueue.pop();
delete Item;
}
// Go through the garbage collection array and clear it
// It's safe to clear things that aren't marked safe since we are clearing cache
if (GCArray.size()) {
// Clean up our GC array
for (auto it = GCArray.begin(); it != GCArray.end();) {
delete *it;
it = GCArray.erase(it);
}
}
GCArray.clear();
LOGMAN_THROW_A(CompileThreadData->LocalIRCache.size() == 0, "Compile service must never have LocalIRCache");
LOGMAN_THROW_A_FMT(CompileThreadData->LocalIRCache.empty(), "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
@@ -85,28 +79,26 @@ namespace FEXCore {
SelectedThread->CPUBackend->ClearCache();
}
CompileService::WorkItem *CompileService::CompileCode(uint64_t RIP) {
// Tell the worker thread to compile code for us
WorkItem *Item = new WorkItem{};
Item->RIP = RIP;
WorkItem* ResultItem = nullptr;
{
// Tell the worker thread to compile code for us
auto Item = std::make_unique<WorkItem>();
Item->RIP = RIP;
// Fill the threads work queue
std::scoped_lock<std::mutex> lk(QueueMutex);
WorkQueue.emplace(Item);
std::scoped_lock lk(QueueMutex);
ResultItem = WorkQueue.emplace(std::move(Item)).get();
}
// Notify the thread that it has more work
StartWork.NotifyAll();
return Item;
return ResultItem;
}
void CompileService::ExecutionThread() {
// Ignore signals coming from the guest
CTX->SignalDelegation->MaskThreadSignals();
// Set our thread name so we can see its relation
char ThreadName[16]{};
snprintf(ThreadName, 16, "%ld-CS", ParentThread->ThreadManager.TID.load());
@@ -118,18 +110,18 @@ namespace FEXCore {
if (ShuttingDown.load()) {
break;
}
std::scoped_lock<std::mutex> lk(CompileMutex);
std::scoped_lock lk(CompileMutex);
size_t WorkItems{};
do {
// Grab a work item
WorkItem *Item{};
std::unique_ptr<WorkItem> Item{};
{
std::scoped_lock<std::mutex> lk(QueueMutex);
std::scoped_lock lk(QueueMutex);
WorkItems = WorkQueue.size();
if (WorkItems) {
Item = WorkQueue.front();
if (WorkItems != 0) {
Item = std::move(WorkQueue.front());
WorkQueue.pop();
}
}
@@ -137,7 +129,7 @@ namespace FEXCore {
// If we had a work item then work on it
if (Item) {
// Make sure it's not in lookup cache by accident
LOGMAN_THROW_A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
LOGMAN_THROW_A_FMT(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
@@ -145,11 +137,11 @@ namespace FEXCore {
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
LOGMAN_THROW_A(Generated == true, "Compile Service doesn't have IR Cache");
LOGMAN_THROW_A_FMT(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
ERROR_AND_DIE("Couldn't compile code for thread at RIP: 0x%lx", Item->RIP);
ERROR_AND_DIE_FMT("Couldn't compile code for thread at RIP: 0x{:x}", Item->RIP);
}
Item->CodePtr = CodePtr;
@@ -159,23 +151,15 @@ namespace FEXCore {
Item->StartAddr = StartAddr;
Item->Length = Length;
GCArray.emplace_back(Item);
Item->ServiceWorkDone.NotifyAll();
auto& GCItem = GCArray.emplace_back(std::move(Item));
GCItem->ServiceWorkDone.NotifyAll();
}
} while (WorkItems != 0);
if (GCArray.size()) {
// Clean up our GC array
for (auto it = GCArray.begin(); it != GCArray.end();) {
if ((*it)->SafeToClear) {
delete *it;
it = GCArray.erase(it);
}
else {
++it;
}
}
}
// Clean up any safe entries in our GC array if we have any.
std::erase_if(GCArray, [](const auto& Entry) {
return Entry->SafeToClear.load(std::memory_order_relaxed);
});
}
}
}
+5 -2
View File
@@ -48,6 +48,9 @@ class CompileService final {
// Public for threading
void ExecutionThread();
bool IsAddressInJITCode(uint64_t Address) const {
return CompileThreadData->CPUBackend->IsAddressInJITCode(Address, false, false);
}
private:
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ParentThread;
@@ -57,8 +60,8 @@ class CompileService final {
std::mutex QueueMutex{};
std::mutex CompileMutex{};
std::queue<WorkItem*> WorkQueue{};
std::vector<WorkItem*> GCArray{};
std::queue<std::unique_ptr<WorkItem>> WorkQueue{};
std::vector<std::unique_ptr<WorkItem>> GCArray{};
Event StartWork{};
std::atomic_bool ShuttingDown{false};
};
+180 -461
View File
@@ -41,6 +41,7 @@ $end_info$
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <algorithm>
#include <array>
@@ -50,6 +51,7 @@ $end_info$
#include <cstdint>
#include <filesystem>
#include <functional>
#include <fstream>
#include <map>
#include <memory>
#include <mutex>
@@ -61,7 +63,6 @@ $end_info$
#include <string.h>
#include <string>
#include <string_view>
#include <sstream>
#include <sys/mman.h>
#include <sys/stat.h>
#include <sys/syscall.h>
@@ -141,63 +142,11 @@ std::string_view const& GetGRegName(unsigned Reg) {
} // namespace FEXCore::Core
namespace FEXCore::Context {
void Context::AOTIRCaptureCacheWriteoutQueue_Flush() {
{
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
for (;;) {
AOTIRCaptureCacheWriteoutLock.lock();
std::function<void()> fn = std::move(AOTIRCaptureCacheWriteoutQueue.front());
bool MaybeEmpty = false;
AOTIRCaptureCacheWriteoutQueue.pop();
MaybeEmpty = AOTIRCaptureCacheWriteoutQueue.size() == 0;
AOTIRCaptureCacheWriteoutLock.unlock();
fn();
if (MaybeEmpty) {
std::shared_lock lk{AOTIRCaptureCacheWriteoutLock};
if (AOTIRCaptureCacheWriteoutQueue.size() == 0) {
AOTIRCaptureCacheWriteoutFlusing.store(false);
return;
}
}
}
LOGMAN_MSG_A("Must never get here");
}
void Context::AOTIRCaptureCacheWriteoutQueue_Append(const std::function<void()> &fn) {
bool Flush = false;
{
std::unique_lock lk{AOTIRCaptureCacheWriteoutLock};
AOTIRCaptureCacheWriteoutQueue.push(fn);
if (AOTIRCaptureCacheWriteoutQueue.size() > 10000) {
Flush = true;
}
}
bool test_val = false;
if (Flush && AOTIRCaptureCacheWriteoutFlusing.compare_exchange_strong(test_val, true)) {
AOTIRCaptureCacheWriteoutQueue_Flush();
}
}
Context::Context() {
Context::Context()
: IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
if (Config.GdbServer) {
StartGdbServer();
}
else {
StopGdbServer();
}
}
Context::~Context() {
@@ -217,10 +166,6 @@ namespace FEXCore::Context {
}
Threads.clear();
}
for (auto &Mod: AOTIRCache) {
FEXCore::Allocator::munmap(Mod.second.mapping, Mod.second.size);
}
}
static FEXCore::Core::CPUState CreateDefaultCPUState() {
@@ -245,12 +190,44 @@ namespace FEXCore::Context {
}
FEXCore::Core::InternalThreadState* Context::InitCore(FEXCore::CodeLoader *Loader) {
// Initialize the CPU core signal handlers
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
FEXCore::CPU::InitializeInterpreterSignalHandlers(this);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
#elif (_M_ARM_64 && JIT_ARM64)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
// Do nothing
break;
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
}
// Initialize GDBServer after the signal handlers are installed
// It may install its own handlers that need to be executed AFTER the CPU cores
if (Config.GdbServer) {
StartGdbServer();
}
else {
StopGdbServer();
}
ThunkHandler.reset(FEXCore::ThunkHandler::Create());
LocalLoader = Loader;
using namespace FEXCore::Core;
FEXCore::CPU::InitializeInterpreterOpHandlers();
FEXCore::Core::CPUState NewThreadState = CreateDefaultCPUState();
FEXCore::Core::InternalThreadState *Thread = CreateThread(&NewThreadState, 0);
@@ -323,7 +300,7 @@ namespace FEXCore::Context {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Pause);
if (Thread->RunningEvents.Running.load()) {
// Only attempt to stop this thread if it is running
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
}
@@ -342,7 +319,6 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Return);
Thread->RunningEvents.WaitingToStart.store(true);
}
for (auto &Thread : Threads) {
@@ -381,13 +357,13 @@ namespace FEXCore::Context {
this->Config.MaxInstPerBlock = 1;
Run();
WaitForThreadsToRun();
WaitForIdleWithTimeout();
WaitForIdle();
this->Config.RunningMode = PreviousRunningMode;
this->Config.MaxInstPerBlock = PreviousMaxIntPerBlock;
}
void Context::Stop(bool IgnoreCurrentThread) {
pid_t tid = gettid();
pid_t tid = FHU::Syscalls::gettid();
FEXCore::Core::InternalThreadState* CurrentThread{};
// Tell all the threads that they should stop
@@ -408,6 +384,13 @@ namespace FEXCore::Context {
if (Thread->RunningEvents.Running.load()) {
StopThread(Thread);
}
// If the thread is waiting to start but immediately killed then there can be a hang
// This occurs in the case of gdb attach with immediate kill
if (Thread->RunningEvents.WaitingToStart.load()) {
Thread->RunningEvents.EarlyExit = true;
Thread->StartRunning.NotifyAll();
}
}
}
@@ -420,20 +403,25 @@ namespace FEXCore::Context {
void Context::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->RunningEvents.Running.exchange(false)) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Stop);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
void Context::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->RunningEvents.Running.load()) {
Thread->SignalReason.store(Event);
tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
FHU::Syscalls::tgkill(Thread->ThreadManager.PID, Thread->ThreadManager.TID, SignalDelegator::SIGNAL_FOR_PAUSE);
}
}
FEXCore::Context::ExitReason Context::RunUntilExit() {
if(!StartPaused)
Run();
if(!StartPaused) {
// We will only have one thread at this point, but just in case run notify everything
std::lock_guard lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->StartRunning.NotifyAll();
}
}
ExecutionThread(ParentThread);
while(true) {
@@ -468,7 +456,6 @@ namespace FEXCore::Context {
};
LocalLoader->AddIR(IRHandler);
}
struct ExecutionThreadHandler {
@@ -496,7 +483,7 @@ namespace FEXCore::Context {
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = ::gettid();
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
SignalDelegation->RegisterTLSState(Thread);
ThunkHandler->RegisterTLSState(Thread);
@@ -532,9 +519,11 @@ namespace FEXCore::Context {
// Create CPU backend
switch (Config.Core) {
#ifdef INTERPRETER_ENABLED
case FEXCore::Config::CONFIG_INTERPRETER:
State->CPUBackend = FEXCore::CPU::CreateInterpreterCore(this, State, CompileThread);
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
State->PassManager->InsertRegisterAllocationPass(DoSRA);
@@ -543,14 +532,16 @@ namespace FEXCore::Context {
#elif (_M_ARM_64 && JIT_ARM64)
State->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, State, CompileThread);
#else
ERROR_AND_DIE("FEXCore has been compiled without a viable JIT core");
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
#endif
break;
case FEXCore::Config::CONFIG_CUSTOM:
State->CPUBackend = CustomCPUFactory(this, State);
break;
default: ERROR_AND_DIE("Unknown core configuration");
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
}
}
@@ -583,7 +574,7 @@ namespace FEXCore::Context {
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
auto It = std::find(Threads.begin(), Threads.end(), Thread);
LOGMAN_THROW_A(It != Threads.end(), "Thread wasn't in Threads");
LOGMAN_THROW_A_FMT(It != Threads.end(), "Thread wasn't in Threads");
Threads.erase(It);
}
@@ -657,6 +648,58 @@ namespace FEXCore::Context {
}
}
static void IRDumper(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
const auto DumpIRStr = Thread->CTX->Config.DumpIR();
if (DumpIRStr =="stderr") {
f = stderr;
}
else if (DumpIRStr =="stdout") {
f = stdout;
}
else {
const auto fileName = fmt::format("{}/{:x}{}", DumpIRStr, GuestRIP, RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.c_str(), "w");
CloseAfter = true;
}
if (f) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fmt::print(f,"IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
if (CloseAfter) {
fclose(f);
}
}
};
static void ValidateIR(FEXCore::Core::InternalThreadState *Thread) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction();
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
auto reparsed = IR::Parse(&out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
std::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::IFmt("one:\n {}", out.str());
LogMan::Msg::IFmt("two:\n {}", out2.str());
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
}
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
uint8_t const *GuestCode{};
GuestCode = reinterpret_cast<uint8_t const*>(GuestRIP);
@@ -666,9 +709,7 @@ namespace FEXCore::Context {
uint64_t TotalInstructions {0};
uint64_t TotalInstructionsLength {0};
if (!Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP)) {
return {};
}
Thread->FrontendDecoder->DecodeInstructionsAtEntry(GuestCode, GuestRIP);
auto CodeBlocks = Thread->FrontendDecoder->GetDecodedBlocks();
@@ -681,7 +722,6 @@ namespace FEXCore::Context {
// Set the block entry point
Thread->OpDispatcher->SetNewBlockIfChanged(Block.Entry);
uint64_t BlockInstructionsLength {};
// Reset any block-specific state
@@ -689,11 +729,6 @@ namespace FEXCore::Context {
uint64_t InstsInBlock = Block.NumInstructions;
if (Block.HasInvalidInstruction) {
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
break;
}
for (size_t i = 0; i < InstsInBlock; ++i) {
FEXCore::X86Tables::X86InstInfo const* TableInfo {nullptr};
FEXCore::X86Tables::DecodedInst const* DecodedInfo {nullptr};
@@ -723,7 +758,7 @@ namespace FEXCore::Context {
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
}
if (TableInfo->OpcodeDispatcher) {
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
Thread->OpDispatcher->HandledLock = false;
Thread->OpDispatcher->ResetDecodeFailure();
@@ -734,7 +769,7 @@ namespace FEXCore::Context {
else {
if (Thread->OpDispatcher->HandledLock != IsLocked) {
HadDispatchError = true;
LogMan::Msg::E("Missing LOCK HANDLER at 0x%lx{'%s'}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
LogMan::Msg::EFmt("Missing LOCK HANDLER at 0x{:x}{{'{}'}}", Block.Entry + BlockInstructionsLength, TableInfo->Name ?: "UND");
}
BlockInstructionsLength += DecodedInfo->InstSize;
TotalInstructionsLength += DecodedInfo->InstSize;
@@ -742,8 +777,9 @@ namespace FEXCore::Context {
}
}
else {
LogMan::Msg::E("Missing OpDispatcher at 0x%lx{'%s'}", Block.Entry + BlockInstructionsLength, TableInfo->Name);
HadDispatchError = true;
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize));
}
// If we had a dispatch error then leave early
@@ -770,74 +806,32 @@ namespace FEXCore::Context {
Thread->OpDispatcher->Finalize();
auto IRDumper = [Thread, GuestRIP](IR::RegisterAllocationData* RA) {
FILE* f = nullptr;
bool CloseAfter = false;
if (Thread->CTX->Config.DumpIR() =="stderr") {
f = stderr;
}
else if (Thread->CTX->Config.DumpIR() =="stdout") {
f = stdout;
}
else {
std::stringstream fileName;
fileName << Thread->CTX->Config.DumpIR() << "/" << std::hex << GuestRIP << (RA ? "-post.ir" : "-pre.ir");
f = fopen(fileName.str().c_str(), "w");
CloseAfter = true;
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, nullptr);
}
if (f) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, RA);
fprintf(f,"IR-%s 0x%lx:\n%s\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str().c_str());
if (CloseAfter) {
fclose(f);
}
}
};
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(nullptr);
}
if (Thread->CTX->Config.ValidateIRarser) {
// Convert to text, Parse, Convert to text again and make sure the texts match
std::stringstream out;
static auto compaction = IR::CreateIRCompaction();
compaction->Run(Thread->OpDispatcher.get());
auto NewIR = Thread->OpDispatcher->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
auto reparsed = IR::Parse(&out);
if (reparsed == nullptr) {
LOGMAN_MSG_A("Failed to parse ir\n");
} else {
std::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::I("one:\n %s", out.str().c_str());
LogMan::Msg::I("two:\n %s", out2.str().c_str());
LOGMAN_MSG_A("Parsed ir doesn't match\n");
}
if (Thread->CTX->Config.ValidateIRarser) {
ValidateIR(Thread);
}
}
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(Thread->OpDispatcher.get());
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
// Debug
{
if (Thread->CTX->Config.DumpIR() != "no") {
IRDumper(Thread, GuestRIP, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::I("IR 0x%lx:\n%s\n@@@@@\n", GuestRIP, out.str().c_str());
if (Thread->OpDispatcher->ShouldDump) {
std::stringstream out;
auto NewIR = Thread->OpDispatcher->ViewIR();
FEXCore::IR::Dump(&out, &NewIR, Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
LogMan::Msg::IFmt("IR 0x{:x}:\n{}\n@@@@@\n", GuestRIP, out.str());
}
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
@@ -855,63 +849,6 @@ namespace FEXCore::Context {
};
}
AOTIRInlineEntry *AOTIRInlineIndex::GetInlineEntry(uint64_t DataOffset) {
uintptr_t This = (uintptr_t)this;
return (AOTIRInlineEntry*)(This + DataBase + DataOffset);
}
AOTIRInlineEntry *AOTIRInlineIndex::Find(uint64_t GuestStart) {
ssize_t l = 0;
ssize_t r = Count - 1;
while (l <= r) {
size_t m = l + (r - l) / 2;
if (Entries[m].GuestStart == GuestStart)
return GetInlineEntry(Entries[m].DataOffset);
else if (Entries[m].GuestStart < GuestStart)
l = m + 1;
else
r = m - 1;
}
return nullptr;
}
IR::RegisterAllocationData *AOTIRInlineEntry::GetRAData() {
return (IR::RegisterAllocationData *)InlineData;
}
IR::IRListView *AOTIRInlineEntry::GetIRData() {
auto RAData = GetRAData();
auto Offset = RAData->Size(RAData->MapCount);
return (IR::IRListView *)&InlineData[Offset];
}
void AOTIRCaptureCacheEntry::AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView *IRList, FEXCore::IR::RegisterAllocationData *RAData) {
auto Inserted = Index.emplace(GuestRIP, Stream->tellp());
if (Inserted.second) {
//GuestHash
Stream->write((const char*)&Hash, sizeof(Hash));
//GuestLength
Stream->write((const char*)&Length, sizeof(Length));
// RAData (inline)
// In file, IsShared is always set
auto Shared = RAData->IsShared;
RAData->IsShared = true;
Stream->write((const char*)RAData, RAData->Size(RAData->MapCount));
RAData->IsShared = Shared;
// IRData (inline)
IRList->Serialize(*Stream);
}
}
Context::CompileCodeResult Context::CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
FEXCore::IR::IRListView *IRList {};
FEXCore::Core::DebugData *DebugData {};
@@ -935,55 +872,17 @@ namespace FEXCore::Context {
GeneratedIR = false;
}
// AOT IR bookkeeping and cache
{
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
if (!file->second.ContainsCode) {
file->second.ContainsCode = true;
FilesWithCode[file->second.fileid] = file->second.filename;
}
}
}
if (IRList == nullptr && Config.AOTIRLoad) {
std::shared_lock lk(AOTIRCacheLock);
auto file = AddrToFile.lower_bound(GuestRIP);
if (file != AddrToFile.begin()) {
--file;
auto Mod = (AOTIRInlineIndex*)file->second.CachedFileEntry;
if (Mod == nullptr) {
file->second.CachedFileEntry = Mod = AOTIRCache[file->second.fileid].Array;
}
if (Mod != nullptr)
{
auto AOTEntry = Mod->Find(GuestRIP - file->second.Start + file->second.Offset);
if (AOTEntry) {
// verify hash
auto MappedStart = GuestRIP;
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
if (hash == AOTEntry->GuestHash) {
IRList = AOTEntry->GetIRData();
//LogMan::Msg::D("using %s + %lx -> %lx\n", file->second.fileid.c_str(), AOTEntry->first, GuestRIP);
RAData = AOTEntry->GetRAData();;
DebugData = new FEXCore::Core::DebugData();
StartAddr = MappedStart;
Length = AOTEntry->GuestLength;
GeneratedIR = true;
} else {
LogMan::Msg::I("AOTIR: hash check failed %lx\n", MappedStart);
}
} else {
//LogMan::Msg::I("AOTIR: Failed to find %lx, %lx, %s\n", GuestRIP, GuestRIP - file->second.Start + file->second.Offset, file->second.fileid.c_str());
}
}
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(GuestRIP, IRList);
if (_GeneratedIR) {
// Setup pointers to internal structures
IRList = IRCopy;
RAData = RACopy;
DebugData = DebugDataCopy;
StartAddr = _StartAddr;
Length = _Length;
GeneratedIR = _GeneratedIR;
}
}
@@ -998,10 +897,6 @@ namespace FEXCore::Context {
StartAddr = _StartAddr;
Length = _Length;
// Initialize metadata
DebugData->GuestCodeSize = TotalInstructionsLength;
DebugData->GuestInstructionCount = TotalInstructions;
// Increment stats
Thread->Stats.BlocksCompiled.fetch_add(1);
@@ -1024,138 +919,17 @@ namespace FEXCore::Context {
};
}
static bool readAll(int fd, void *data, size_t size) {
int rv = read(fd, data, size);
if (rv != size)
return false;
else
return true;
}
bool Context::LoadAOTIRCache(int streamfd) {
uint64_t tag;
if (!readAll(streamfd, (char*)&tag, sizeof(tag)) || tag != 0xDEADBEEFC0D30004)
return false;
std::string Module;
uint64_t ModSize;
uint64_t IndexSize;
lseek(streamfd, -sizeof(ModSize), SEEK_END);
if (!readAll(streamfd, (char*)&ModSize, sizeof(ModSize)))
return false;
Module.resize(ModSize);
lseek(streamfd, -sizeof(ModSize) - ModSize, SEEK_END);
if (!readAll(streamfd, (char*)&Module[0], Module.size()))
return false;
lseek(streamfd, -sizeof(ModSize) - ModSize - sizeof(IndexSize), SEEK_END);
if (!readAll(streamfd, (char*)&IndexSize, sizeof(IndexSize)))
return false;
struct stat fileinfo;
if (fstat(streamfd, &fileinfo) < 0)
return false;
size_t Size = (fileinfo.st_size + 4095) & ~4095;
size_t IndexOffset = fileinfo.st_size - IndexSize -sizeof(ModSize) - ModSize - sizeof(IndexSize);
void *FilePtr = FEXCore::Allocator::mmap(nullptr, Size, PROT_READ, MAP_SHARED, streamfd, 0);
if (FilePtr == MAP_FAILED)
return false;
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
AOTIRCache.insert({Module, {Array, FilePtr, Size}});
LogMan::Msg::D("AOTIR: Module %s has %ld functions", Module.c_str(), Array->Count);
return true;
}
void Context::WriteFilesWithCode(std::function<void(const std::string& fileid, const std::string& filename)> Writer) {
std::shared_lock lk(AOTIRCacheLock);
for( const auto &File: FilesWithCode) {
Writer(File.first, File.second);
}
}
void Context::FinalizeAOTIRCache() {
AOTIRCaptureCacheWriteoutQueue_Flush();
std::unique_lock lk(AOTIRCacheLock);
for (auto& [String, Entry] : AOTIRCaptureCache) {
if (!Entry.Stream) {
continue;
}
const auto ModSize = String.size();
auto &stream = Entry.Stream;
// pad to 32 bytes
constexpr char Zero = 0;
while(stream->tellp() & 31)
stream->write(&Zero, 1);
// AOTIRInlineIndex
const auto FnCount = Entry.Index.size();
const size_t DataBase = -stream->tellp();
stream->write((const char*)&FnCount, sizeof(FnCount));
stream->write((const char*)&DataBase, sizeof(DataBase));
for (const auto& [GuestStart, DataOffset] : Entry.Index) {
//AOTIRInlineIndexEntry
// GuestStart
stream->write((const char*)&GuestStart, sizeof(GuestStart));
// DataOffset
stream->write((const char*)&DataOffset, sizeof(DataOffset));
}
// End of file header
const auto IndexSize = FnCount * sizeof(AOTIRInlineIndexEntry) + sizeof(DataBase) + sizeof(FnCount);
stream->write((const char*)&IndexSize, sizeof(IndexSize));
stream->write(String.c_str(), ModSize);
stream->write((const char*)&ModSize, sizeof(ModSize));
}
}
void Context::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto NewBlock = CompileBlock(Frame, GuestRIP);
if (NewBlock == 0) {
LogMan::Msg::E("CompileBlockJit: Failed to compile code %lX - aborting process", GuestRIP);
LogMan::Msg::EFmt("CompileBlockJit: Failed to compile code {:X} - aborting process", GuestRIP);
// Return similar behaviour of SIGILL abort
Frame->Thread->StatusCode = 128 + SIGILL;
Stop(false /* Ignore current thread */);
}
}
Context::AddrToFileMapType::iterator Context::FindAddrForFile(uint64_t Entry, uint64_t Length) {
// Thread safety here! We are returning an iterator to the map object
// This needs the AOTIRCacheLock locked prior to coming in to the function
auto file = AddrToFile.lower_bound(Entry);
if (file != AddrToFile.begin()) {
--file;
if (file->second.Start <= Entry && (file->second.Start + file->second.Len) >= (Entry + Length)) {
return file;
}
}
return AddrToFile.end();
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
@@ -1180,7 +954,7 @@ namespace FEXCore::Context {
Thread->CompileService->Initialize();
}
auto WorkItem = Thread->CompileService->CompileCode(GuestRIP);
auto* WorkItem = Thread->CompileService->CompileCode(GuestRIP);
WorkItem->ServiceWorkDone.Wait();
// Return here with the data in place
CodePtr = WorkItem->CodePtr;
@@ -1227,56 +1001,19 @@ namespace FEXCore::Context {
}
}
// Both generated ir and LibraryJITName need a named region lookup
if (GeneratedIR || Config.LibraryJITNaming()) {
std::shared_lock lk(AOTIRCacheLock);
auto file = FindAddrForFile(StartAddr, Length);
// Only go down this path if we actually found a library region
if (file != AddrToFile.end()) {
if (DebugData && Config.LibraryJITNaming()) {
Symbols.RegisterNamedRegion(CodePtr, DebugData->HostCodeSize, file->second.filename);
}
// Add to AOT cache if aot generation is enabled
if (GeneratedIR && RAData &&
(Config.AOTIRCapture() || Config.AOTIRGenerate())) {
auto hash = XXH3_64bits((void*)StartAddr, Length);
auto LocalRIP = GuestRIP - file->second.Start + file->second.Offset;
auto LocalStartAddr = StartAddr - file->second.Start + file->second.Offset;
auto fileid = file->second.fileid;
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRList, RAData, fileid]() {
auto *AotFile = &AOTIRCaptureCache[fileid];
if (!AotFile->Stream) {
AotFile->Stream = AOTIRWriter(fileid);
uint64_t tag = 0xDEADBEEFC0D30004;
AotFile->Stream->write((char*)&tag, sizeof(tag));
}
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRList, RAData);
});
if (Config.AOTIRGenerate()) {
// cleanup memory and early exit here -- we're not running the application
if (DecrementRefCount)
--Thread->CompileBlockReentrantRefCount;
Thread->CPUBackend->ClearCache();
return (uintptr_t)CodePtr;
}
}
}
// Insert to caches if we generated IR
if (GeneratedIR) {
// Add to thread local ir cache
Core::LocalIREntry Entry = {StartAddr, Length, decltype(Entry.IR)(IRList), decltype(Entry.RAData)(RAData), decltype(Entry.DebugData)(DebugData)};
Thread->LocalIRCache.insert({GuestRIP, std::move(Entry)});
}
if (IRCaptureCache.PostCompileCode(
Thread,
CodePtr,
GuestRIP,
StartAddr,
Length,
RAData,
IRList,
DebugData,
GeneratedIR,
DecrementRefCount)) {
// Early exit
return (uintptr_t)CodePtr;
}
if (DecrementRefCount)
@@ -1304,14 +1041,17 @@ namespace FEXCore::Context {
Thread->StartRunning.Wait();
}
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
if (!Thread->RunningEvents.EarlyExit.load()) {
Thread->RunningEvents.WaitingToStart = false;
Thread->RunningEvents.Running = true;
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_NONE;
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = true;
Thread->RunningEvents.WaitingToStart = false;
Thread->RunningEvents.Running = false;
Thread->CPUBackend->ExecuteDispatch(Thread->CurrentFrame);
Thread->RunningEvents.Running = false;
}
// If it is the parent thread that died then just leave
// XXX: This doesn't make sense when the parent thread doesn't outlive its children
@@ -1403,38 +1143,17 @@ namespace FEXCore::Context {
}
void Context::AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename) {
// TODO: Support overlapping maps and region splitting
auto base_filename = std::filesystem::path(filename).filename().string();
if (!base_filename.empty()) {
auto filename_hash = XXH3_64bits(filename.c_str(), filename.size());
auto fileid = base_filename + "-" + std::to_string(filename_hash) + "-";
// append optimization flags to the fileid
fileid += (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) ? "S" : "s";
fileid += Config.TSOEnabled ? "T" : "t";
fileid += Config.ABILocalFlags ? "L" : "l";
fileid += Config.ABINoPF ? "p" : "P";
std::unique_lock lk(AOTIRCacheLock);
AddrToFile.insert({ Base, { Base, Size, Offset, fileid, filename, nullptr, false} });
if (Config.AOTIRLoad && !AOTIRCache.contains(fileid) && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
if (streamfd != -1) {
LoadAOTIRCache(streamfd);
close(streamfd);
}
}
IRCaptureCache.AddNamedRegion(Base, Size, Offset, filename);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void Context::RemoveNamedRegion(uintptr_t Base, uintptr_t Size) {
std::unique_lock lk(AOTIRCacheLock);
// TODO: Support partial removing
AddrToFile.erase(Base);
IRCaptureCache.RemoveNamedRegion(Base, Size);
if (DebugServer) {
DebugServer->AlertLibrariesChanged();
}
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
@@ -11,6 +11,8 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <array>
#include <bit>
#include <cmath>
@@ -25,6 +27,7 @@
#include "code-buffer-vixl.h"
#include "platform-vixl.h"
#include <sys/syscall.h>
#include <unistd.h>
namespace FEXCore::CPU {
@@ -36,7 +39,7 @@ static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config)
: Dispatcher(ctx, Thread), Arm64Emitter(MAX_DISPATCHER_CODE_SIZE) {
: Dispatcher(ctx, Thread), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE) {
SRAEnabled = config.StaticRegisterAssignment;
SetAllowAssembler(true);
@@ -50,10 +53,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// Ptr();
// }
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
Literal l_VirtualMemory {VirtualMemorySize};
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_L1Ptr {Thread->LookupCache->GetL1Pointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
Literal l_CompileBlock {GetCompileBlockPtr()};
@@ -96,7 +96,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
auto RipReg = x2;
// L1 Cache
ldr(x0, &l_L1Ptr);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -118,11 +118,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ldr(x0, &l_PagePtr);
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
if (std::popcount(VirtualMemorySize) == 1) {
and_(x3, RipReg, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, VirtualMemorySize - 1);
}
else {
ldr(x3, &l_VirtualMemory);
LoadConstant(x3, VirtualMemorySize);
and_(x3, RipReg, x3);
}
@@ -156,7 +157,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(x0, &l_L1Ptr);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
@@ -205,11 +206,31 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ret();
}
constexpr bool SignalSafeCompile = true;
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x0, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
}
ldr(x0, &l_ExitFunctionLinkThis);
mov(x1, STATE);
mov(x2, lr);
@@ -217,6 +238,24 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
ldr(x3, &l_ExitFunctionLink);
blr(x3);
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
mov(x4, x0);
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
mov(x0, x4);
}
if (SRAEnabled)
FillStaticRegs();
br(x0);
@@ -226,16 +265,53 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
{
bind(&NoBlock);
if (SRAEnabled)
SpillStaticRegs();
if (SignalSafeCompile) {
// When compiling code, mask all signals to reduce the chance of reentrant allocations
// Args:
// X0: SETMASK
// X1: Pointer to mask value (uint64_t)
// X2: Pointer to old mask value (uint64_t)
// X3: Size of mask, sizeof(uint64_t)
// X8: Syscall
LoadConstant(x0, ~0ULL);
stp(x0, x2, MemOperand(sp, -16, PreIndex));
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
add(x2, sp, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Reload x2 to bring back RIP
ldr(x2, MemOperand(sp, 8, Offset));
}
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileBlock);
if (SRAEnabled)
SpillStaticRegs();
// X2 contains our guest RIP
blr(x3); // { CTX, Frame, RIP}
if (SignalSafeCompile) {
// Now restore the signal mask
// Living in the same location
LoadConstant(x0, SIG_SETMASK);
add(x1, sp, 0);
LoadConstant(x2, 0);
LoadConstant(x3, 8);
LoadConstant(x8, SYS_rt_sigprocmask);
svc(0);
// Bring stack back
add(sp, sp, 16);
}
if (SRAEnabled)
FillStaticRegs();
@@ -250,6 +326,42 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
hlt(0);
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
hlt(0);
}
{
// Guest Overflow handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = GetCursorAddress<uint64_t>();
if (SRAEnabled)
SpillStaticRegs();
LoadConstant(x0, reinterpret_cast<uint64_t>(&SynchronousFaultData));
LoadConstant(w1, 1);
strb(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)));
LoadConstant(w1, X86State::X86_TRAPNO_OF);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)));
LoadConstant(w1, 0x80);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)));
LoadConstant(x1, 0);
str(w1, MemOperand(x0, offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)));
// hlt/udf = SIGILL
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
ldr(x1, MemOperand(x1));
}
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (SRAEnabled)
@@ -324,9 +436,89 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
b(&LoopTop);
}
place(&l_VirtualMemory);
// Long division helpers
uint64_t LUDIVHandler{};
uint64_t LDIVHandler{};
uint64_t LUREMHandler{};
uint64_t LREMHandler{};
{
LUDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LDIVHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIV)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LUREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LREMHandler = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREM)));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
place(&l_PagePtr);
place(&l_L1Ptr);
place(&l_CTX);
place(&l_Sleep);
place(&l_CompileBlock);
@@ -341,16 +533,38 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::
GetBuffer()->SetExecutable();
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr), Name);
}
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(DispatchPtr), End - reinterpret_cast<uint64_t>(DispatchPtr));
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Pointers.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
Pointers.LUDIVHandler = LUDIVHandler;
Pointers.LDIVHandler = LDIVHandler;
Pointers.LUREMHandler = LUREMHandler;
Pointers.LREMHandler = LREMHandler;
}
}
void Arm64Dispatcher::SpillSRA(void *ucontext) {
void Arm64Dispatcher::SpillSRA(void *ucontext, uint32_t IgnoreMask) {
for(int i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
// Skip this one, it's already spilled
continue;
}
ThreadState->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
@@ -18,7 +18,7 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
Arm64Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, DispatcherConfig &config);
protected:
void SpillSRA(void *ucontext) override;
void SpillSRA(void *ucontext, uint32_t IgnoreMask) override;
};
}
@@ -1,8 +1,8 @@
#include "Common/MathUtils.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/CompileService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
@@ -12,12 +12,13 @@
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <atomic>
#include <condition_variable>
#include <bits/types/siginfo_t.h>
#include <csignal>
#include <cstring>
#include <signal.h>
#include <string.h>
namespace FEXCore::CPU {
@@ -27,15 +28,19 @@ void Dispatcher::SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuS
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
Thread->RunningEvents.ThreadSleeping = true;
// Go to sleep
Thread->StartRunning.Wait();
Thread->RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
Thread->RunningEvents.ThreadSleeping = false;
ctx->IdleWaitCV.notify_all();
}
void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
ArchHelpers::Context::ContextBackup* Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// We can end up getting a signal at any point in our host state
// Jump to a handler that saves all state so we can safely return
uint64_t OldSP = ArchHelpers::Context::GetSp(ucontext);
@@ -65,13 +70,35 @@ void Dispatcher::StoreThreadState(int Signal, void *ucontext) {
// Set the new SP
ArchHelpers::Context::SetSp(ucontext, NewSP);
SignalFrames.push(NewSP);
// Signal frames are only used on the interpreter
// The JITS require the stack to be setup correctly on rt_sigreturn
if (CTX->Config.Core() == FEXCore::Config::CONFIG_INTERPRETER) {
SignalFrames.push(NewSP);
}
Context->Flags = 0;
Context->FPStateLocation = 0;
Context->UContextLocation = 0;
Context->SigInfoLocation = 0;
// Store fault to top status and then reset it
Context->FaultToTopAndGeneratedException = SynchronousFaultData.FaultToTopAndGeneratedException;
SynchronousFaultData.FaultToTopAndGeneratedException = false;
return Context;
}
void Dispatcher::RestoreThreadState(void *ucontext) {
LOGMAN_THROW_A(!SignalFrames.empty(), "Trying to restore a signal frame when we don't have any");
uint64_t OldSP = SignalFrames.top();
SignalFrames.pop();
uint64_t OldSP{};
if (CTX->Config.Core() == FEXCore::Config::CONFIG_IRJIT) {
OldSP = ArchHelpers::Context::GetSp(ucontext);
}
else {
LOGMAN_THROW_A_FMT(!SignalFrames.empty(), "Trying to restore a signal frame when we don't have any");
OldSP = SignalFrames.top();
SignalFrames.pop();
}
uintptr_t NewSP = OldSP;
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
@@ -80,6 +107,139 @@ void Dispatcher::RestoreThreadState(void *ucontext) {
// Now restore host state
ArchHelpers::Context::RestoreContext(ucontext, Context);
if (Context->UContextLocation) {
auto Frame = ThreadState->CurrentFrame;
if (Context->Flags &ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT) {
// XXX: Unsupported since it needs state reconstruction
// If we are in the JIT then SRA might need to be restored to values from the context
// We can't currently support this since it might result in tearing without real state reconstruction
}
if (!(Context->Flags & ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_32BIT)) {
auto *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(Context->UContextLocation);
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] ||
Context->FaultToTopAndGeneratedException) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP];
// XXX: Full context setting
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
Frame->State.flags[1] = 1;
Frame->State.flags[9] = 1;
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_##x];
COPY_REG(R8);
COPY_REG(R9);
COPY_REG(R10);
COPY_REG(R11);
COPY_REG(R12);
COPY_REG(R13);
COPY_REG(R14);
COPY_REG(R15);
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86_64::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
memcpy(Frame->State.mm, fpstate->_st, sizeof(Frame->State.mm));
memcpy(Frame->State.xmm, fpstate->_xmm, sizeof(Frame->State.xmm));
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.FTW = fpstate->ftw;
// Deconstruct FSW
Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] = (fpstate->fsw >> 8) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] = (fpstate->fsw >> 9) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] = (fpstate->fsw >> 10) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] = (fpstate->fsw >> 14) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] = (fpstate->fsw >> 11) & 0b111;
}
}
else {
auto *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(Context->UContextLocation);
[[maybe_unused]] auto *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(Context->SigInfoLocation);
// If the guest modified the RIP then we need to take special precautions here
if (Context->OriginalRIP != guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] ||
Context->FaultToTopAndGeneratedException) {
// Hack! Go back to the top of the dispatcher top
// This is only safe inside the JIT rather than anything outside of it
ArchHelpers::Context::SetPc(ucontext, AbsoluteLoopTopAddressFillSRA);
// Set our state register to point to our guest thread data
ArchHelpers::Context::SetState(ucontext, reinterpret_cast<uint64_t>(Frame));
// XXX: Full context setting
// First 32-bytes of flags is EFLAGS broken out
uint32_t eflags = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL];
for (size_t i = 0; i < 32; ++i) {
Frame->State.flags[i] = (eflags & (1U << i)) ? 1 : 0;
}
Frame->State.flags[1] = 1;
Frame->State.flags[9] = 1;
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP];
Frame->State.cs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS];
Frame->State.ds = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS];
Frame->State.es = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES];
Frame->State.fs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS];
Frame->State.gs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS];
Frame->State.ss = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS];
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_##x];
COPY_REG(RDI);
COPY_REG(RSI);
COPY_REG(RBP);
COPY_REG(RBX);
COPY_REG(RDX);
COPY_REG(RAX);
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
// Copy float registers
for (size_t i = 0; i < 8; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
memcpy(&Frame->State.mm[i], &fpstate->_st[i], 10);
}
// Extended XMM state
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
// FCW store default
Frame->State.FCW = fpstate->fcw;
Frame->State.FTW = fpstate->ftw;
// Deconstruct FSW
Frame->State.flags[FEXCore::X86State::X87FLAG_C0_LOC] = (fpstate->fsw >> 8) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C1_LOC] = (fpstate->fsw >> 9) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C2_LOC] = (fpstate->fsw >> 10) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_C3_LOC] = (fpstate->fsw >> 14) & 1;
Frame->State.flags[FEXCore::X86State::X87FLAG_TOP_LOC] = (fpstate->fsw >> 11) & 0b111;
}
}
}
}
static uint32_t ConvertSignalToTrapNo(int Signal, siginfo_t *HostSigInfo) {
@@ -115,7 +275,8 @@ static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
}
bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
StoreThreadState(Signal, ucontext);
auto ContextBackup = StoreThreadState(Signal, ucontext);
auto Frame = ThreadState->CurrentFrame;
// Ref count our faults
@@ -140,15 +301,42 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// Otherwise we might load garbage
if (SRAEnabled) {
if (IsAddressInJITCode(OldPC, false)) {
uint32_t IgnoreMask{};
#ifdef _M_ARM_64
if (Frame->InSyscallInfo != 0) {
// We are in a syscall, this means we are in a weird register state
// We need to spill SRA but only some of it, since some values have already been spilled
// Lower 16 bits tells us which registers are already spilled to the context
// So we ignore spilling those ones
uint16_t NumRegisters = std::popcount(Frame->InSyscallInfo & 0xFFFF);
if (NumRegisters >= 4) {
// Unhandled case
IgnoreMask = 0;
}
else {
IgnoreMask = Frame->InSyscallInfo & 0xFFFF;
}
}
else {
// We must spill everything
IgnoreMask = 0;
}
#endif
// We are in jit, SRA must be spilled
SpillSRA(ucontext);
SpillSRA(ucontext, IgnoreMask);
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_INJIT;
} else {
if (!IsAddressInJITCode(OldPC, true)) {
// This is likely to cause issues but in some cases it isn't fatal
// This can also happen if we have put a signal on hold, then we just reenabled the signal
// So we are in the syscall handler
// Only throw a log message in this case
LogMan::Msg::E("Signals in dispatcher have unsynchronized context");
if constexpr (false) {
// XXX: Messages in the signal handler can cause us to crash
LogMan::Msg::EFmt("Signals in dispatcher have unsynchronized context");
}
}
}
}
@@ -180,8 +368,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// siginfo_t
siginfo_t *HostSigInfo = reinterpret_cast<siginfo_t*>(info);
if (GuestAction->sa_flags & SA_SIGINFO) {
// Backup where we think the RIP currently is
ContextBackup->OriginalRIP = Frame->State.rip;
if (GuestAction->sa_flags & SA_SIGINFO) {
// Setup ucontext a bit
if (Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::_libc_fpstate);
@@ -196,6 +386,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
NewGuestSP = AlignDown(NewGuestSP, alignof(siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
ContextBackup->FPStateLocation = FPStateLocation;
ContextBackup->UContextLocation = UContextLocation;
ContextBackup->SigInfoLocation = SigInfoLocation;
FEXCore::x86_64::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86_64::ucontext_t*>(UContextLocation);
siginfo_t *guest_siginfo = reinterpret_cast<siginfo_t*>(SigInfoLocation);
@@ -209,8 +403,23 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CSGSFS] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
// aarch64 and x86_64 siginfo_t matches. We can just copy this over
// SI_USER could also potentially have random data in it, needs to be bit perfect
// For guest faults we don't have a real way to reconstruct state to a real guest RIP
*guest_siginfo = *HostSigInfo;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = SynchronousFaultData.err_code;
// Overwrite si_code
guest_siginfo->si_code = SynchronousFaultData.si_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
}
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_OLDMASK] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_CR2] = 0;
@@ -255,15 +464,12 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_stack.ss_sp = GuestStack->ss_sp;
guest_uctx->uc_stack.ss_size = GuestStack->ss_size;
// aarch64 and x86_64 siginfo_t matches. We can just copy this over
// SI_USER could also potentially have random data in it, needs to be bit perfect
// For guest faults we don't have a real way to reconstruct state to a real guest RIP
*guest_siginfo = *HostSigInfo;
Frame->State.gregs[X86State::REG_RSI] = SigInfoLocation;
Frame->State.gregs[X86State::REG_RDX] = UContextLocation;
}
else {
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_32BIT;
NewGuestSP -= sizeof(FEXCore::x86::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::_libc_fpstate));
uint64_t FPStateLocation = NewGuestSP;
@@ -276,6 +482,10 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::siginfo_t));
uint64_t SigInfoLocation = NewGuestSP;
ContextBackup->FPStateLocation = FPStateLocation;
ContextBackup->UContextLocation = UContextLocation;
ContextBackup->SigInfoLocation = SigInfoLocation;
FEXCore::x86::ucontext_t *guest_uctx = reinterpret_cast<FEXCore::x86::ucontext_t*>(UContextLocation);
FEXCore::x86::siginfo_t *guest_siginfo = reinterpret_cast<FEXCore::x86::siginfo_t*>(SigInfoLocation);
@@ -290,8 +500,16 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS] = Frame->State.fs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = SynchronousFaultData.TrapNo;
guest_siginfo->si_code = SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = SynchronousFaultData.err_code;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_siginfo->si_code = HostSigInfo->si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
}
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS] = Frame->State.cs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL] = 0;
@@ -339,7 +557,6 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
// These three elements are in every siginfo
guest_siginfo->si_signo = HostSigInfo->si_signo;
guest_siginfo->si_errno = HostSigInfo->si_errno;
guest_siginfo->si_code = HostSigInfo->si_code;
switch (Signal) {
case SIGSEGV:
@@ -368,7 +585,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
guest_siginfo->_sifields._timer.sigval.sival_int = HostSigInfo->si_int;
break;
default:
LogMan::Msg::E("Unhandled siginfo_t for signal: %d\n", Signal);
LogMan::Msg::EFmt("Unhandled siginfo_t for signal: {}\n", Signal);
break;
}
@@ -402,7 +619,7 @@ bool Dispatcher::HandleGuestSignal(int Signal, void *info, void *ucontext, Guest
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SignalReturn;
LOGMAN_THROW_A(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
LOGMAN_THROW_A_FMT(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[FEXCore::X86State::REG_RSP] = NewGuestSP;
}
@@ -454,7 +671,8 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddress);
}
@@ -486,11 +704,21 @@ bool Dispatcher::HandleSignalPause(int Signal, void *info, void *ucontext) {
} else {
if (SRAEnabled) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true), "Signals in dispatcher have unsynchronized context");
LOGMAN_THROW_A_FMT(!IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext), true),
"Signals in dispatcher have unsynchronized context");
}
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddress);
}
// We need to be a little bit careful here
// If we were already paused (due to GDB) and we are immediately stopping (due to gdb kill)
// Then we need to ensure we don't double decrement our idle thread counter
if (ThreadState->RunningEvents.ThreadSleeping) {
// If the thread was sleeping then its idle counter was decremented
// Reincrement it here to not break logic
++ThreadState->CTX->IdleWaitRefCount;
}
ThreadState->SignalReason.store(FEXCore::Core::SignalEvent::Nothing);
return true;
}
@@ -531,15 +759,19 @@ void Dispatcher::RemoveCodeBuffer(uint8_t* start_to_remove) {
}
}
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher) const {
bool Dispatcher::IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher, bool IncludeCompileService) const {
for (auto [start, end] : CodeBuffers) {
if (Address >= start && Address < end) {
return true;
}
}
if (IncludeDispatcher) {
return IsAddressInDispatcher(Address);
if (IncludeDispatcher && IsAddressInDispatcher(Address)) {
return true;
}
if (IncludeCompileService && ThreadState->CompileService && ThreadState->CompileService->IsAddressInJITCode(Address)) {
return true;
}
return false;
}
@@ -3,9 +3,10 @@
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/MContext.h"
#include <bits/types/stack_t.h>
#include <cstdint>
#include <signal.h>
#include <stddef.h>
#include <stack>
#include <tuple>
@@ -47,11 +48,20 @@ public:
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint64_t PauseReturnInstruction{};
/** @} */
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
} SynchronousFaultData;
uint64_t Start{};
uint64_t End{};
@@ -67,7 +77,7 @@ public:
void RemoveCodeBuffer(uint8_t* start);
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true) const;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const;
bool IsAddressInDispatcher(uint64_t Address) const {
return Address >= Start && Address < End;
}
@@ -77,12 +87,12 @@ protected:
: CTX {ctx}
, ThreadState {Thread} {}
void StoreThreadState(int Signal, void *ucontext);
ArchHelpers::Context::ContextBackup* StoreThreadState(int Signal, void *ucontext);
void RestoreThreadState(void *ucontext);
std::stack<uint64_t> SignalFrames;
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(void *ucontext) {}
virtual void SpillSRA(void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
@@ -11,6 +11,7 @@
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <cmath>
#include <memory>
@@ -94,7 +95,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
// L1 Cache
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rax, rdx);
and_(rax, LookupCache::L1_ENTRIES_MASK);
@@ -113,8 +114,9 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
mov(rax, rdx);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
mov(rbx, VirtualMemorySize - 1);
and_(rax, rbx);
shr(rax, 12);
@@ -141,8 +143,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
je(NoBlock);
// Update L1
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(r13, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rcx, rdx);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
@@ -251,8 +252,7 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
// XXX: XMM?
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(&SignalHandlerRefCounter));
add(dword [rax], 1);
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
// Now push the callback return trampoline to the guest stack
// Guest will be misaligned because calling a thunk won't correct the guest's stack once we call the callback from the host
@@ -274,10 +274,33 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
{
// Signal return handler
SignalHandlerReturnAddress = getCurr<uint64_t>();
ud2();
}
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = getCurr<uint64_t>();
ud2();
}
{
// Guest Overflow handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = getCurr<uint64_t>();
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
mov(rax, reinterpret_cast<uint64_t>(&SynchronousFaultData));
add(byte [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, FaultToTopAndGeneratedException)], 1);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, TrapNo)], X86State::X86_TRAPNO_OF);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, err_code)], 0);
mov(dword [rax + offsetof(Dispatcher::SynchronousFaultDataStruct, si_code)], 0x80);
hlt();
}
{
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
@@ -307,12 +330,26 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, FEXCore::Core::Inte
End = Start + getSize();
if (CTX->Config.BlockJITNaming()) {
std::string Name = "Dispatch_" + std::to_string(::gettid());
std::string Name = "Dispatch_" + std::to_string(FHU::Syscalls::gettid());
CTX->Symbols.Register(reinterpret_cast<void*>(Start), End-Start, Name);
}
if (CTX->Config.GlobalJITNaming()) {
CTX->Symbols.RegisterJITSpace(reinterpret_cast<void*>(Start), End-Start);
}
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
Pointers.DispatcherLoopTop = AbsoluteLoopTopAddress;
Pointers.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Pointers.ThreadStopHandler = ThreadStopHandlerAddress;
Pointers.ThreadPauseHandler = ThreadPauseHandlerAddress;
Pointers.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Pointers.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Pointers.SignalReturnHandler = SignalHandlerReturnAddress;
Pointers.L1Pointer = Thread->LookupCache->GetL1Pointer();
}
}
X86Dispatcher::~X86Dispatcher() {
+74 -52
View File
@@ -192,7 +192,7 @@ Decoder::~Decoder() {
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_A(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
@@ -209,7 +209,7 @@ uint64_t Decoder::ReadData(uint8_t Size) {
}
if (Size > sizeof(uint64_t)) {
LOGMAN_MSG_A("Unknown data size to read");
LOGMAN_MSG_A_FMT("Unknown data size to read");
return 0;
}
@@ -351,10 +351,9 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
Operand->Data.SIB.Index = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X ? 1 : 0, SIB.index, false, false, false, false, 0b100);
Operand->Data.SIB.Base = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
uint64_t Literal {0};
LOGMAN_THROW_A(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
LOGMAN_THROW_A_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
Literal = ReadData(Displacement);
uint64_t Literal = ReadData(Displacement);
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -364,8 +363,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
// Explained in Table 1-14. "Operand Addressing Using ModRM and SIB Bytes"
if (ModRM.rm == 0b101) {
// 32bit Displacement
uint32_t Literal;
Literal = ReadData(4);
const uint32_t Literal = ReadData(4);
Operand->Type = DecodedOperand::OpType::RIPRelative;
Operand->Data.RIPLiteral.Value.u = Literal;
@@ -378,8 +376,7 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
}
else {
uint8_t DisplacementSize = ModRM.mod == 1 ? 1 : 4;
uint32_t Literal{};
Literal = ReadData(DisplacementSize);
uint32_t Literal = ReadData(DisplacementSize);
if (DisplacementSize == 1) {
Literal = static_cast<int8_t>(Literal);
}
@@ -397,26 +394,26 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
LOGMAN_THROW_A(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
uint8_t DestSize{};
const bool HasWideningDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_WIDENING_SIZE_LAST) != 0 ||
Options.w;
(Options.w && CTX->Config.Is64BitMode);
const bool HasNarrowingDisplacement = (FEXCore::X86Tables::DecodeFlags::GetOpAddr(DecodeInst->Flags, 0) & FEXCore::X86Tables::DecodeFlags::FLAG_OPERAND_SIZE_LAST) != 0;
bool HasXMMSrc = !!(Info->Flags & FEXCore::X86Tables::InstFlags::FLAGS_XMM_FLAGS) &&
@@ -535,7 +532,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LOGMAN_THROW_A(!HasMODRM, "This instruction shouldn't have ModRM!");
LOGMAN_THROW_A_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
// This also means that the destination is always a GPR on these ones
@@ -586,8 +583,11 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
return false;
}
else {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&NonGPR, ModRM);
// Only decode if we haven't pre-decoded
if (NonGPR.IsNone()) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&NonGPR, ModRM);
}
}
return true;
@@ -641,7 +641,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (Bytes != 0) {
LOGMAN_THROW_A(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
@@ -666,7 +666,8 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LOGMAN_THROW_A(Bytes == 0, "Inst at 0x%lx: 0x%04x '%s' Had an instruction of size %d with %d remaining", DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name, InstructionSize, Bytes);
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining",
DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
}
@@ -677,21 +678,22 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
// XXX: Once we support 32bit x86 then this will be necessary to support
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
LogMan::Msg::DFmt("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Msg::DFmt("Invalid or Unknown instruction: {} 0x{:04x} 0x{:x}", Info->Name ?: "UND", Op, DecodeInst->PC);
return false;
}
LOGMAN_THROW_A(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX,
"REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
@@ -747,7 +749,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
3,
};
uint8_t Field = RegToField[ModRM.reg];
LOGMAN_THROW_A(Field != 255, "Invalid field selected!");
LOGMAN_THROW_A_FMT(Field != 255, "Invalid field selected!");
LocalOp = (Field << 3) | ModRM.rm;
return NormalOp(&SecondModRMTableOps[LocalOp], LocalOp);
@@ -775,6 +777,11 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
const uint8_t Byte1 = ReadByte();
DecodedHeader options{};
if ((Byte1 & 0b10000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.R shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
}
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
@@ -785,8 +792,15 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
options.w = (Byte2 & 0b10000000) != 0;
if ((Byte1 & 0b01000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.X shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
}
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::E("We don't understand a map_select of: %d", map_select);
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
return false;
}
}
@@ -846,7 +860,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
case 0x0F: {// Escape Op
uint8_t EscapeOp = ReadByte();
switch (EscapeOp) {
case 0x0F: { // 3DNow!
case 0x0F: [[unlikely]] { // 3DNow!
// 3DNow! Instruction Encoding: 0F 0F [ModRM] [SIB] [Displacement] [Opcode]
// Decode ModRM
uint8_t ModRMByte = ReadByte();
@@ -859,8 +873,12 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
const bool Has16BitAddressing = !CTX->Config.Is64BitMode &&
DecodeInst->Flags & DecodeFlags::FLAG_ADDRESS_SIZE;
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
// All 3DNow! instructions have the second argument as the rm handler
// We need to decode it upfront to get the displacement out of the way
if (ModRM.mod != 0b11) {
auto Disp = DecodeModRMs_Disp[Has16BitAddressing];
(this->*Disp)(&DecodeInst->Src[0], ModRM);
}
// Take a peek at the op just past the displacement
uint8_t LocalOp = ReadByte();
@@ -869,14 +887,20 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
}
case 0x38: { // F38 Table!
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = 1;
constexpr uint16_t PF_38_F2 = 2;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
uint16_t Prefix = PF_38_NONE;
if (DecodeInst->LastEscapePrefix == 0xF2) // REPNE
Prefix = PF_38_F2;
else if (DecodeInst->LastEscapePrefix == 0x66) // Operand Size
Prefix = PF_38_66;
if (DecodeInst->Flags & DecodeFlags::FLAG_OPERAND_SIZE) {
Prefix |= PF_38_66;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REPNE_PREFIX) {
Prefix |= PF_38_F2;
}
if (DecodeInst->Flags & DecodeFlags::FLAG_REP_PREFIX) {
Prefix |= PF_38_F3;
}
uint16_t LocalOp = (Prefix << 8) | ReadByte();
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
@@ -991,7 +1015,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
LOGMAN_THROW_A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
@@ -1044,13 +1068,13 @@ void Decoder::BranchTargetInMultiblockRange() {
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
break;
}
case 0xE9:
case 0xEB: // Both are unconditional JMP instructions
LOGMAN_THROW_A(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
Conditional = false;
break;
@@ -1116,7 +1140,7 @@ const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC) {
Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
@@ -1130,7 +1154,6 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
EntryPoint = PC;
InstStream = _InstStream;
bool ErrorDuringDecoding = false;
uint64_t TotalInstructions{};
// If we don't have symbols available then we become a bit optimistic about multiblock ranges
@@ -1162,20 +1185,15 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
InstStream = AdjustAddrForSpecialRegion(_InstStream, EntryPoint, RIPToDecode);
while (1) {
ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
bool ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
LogMan::Msg::D("Couldn't Decode something at 0x%lx, Started at 0x%lx", PC + PCOffset, PC);
if (Blocks.size() == 1) {
return false;
}
LOGMAN_THROW_A(Blocks.size() != 1, "Decode Error in entry block");
LogMan::Msg::DFmt("Couldn't Decode something at 0x{:x}, Started at 0x{:x}", PC + PCOffset, PC);
// Put an invalid instruction in the stream so the core can raise SIGILL if hit
CurrentBlockDecoding.HasInvalidInstruction = true;
if (ErrorDuringDecoding && Blocks.size() != 1) {
ErrorDuringDecoding = false;
}
break;
// Error while decoding instruction. We don't know the table or instruction size
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
}
DecodedMinAddress = std::min(DecodedMinAddress, RIPToDecode + PCOffset);
@@ -1184,6 +1202,11 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
++BlockNumberOfInstructions;
++DecodedSize;
// Can not continue this block at all on invalid instruction
if (CurrentBlockDecoding.HasInvalidInstruction) {
break;
}
bool CanContinue = false;
if (!(DecodeInst->TableInfo->Flags &
(FEXCore::X86Tables::InstFlags::FLAGS_BLOCK_END | FEXCore::X86Tables::InstFlags::FLAGS_SETS_RIP))) {
@@ -1228,7 +1251,6 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
std::sort(Blocks.begin(), Blocks.end(), [](const FEXCore::Frontend::Decoder::DecodedBlocks& a, const FEXCore::Frontend::Decoder::DecodedBlocks& b) {
return a.Entry < b.Entry;
});
return !ErrorDuringDecoding;
}
}
+3 -3
View File
@@ -27,7 +27,7 @@ public:
Decoder(FEXCore::Context::Context *ctx);
~Decoder();
bool DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC);
std::vector<DecodedBlocks> const *GetDecodedBlocks() const {
return &Blocks;
@@ -35,7 +35,7 @@ public:
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(std::set<uint64_t> *v) { ExternalBranches = v; }
private:
@@ -91,7 +91,7 @@ private:
void DecodeModRM_16(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
const std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp {
static constexpr std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp{
&FEXCore::Frontend::Decoder::DecodeModRM_64,
&FEXCore::Frontend::Decoder::DecodeModRM_16,
};
File diff suppressed because it is too large. Load diff
+21
View File
@@ -6,8 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/Config/Config.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <atomic>
#include <istream>
#include <memory>
#include <mutex>
@@ -27,9 +29,14 @@ public:
// Public for threading
void GdbServerLoop();
void AlertLibrariesChanged() {
LibraryMapChanged = true;
}
private:
void Break(int signal);
void OpenListenSocket();
std::unique_ptr<std::iostream> OpenSocket();
void StartThread();
std::string ReadPacket(std::iostream &stream);
@@ -37,6 +44,9 @@ private:
void SendACK(std::ostream &stream, bool NACK);
Event ThreadBreakEvent{};
void WaitForThreadWakeup();
struct HandledPacketType {
std::string Response{};
enum ResponseType {
@@ -60,6 +70,8 @@ private:
HandledPacketType handleBreakpoint(const std::string &packet);
HandledPacketType handleProgramOffsets();
HandledPacketType ThreadAction(char action, uint32_t tid);
std::string readRegs();
HandledPacketType readReg(const std::string& packet);
@@ -69,8 +81,17 @@ private:
std::mutex sendMutex;
bool SettingNoAckMode{false};
bool NoAckMode{false};
bool NonStopMode{false};
std::string ThreadString{};
std::string OSDataString{};
void buildLibraryMap();
std::atomic<bool> LibraryMapChanged = true;
std::string LibraryMapString{};
// Used to keep track of which signals to pass to the guest
std::array<bool, SignalDelegator::MAX_SIGNALS + 1> PassSignals{};
uint32_t CurrentDebuggingThread{};
int ListenSocket{};
FEX_CONFIG_OPT(Filename, APP_FILENAME);
};
+104
View File
@@ -1,3 +1,4 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#ifdef _M_ARM_64
@@ -13,14 +14,117 @@
namespace FEXCore {
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
static uint32_t GetDCZID() {
uint64_t Result{};
__asm("mrs %[Res], DCZID_EL0"
: [Res] "=r" (Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result{};
__asm ("mrs %[Res], FPCR"
: [Res] "=r" (Result));
return Result;
}
static void SetFPCR(uint64_t Value) {
__asm ("msr FPCR, %[Value]"
:: [Value] "r" (Value));
}
#else
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
#endif
HostFeatures::HostFeatures() {
#ifdef _M_ARM_64
auto Features = vixl::CPUFeatures::InferFromOS();
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
SupportsRAND = Features.Has(vixl::CPUFeatures::Feature::kRNG);
// Only supported when FEAT_AFP is supported
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
// RCPC is bugged on Snapdragon 865
// Causes glibc cond16 test to immediately throw assert
// __pthread_mutex_cond_lock: Assertion `mutex->__data.__owner == 0'
SupportsRCPC = false; //Features.Has(vixl::CPUFeatures::Feature::kRCpc);
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#endif
#ifdef _M_X86_64
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
constexpr uint32_t ExceptionEnableTraps =
(1U << 8) | // Invalid Operation float exception trap enable
(1U << 9) | // Divide by zero float exception trap enable
(1U << 10) | // Overflow float exception trap enable
(1U << 11) | // Underflow float exception trap enable
(1U << 12) | // Inexact float exception trap enable
(1U << 15); // Input Denormal float exception trap enable
uint32_t OriginalFPCR = GetFPCR();
uint32_t FPCR = OriginalFPCR | ExceptionEnableTraps;
SetFPCR(FPCR);
FPCR = GetFPCR();
SupportsFloatExceptions = (FPCR & ExceptionEnableTraps) == ExceptionEnableTraps;
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
#endif
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
SupportsCLZERO = DCZID_Bytes == CPUIDEmu::CACHELINE_SIZE;
}
}
}
+19
View File
@@ -1,9 +1,28 @@
#pragma once
#include <cstdint>
namespace FEXCore {
class HostFeatures final {
public:
HostFeatures();
/**
* @brief Backend features that change how codegen is generated from IR
*
* Specifically things that affect the IR->Codegen process
* Not the x86->IR process
*/
uint32_t DCacheLineSize{};
uint32_t ICacheLineSize{};
bool SupportsAES{};
bool SupportsCRC{};
bool SupportsCLZERO{};
bool SupportsAtomics{};
bool SupportsRCPC{};
bool SupportsRAND{};
// Float exception behaviour
bool SupportsFlushInputsToZero{};
bool SupportsFloatExceptions{};
};
}
+64 -64
View File
@@ -13,12 +13,12 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
uint64_t *Src = GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Result{};
@@ -27,7 +27,7 @@ DEF_OP(TruncElementPair) {
GD = Result;
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -38,7 +38,13 @@ DEF_OP(Constant) {
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
GD = Data->CurrentEntry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
GD = (Data->CurrentEntry + Op->Offset) & Mask;
}
DEF_OP(InlineConstant) {
@@ -493,6 +499,54 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
if (OpSize != 4 && OpSize != 8) {
LOGMAN_MSG_A_FMT("Unknown PDep Size: {}\n", OpSize);
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
uint64_t Result = 0;
for (uint64_t Index = 0; Mask > 0; Index++) {
const uint64_t Offset = std::countr_zero(Mask);
Mask &= Mask - 1;
Result |= ((Input >> Index) & 1) << Offset;
}
GD = Result;
}
DEF_OP(PExt) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
if (OpSize != 4 && OpSize != 8) {
LOGMAN_MSG_A_FMT("Unknown PExt Size: {}\n", OpSize);
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
uint64_t Result = 0;
for (uint64_t Offset = 0; Mask > 0; Offset++) {
const uint64_t Index = std::countr_zero(Mask);
Mask &= Mask - 1;
Result |= ((Input >> Index) & 1) << Offset;
}
GD = Result;
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -787,9 +841,8 @@ DEF_OP(Bfi) {
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for BFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
uint64_t SourceMask = (1ULL << Op->Width) - 1;
if (Op->Width == 64)
SourceMask = ~0ULL;
@@ -800,9 +853,8 @@ DEF_OP(Bfe) {
DEF_OP(Sbfe) {
auto Op = IROp->C<IR::IROp_Sbfe>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 8, "OpSize is too large for SBFE: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for SBFE: {}", IROp->Size);
int64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t ShiftLeftAmount = (64 - (Op->Width + Op->lsb));
uint64_t ShiftRightAmount = ShiftLeftAmount + Op->lsb;
@@ -841,15 +893,14 @@ DEF_OP(Select) {
DEF_OP(VExtractToGPR) {
auto Op = IROp->C<IR::IROp_VExtractToGPR>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractToGPR: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractToGPR: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Idx * 8;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
SourceMask = ~0ULL;
@@ -860,7 +911,7 @@ DEF_OP(VExtractToGPR) {
}
else {
uint64_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Idx * 8;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
SourceMask = ~0ULL;
@@ -974,55 +1025,4 @@ DEF_OP(FCmp) {
#undef DEF_OP
void InterpreterOps::RegisterALUHandlers() {
#define REGISTER_OP(op, x) FEXCore::CPU::InterpreterOps::OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
REGISTER_OP(NEG, Neg);
REGISTER_OP(MUL, Mul);
REGISTER_OP(UMUL, UMul);
REGISTER_OP(DIV, Div);
REGISTER_OP(UDIV, UDiv);
REGISTER_OP(REM, Rem);
REGISTER_OP(UREM, URem);
REGISTER_OP(MULH, MulH);
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
REGISTER_OP(LUREM, LURem);
REGISTER_OP(NOT, Not);
REGISTER_OP(POPCOUNT, Popcount);
REGISTER_OP(FINDLSB, FindLSB);
REGISTER_OP(FINDMSB, FindMSB);
REGISTER_OP(FINDTRAILINGZEROS, FindTrailingZeros);
REGISTER_OP(COUNTLEADINGZEROES, CountLeadingZeroes);
REGISTER_OP(REV, Rev);
REGISTER_OP(BFI, Bfi);
REGISTER_OP(BFE, Bfe);
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
+114 -133
View File
@@ -310,33 +310,32 @@ uint64_t AtomicCompareAndSwap(uint64_t expected, uint64_t desired, uint64_t *add
#endif
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
// Size is the size of each pair element
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint64_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint64_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint64_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 8: {
std::atomic<__uint128_t> *MemData = *GetSrc<std::atomic<__uint128_t> **>(Data->SSAData, Op->Header.Args[2]);
std::atomic<__uint128_t> *MemData = *GetSrc<std::atomic<__uint128_t> **>(Data->SSAData, Op->Addr);
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Expected);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Desired);
__uint128_t Expected = Src1;
bool Result = MemData->compare_exchange_strong(Expected, Src2);
memcpy(GDP, Result ? &Src1 : &Expected, 16);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", OpSize); break;
default: LOGMAN_MSG_A_FMT("Unknown CAS size: {}", IROp->ElementSize); break;
}
}
@@ -347,33 +346,33 @@ DEF_OP(CAS) {
switch (OpSize) {
case 1: {
GD = AtomicCompareAndSwap(
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint8_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint8_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint8_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint8_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 2: {
GD = AtomicCompareAndSwap(
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint16_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint16_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint16_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint16_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 4: {
GD = AtomicCompareAndSwap(
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint32_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint32_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint32_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint32_t**>(Data->SSAData, Op->Addr)
);
break;
}
case 8: {
GD = AtomicCompareAndSwap(
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]),
*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]),
*GetSrc<uint64_t**>(Data->SSAData, Op->Header.Args[2])
*GetSrc<uint64_t*>(Data->SSAData, Op->Expected),
*GetSrc<uint64_t*>(Data->SSAData, Op->Desired),
*GetSrc<uint64_t**>(Data->SSAData, Op->Addr)
);
break;
}
@@ -385,26 +384,26 @@ DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData += Src;
break;
}
@@ -416,26 +415,26 @@ DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData -= Src;
break;
}
@@ -447,26 +446,26 @@ DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData &= Src;
break;
}
@@ -478,26 +477,26 @@ DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData |= Src;
break;
}
@@ -509,26 +508,26 @@ DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
*MemData ^= Src;
break;
}
@@ -540,29 +539,29 @@ DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->exchange(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->exchange(Src);
GD = Previous;
break;
@@ -575,29 +574,29 @@ DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_add(Src);
GD = Previous;
break;
@@ -610,29 +609,29 @@ DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_sub(Src);
GD = Previous;
break;
@@ -645,29 +644,29 @@ DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_and(Src);
GD = Previous;
break;
@@ -680,29 +679,29 @@ DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_or(Src);
GD = Previous;
break;
@@ -715,29 +714,29 @@ DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
switch (IROp->Size) {
case 1: {
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Header.Args[0]);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint8_t> *MemData = *GetSrc<std::atomic<uint8_t> **>(Data->SSAData, Op->Addr);
uint8_t Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uint8_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 2: {
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Header.Args[0]);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint16_t> *MemData = *GetSrc<std::atomic<uint16_t> **>(Data->SSAData, Op->Addr);
uint16_t Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Value);
uint16_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 4: {
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Header.Args[0]);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint32_t> *MemData = *GetSrc<std::atomic<uint32_t> **>(Data->SSAData, Op->Addr);
uint32_t Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Value);
uint32_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
}
case 8: {
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
std::atomic<uint64_t> *MemData = *GetSrc<std::atomic<uint64_t> **>(Data->SSAData, Op->Addr);
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
uint64_t Previous = MemData->fetch_xor(Src);
GD = Previous;
break;
@@ -751,22 +750,22 @@ DEF_OP(AtomicFetchNeg) {
switch (IROp->Size) {
case 1: {
using Type = uint8_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 2: {
using Type = uint16_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 4: {
using Type = uint32_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
case 8: {
using Type = uint64_t;
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Header.Args[0]));
GD = AtomicFetchNeg(*GetSrc<Type**>(Data->SSAData, Op->Addr));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
@@ -774,23 +773,5 @@ DEF_OP(AtomicFetchNeg) {
}
#undef DEF_OP
void InterpreterOps::RegisterAtomicHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
REGISTER_OP(ATOMICSUB, AtomicSub);
REGISTER_OP(ATOMICAND, AtomicAnd);
REGISTER_OP(ATOMICOR, AtomicOr);
REGISTER_OP(ATOMICXOR, AtomicXor);
REGISTER_OP(ATOMICSWAP, AtomicSwap);
REGISTER_OP(ATOMICFETCHADD, AtomicFetchAdd);
REGISTER_OP(ATOMICFETCHSUB, AtomicFetchSub);
REGISTER_OP(ATOMICFETCHAND, AtomicFetchAnd);
REGISTER_OP(ATOMICFETCHOR, AtomicFetchOr);
REGISTER_OP(ATOMICFETCHXOR, AtomicFetchXor);
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -13,6 +13,7 @@ $end_info$
#include <FEXCore/HLE/SyscallHandler.h>
#include <cstdint>
#include <unistd.h>
namespace FEXCore::CPU {
[[noreturn]]
@@ -23,7 +24,7 @@ static void SignalReturn(FEXCore::Core::InternalThreadState *Thread) {
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -32,10 +33,6 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
SignalReturn(Data->State);
}
@@ -104,6 +101,34 @@ DEF_OP(Syscall) {
GD = Res;
}
DEF_OP(InlineSyscall) {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
FEXCore::HLE::SyscallArguments Args;
for (size_t j = 0; j < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++j) {
if (Op->Header.Args[j].IsInvalid()) break;
Args.Argument[j] = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[j]);
}
// We don't want the errno handling but I also don't want to write inline ASM atm
uint64_t Res = syscall(
Op->HostSyscallNumber,
Args.Argument[0],
Args.Argument[1],
Args.Argument[2],
Args.Argument[3],
Args.Argument[4],
Args.Argument[5],
Args.Argument[6]
);
if (Res == -1) {
Res = -errno;
}
GD = Res;
}
DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
@@ -137,21 +162,5 @@ DEF_OP(CPUID) {
}
#undef DEF_OP
void InterpreterOps::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,7 +11,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
uint8_t OpSize = IROp->Size;
@@ -19,7 +19,7 @@ DEF_OP(VInsGPR) {
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Offset = Op->Index * Op->Header.ElementSize * 8;
uint64_t Offset = Op->DestIdx * Op->Header.ElementSize * 8;
__uint128_t Mask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
if (Op->Header.ElementSize == 8) {
Mask = ~0ULL;
@@ -220,18 +220,5 @@ DEF_OP(Vector_FToI) {
}
#undef DEF_OP
void InterpreterOps::RegisterConversionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -298,8 +298,65 @@ namespace AES {
}
}
namespace CRC32 {
// CRC32 per byte lookup table.
constexpr std::array<uint32_t, 256> CRC32CTable = []() consteval {
std::array<uint32_t, 256> Table{};
// Clang 11.x doesn't support bitreverse as a consteval
// constexpr uint32_t Polynomial = 0x1EDC6F41;
constexpr uint32_t PolynomialRev = 0x82F63B78; //__builtin_bitreverse32(Polynomial);
for (size_t Char = 0; Char < std::size(Table); ++Char) {
uint32_t CurrentChar = Char;
for (size_t i = 0; i < 8; ++i) {
if (CurrentChar & 1) {
CurrentChar = (CurrentChar >> 1) ^ PolynomialRev;
}
else {
CurrentChar >>= 1;
}
}
Table[Char] = CurrentChar;
}
return Table;
}();
uint32_t crc32cb(uint32_t Accumulator, uint8_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ data] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32ch(uint32_t Accumulator, uint16_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cw(uint32_t Accumulator, uint32_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
uint32_t crc32cx(uint32_t Accumulator, uint64_t data) {
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 0) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 8) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 16) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 24) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 32) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 40) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 48) & 0xFF)] ^ Accumulator >> 8;
Accumulator = CRC32CTable[(uint8_t)Accumulator ^ ((data >> 56) & 0xFF)] ^ Accumulator >> 8;
return Accumulator;
}
}
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -429,15 +486,33 @@ DEF_OP(AESKeyGenAssist) {
memcpy(GDP, &Tmp, sizeof(Tmp));
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
uint32_t Src1 = *GetSrc<uint32_t*>(Data->SSAData, Op->Src1);
uint8_t *Src2 = GetSrc<uint8_t*>(Data->SSAData, Op->Src2);
uint32_t Tmp{};
switch (Op->SrcSize) {
case 1:
Tmp = CRC32::crc32cb(Src1, *(uint8_t*)Src2);
break;
case 2:
Tmp = CRC32::crc32ch(Src1, *(uint16_t*)Src2);
break;
case 4:
Tmp = CRC32::crc32cw(Src1, *(uint32_t*)Src2);
break;
case 8:
Tmp = CRC32::crc32cx(Src1, *(uint64_t*)Src2);
break;
default:
LOGMAN_MSG_A_FMT("Unknown CRC32C size: {}", Op->SrcSize);
break;
}
memcpy(GDP, &Tmp, sizeof(Tmp));
}
#undef DEF_OP
void InterpreterOps::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -13,7 +13,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(F80LOADFCW) {
FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle(*GetSrc<uint16_t*>(Data->SSAData, IROp->Args[0]));
}
@@ -158,7 +158,7 @@ DEF_OP(F80CVTINT) {
DEF_OP(F80CVTTO) {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
switch (Op->SrcSize) {
case 4: {
float Src = *GetSrc<float *>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
@@ -171,14 +171,14 @@ DEF_OP(F80CVTTO) {
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->SrcSize);
}
}
DEF_OP(F80CVTTOINT) {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
switch (Op->SrcSize) {
case 2: {
int16_t Src = *GetSrc<int16_t*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp = Src;
@@ -191,7 +191,7 @@ DEF_OP(F80CVTTOINT) {
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->Size);
default: LogMan::Msg::DFmt("Unhandled size: {}", Op->SrcSize);
}
}
@@ -323,7 +323,7 @@ DEF_OP(F80BCDLOAD) {
DEF_OP(F80BCDSTORE) {
auto Op = IROp->C<IR::IROp_F80BCDStore>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src1 = X80SoftFloat::FRNDINT(*GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]));
bool Negative = Src1.Sign;
// Clear the Sign bit
@@ -357,33 +357,5 @@ DEF_OP(F80BCDSTORE) {
}
#undef DEF_OP
void InterpreterOps::RegisterF80Handlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(F80LOADFCW, F80LOADFCW);
REGISTER_OP(F80ADD, F80ADD);
REGISTER_OP(F80SUB, F80SUB);
REGISTER_OP(F80MUL, F80MUL);
REGISTER_OP(F80DIV, F80DIV);
REGISTER_OP(F80FYL2X, F80FYL2X);
REGISTER_OP(F80ATAN, F80ATAN);
REGISTER_OP(F80FPREM1, F80FPREM1);
REGISTER_OP(F80FPREM, F80FPREM);
REGISTER_OP(F80SCALE, F80SCALE);
REGISTER_OP(F80CVT, F80CVT);
REGISTER_OP(F80CVTINT, F80CVTINT);
REGISTER_OP(F80CVTTO, F80CVTTO);
REGISTER_OP(F80CVTTOINT, F80CVTTOINT);
REGISTER_OP(F80ROUND, F80ROUND);
REGISTER_OP(F80F2XM1, F80F2XM1);
REGISTER_OP(F80TAN, F80TAN);
REGISTER_OP(F80SQRT, F80SQRT);
REGISTER_OP(F80SIN, F80SIN);
REGISTER_OP(F80COS, F80COS);
REGISTER_OP(F80XTRACT_EXP, F80XTRACT_EXP);
REGISTER_OP(F80XTRACT_SIG, F80XTRACT_SIG);
REGISTER_OP(F80CMP, F80CMP);
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -227,6 +227,8 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
static X80SoftFloat handle(X80SoftFloat Src1) {
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(Src1);
// Clear the Sign bit
Src1.Sign = 0;
@@ -11,17 +11,11 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
GD = (*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]) >> Op->Flag) & 1;
}
#undef DEF_OP
void InterpreterOps::RegisterFlagHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -37,7 +37,9 @@ public:
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
static void InitializeInterpreterOpHandlers();
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
bool NeedsRetainedIRCopy() const override { return true; }
private:
FEXCore::Context::Context *CTX;
@@ -11,7 +11,6 @@
#include <FEXCore/Utils/LogManager.h>
#include <memory>
#include <bits/types/stack_t.h>
#include <signal.h>
#include <stdint.h>
#include <unordered_map>
@@ -35,24 +34,6 @@ static void InterpreterExecution(FEXCore::Core::CpuStateFrame *Frame) {
InterpreterOps::InterpretIR(Thread, Thread->CurrentFrame->State.rip, LocalEntry->second.IR.get(), LocalEntry->second.DebugData.get());
}
void InitializeInterpreterOpHandlers() {
for (uint32_t i = 0; i <= FEXCore::IR::IROps::OP_LAST; ++i) {
InterpreterOps::OpHandlers[i] = &InterpreterOps::Op_Unhandled;
}
InterpreterOps::RegisterALUHandlers();
InterpreterOps::RegisterAtomicHandlers();
InterpreterOps::RegisterBranchHandlers();
InterpreterOps::RegisterConversionHandlers();
InterpreterOps::RegisterFlagHandlers();
InterpreterOps::RegisterMemoryHandlers();
InterpreterOps::RegisterMiscHandlers();
InterpreterOps::RegisterMoveHandlers();
InterpreterOps::RegisterVectorHandlers();
InterpreterOps::RegisterEncryptionHandlers();
InterpreterOps::RegisterF80Handlers();
}
InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: CTX {ctx}
, State {Thread} {
@@ -60,26 +41,28 @@ InterpreterCore::InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::
if (!CompileThread &&
CTX->Config.Core == FEXCore::Config::CONFIG_INTERPRETER) {
CreateAsmDispatch(ctx, Thread);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
}
}
void InterpreterCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(true, Signal, info, ucontext);
}, true);
#endif
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
InterpreterCore *Core = reinterpret_cast<InterpreterCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -91,4 +74,8 @@ std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx
return std::make_unique<InterpreterCore>(ctx, Thread, CompileThread);
}
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX) {
InterpreterCore::InitializeSignalHandlers(CTX);
}
}
@@ -13,10 +13,10 @@ namespace FEXCore::Core {
namespace FEXCore::CPU {
class CPUBackend;
void InitializeInterpreterOpHandlers();
[[nodiscard]] std::unique_ptr<CPUBackend> CreateInterpreterCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeInterpreterSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
@@ -161,19 +161,19 @@
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID()];
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, uint32_t Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op];
Res GetDest(void* SSAData, FEXCore::IR::NodeID Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID()];
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
@@ -0,0 +1,272 @@
#include "FEXCore/Core/CoreState.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/F80Ops.h"
#include <cstddef>
#include <cstdint>
namespace FEXCore::CPU {
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_UNKNOWN, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F32, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F64, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_I16, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_VOID_U16, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_I32, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F32_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I16_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I32_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I64_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_I64_F80_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F80, (void*)fn, HandlerIndex};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F80_F80_F80, (void*)fn, HandlerIndex};
}
void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80LOADFCW] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle, Core::OPINDEX_F80LOADFCW).fn);
Info[Core::OPINDEX_F80CVTTO_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4).fn);
Info[Core::OPINDEX_F80CVTTO_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8).fn);
Info[Core::OPINDEX_F80CVT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4).fn);
Info[Core::OPINDEX_F80CVT_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8).fn);
Info[Core::OPINDEX_F80CVTINT_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2).fn);
Info[Core::OPINDEX_F80CVTINT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4).fn);
Info[Core::OPINDEX_F80CVTINT_8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4).fn);
Info[Core::OPINDEX_F80CVTINT_TRUNC8] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8).fn);
Info[Core::OPINDEX_F80CMP_0] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>, Core::OPINDEX_F80CMP_0).fn);
Info[Core::OPINDEX_F80CMP_1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>, Core::OPINDEX_F80CMP_1).fn);
Info[Core::OPINDEX_F80CMP_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>, Core::OPINDEX_F80CMP_2).fn);
Info[Core::OPINDEX_F80CMP_3] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>, Core::OPINDEX_F80CMP_3).fn);
Info[Core::OPINDEX_F80CMP_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>, Core::OPINDEX_F80CMP_4).fn);
Info[Core::OPINDEX_F80CMP_5] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>, Core::OPINDEX_F80CMP_5).fn);
Info[Core::OPINDEX_F80CMP_6] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>, Core::OPINDEX_F80CMP_6).fn);
Info[Core::OPINDEX_F80CMP_7] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>, Core::OPINDEX_F80CMP_7).fn);
Info[Core::OPINDEX_F80CVTTOINT_2] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2).fn);
Info[Core::OPINDEX_F80CVTTOINT_4] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4).fn);
// Unary
Info[Core::OPINDEX_F80ROUND] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ROUND>::handle, Core::OPINDEX_F80ROUND).fn);
Info[Core::OPINDEX_F80F2XM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80F2XM1>::handle, Core::OPINDEX_F80F2XM1).fn);
Info[Core::OPINDEX_F80TAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80TAN>::handle, Core::OPINDEX_F80TAN).fn);
Info[Core::OPINDEX_F80SQRT] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SQRT>::handle, Core::OPINDEX_F80SQRT).fn);
Info[Core::OPINDEX_F80SIN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SIN>::handle, Core::OPINDEX_F80SIN).fn);
Info[Core::OPINDEX_F80COS] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80COS>::handle, Core::OPINDEX_F80COS).fn);
Info[Core::OPINDEX_F80XTRACT_EXP] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_EXP>::handle, Core::OPINDEX_F80XTRACT_EXP).fn);
Info[Core::OPINDEX_F80XTRACT_SIG] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_SIG>::handle, Core::OPINDEX_F80XTRACT_SIG).fn);
Info[Core::OPINDEX_F80BCDSTORE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDSTORE>::handle, Core::OPINDEX_F80BCDSTORE).fn);
Info[Core::OPINDEX_F80BCDLOAD] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDLOAD>::handle, Core::OPINDEX_F80BCDLOAD).fn);
// Binary
Info[Core::OPINDEX_F80ADD] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ADD>::handle, Core::OPINDEX_F80ADD).fn);
Info[Core::OPINDEX_F80SUB] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SUB>::handle, Core::OPINDEX_F80SUB).fn);
Info[Core::OPINDEX_F80MUL] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80MUL>::handle, Core::OPINDEX_F80MUL).fn);
Info[Core::OPINDEX_F80DIV] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80DIV>::handle, Core::OPINDEX_F80DIV).fn);
Info[Core::OPINDEX_F80FYL2X] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FYL2X>::handle, Core::OPINDEX_F80FYL2X).fn);
Info[Core::OPINDEX_F80ATAN] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80ATAN>::handle, Core::OPINDEX_F80ATAN).fn);
Info[Core::OPINDEX_F80FPREM1] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM1>::handle, Core::OPINDEX_F80FPREM1).fn);
Info[Core::OPINDEX_F80FPREM] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80FPREM>::handle, Core::OPINDEX_F80FPREM).fn);
Info[Core::OPINDEX_F80SCALE] = reinterpret_cast<uint64_t>(GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80SCALE>::handle, Core::OPINDEX_F80SCALE).fn);
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle, Core::OPINDEX_F80LOADFCW);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->SrcSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2);
}
return true;
}
case 4: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4);
}
return true;
}
case 8: {
if (Op->Truncate) {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8);
}
else {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8);
}
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = GetFallbackInfo(handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags));
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->SrcSize) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
default:
break;
}
return false;
}
}
@@ -21,234 +21,309 @@
#include <alloca.h>
#include <algorithm>
#include <array>
#include <atomic>
#include <bit>
#include <cmath>
#include <cstddef>
#include <cstdint>
#include <cstdlib>
#include <cstring>
#include <ctime>
#include <limits>
#include <memory>
#include <stddef.h>
#include <stdlib.h>
#include <string.h>
#include <time.h>
namespace FEXCore::CPU {
std::array<InterpreterOps::OpHandler, FEXCore::IR::IROps::OP_LAST + 1> InterpreterOps::OpHandlers;
void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node) {
using OpHandler = void (*)(IR::IROp_Header *IROp, InterpreterOps::IROpData *Data, IR::NodeID Node);
using OpHandlerArray = std::array<OpHandler, IR::IROps::OP_LAST + 1>;
constexpr OpHandlerArray InterpreterOpHandlers = [] {
OpHandlerArray Handlers{};
for (auto& Entry : Handlers) {
Entry = &InterpreterOps::Op_Unhandled;
}
#define REGISTER_OP(op, x) Handlers[IR::IROps::OP_##op] = &InterpreterOps::Op_##x
// ALU ops
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
REGISTER_OP(NEG, Neg);
REGISTER_OP(MUL, Mul);
REGISTER_OP(UMUL, UMul);
REGISTER_OP(DIV, Div);
REGISTER_OP(UDIV, UDiv);
REGISTER_OP(REM, Rem);
REGISTER_OP(UREM, URem);
REGISTER_OP(MULH, MulH);
REGISTER_OP(UMULH, UMulH);
REGISTER_OP(OR, Or);
REGISTER_OP(AND, And);
REGISTER_OP(ANDN, Andn);
REGISTER_OP(XOR, Xor);
REGISTER_OP(LSHL, Lshl);
REGISTER_OP(LSHR, Lshr);
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
REGISTER_OP(LUREM, LURem);
REGISTER_OP(NOT, Not);
REGISTER_OP(POPCOUNT, Popcount);
REGISTER_OP(FINDLSB, FindLSB);
REGISTER_OP(FINDMSB, FindMSB);
REGISTER_OP(FINDTRAILINGZEROS, FindTrailingZeros);
REGISTER_OP(COUNTLEADINGZEROES, CountLeadingZeroes);
REGISTER_OP(REV, Rev);
REGISTER_OP(BFI, Bfi);
REGISTER_OP(BFE, Bfe);
REGISTER_OP(SBFE, Sbfe);
REGISTER_OP(SELECT, Select);
REGISTER_OP(VEXTRACTTOGPR, VExtractToGPR);
REGISTER_OP(FLOAT_TOGPR_ZS, Float_ToGPR_ZS);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
// Atomic ops
REGISTER_OP(CASPAIR, CASPair);
REGISTER_OP(CAS, CAS);
REGISTER_OP(ATOMICADD, AtomicAdd);
REGISTER_OP(ATOMICSUB, AtomicSub);
REGISTER_OP(ATOMICAND, AtomicAnd);
REGISTER_OP(ATOMICOR, AtomicOr);
REGISTER_OP(ATOMICXOR, AtomicXor);
REGISTER_OP(ATOMICSWAP, AtomicSwap);
REGISTER_OP(ATOMICFETCHADD, AtomicFetchAdd);
REGISTER_OP(ATOMICFETCHSUB, AtomicFetchSub);
REGISTER_OP(ATOMICFETCHAND, AtomicFetchAnd);
REGISTER_OP(ATOMICFETCHOR, AtomicFetchOr);
REGISTER_OP(ATOMICFETCHXOR, AtomicFetchXor);
REGISTER_OP(ATOMICFETCHNEG, AtomicFetchNeg);
// Branch ops
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
// Conversion ops
REGISTER_OP(VINSGPR, VInsGPR);
REGISTER_OP(VCASTFROMGPR, VCastFromGPR);
REGISTER_OP(FLOAT_FROMGPR_S, Float_FromGPR_S);
REGISTER_OP(FLOAT_FTOF, Float_FToF);
REGISTER_OP(VECTOR_STOF, Vector_SToF);
REGISTER_OP(VECTOR_FTOZS, Vector_FToZS);
REGISTER_OP(VECTOR_FTOS, Vector_FToS);
REGISTER_OP(VECTOR_FTOF, Vector_FToF);
REGISTER_OP(VECTOR_FTOI, Vector_FToI);
// Flag ops
REGISTER_OP(GETHOSTFLAG, GetHostFlag);
// Memory ops
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
REGISTER_OP(FILLREGISTER, FillRegister);
REGISTER_OP(LOADFLAG, LoadFlag);
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
// Misc ops
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
// Vector ops
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
REGISTER_OP(VAND, VAnd);
REGISTER_OP(VBIC, VBic);
REGISTER_OP(VOR, VOr);
REGISTER_OP(VXOR, VXor);
REGISTER_OP(VADD, VAdd);
REGISTER_OP(VSUB, VSub);
REGISTER_OP(VUQADD, VUQAdd);
REGISTER_OP(VUQSUB, VUQSub);
REGISTER_OP(VSQADD, VSQAdd);
REGISTER_OP(VSQSUB, VSQSub);
REGISTER_OP(VADDP, VAddP);
REGISTER_OP(VADDV, VAddV);
REGISTER_OP(VUMINV, VUMinV);
REGISTER_OP(VURAVG, VURAvg);
REGISTER_OP(VABS, VAbs);
REGISTER_OP(VPOPCOUNT, VPopcount);
REGISTER_OP(VFADD, VFAdd);
REGISTER_OP(VFADDP, VFAddP);
REGISTER_OP(VFSUB, VFSub);
REGISTER_OP(VFMUL, VFMul);
REGISTER_OP(VFDIV, VFDiv);
REGISTER_OP(VFMIN, VFMin);
REGISTER_OP(VFMAX, VFMax);
REGISTER_OP(VFRECP, VFRecp);
REGISTER_OP(VFSQRT, VFSqrt);
REGISTER_OP(VFRSQRT, VFRSqrt);
REGISTER_OP(VNEG, VNeg);
REGISTER_OP(VFNEG, VFNeg);
REGISTER_OP(VNOT, VNot);
REGISTER_OP(VUMIN, VUMin);
REGISTER_OP(VSMIN, VSMin);
REGISTER_OP(VUMAX, VUMax);
REGISTER_OP(VSMAX, VSMax);
REGISTER_OP(VZIP, VZip);
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
REGISTER_OP(VCMPGT, VCMPGT);
REGISTER_OP(VCMPGTZ, VCMPGTZ);
REGISTER_OP(VCMPLTZ, VCMPLTZ);
REGISTER_OP(VFCMPEQ, VFCMPEQ);
REGISTER_OP(VFCMPNEQ, VFCMPNEQ);
REGISTER_OP(VFCMPLT, VFCMPLT);
REGISTER_OP(VFCMPGT, VFCMPGT);
REGISTER_OP(VFCMPLE, VFCMPLE);
REGISTER_OP(VFCMPORD, VFCMPORD);
REGISTER_OP(VFCMPUNO, VFCMPUNO);
REGISTER_OP(VUSHL, VUShl);
REGISTER_OP(VUSHR, VUShr);
REGISTER_OP(VSSHR, VSShr);
REGISTER_OP(VUSHLS, VUShlS);
REGISTER_OP(VUSHRS, VUShrS);
REGISTER_OP(VSSHRS, VSShrS);
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VINSSCALARELEMENT, VInsScalarElement);
REGISTER_OP(VEXTRACTELEMENT, VExtractElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
REGISTER_OP(VUSHRNI, VUShrNI);
REGISTER_OP(VUSHRNI2, VUShrNI2);
REGISTER_OP(VBITCAST, VBitcast);
REGISTER_OP(VSXTL, VSXTL);
REGISTER_OP(VSXTL2, VSXTL2);
REGISTER_OP(VUXTL, VUXTL);
REGISTER_OP(VUXTL2, VUXTL2);
REGISTER_OP(VSQXTN, VSQXTN);
REGISTER_OP(VSQXTN2, VSQXTN2);
REGISTER_OP(VSQXTUN, VSQXTUN);
REGISTER_OP(VSQXTUN2, VSQXTUN2);
REGISTER_OP(VUMUL, VUMul);
REGISTER_OP(VSMUL, VSMul);
REGISTER_OP(VUMULL, VUMull);
REGISTER_OP(VSMULL, VSMull);
REGISTER_OP(VUMULL2, VUMull2);
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
// Encryption ops
REGISTER_OP(VAESIMC, AESImc);
REGISTER_OP(VAESENC, AESEnc);
REGISTER_OP(VAESENCLAST, AESEncLast);
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
// F80 ops
REGISTER_OP(F80LOADFCW, F80LOADFCW);
REGISTER_OP(F80ADD, F80ADD);
REGISTER_OP(F80SUB, F80SUB);
REGISTER_OP(F80MUL, F80MUL);
REGISTER_OP(F80DIV, F80DIV);
REGISTER_OP(F80FYL2X, F80FYL2X);
REGISTER_OP(F80ATAN, F80ATAN);
REGISTER_OP(F80FPREM1, F80FPREM1);
REGISTER_OP(F80FPREM, F80FPREM);
REGISTER_OP(F80SCALE, F80SCALE);
REGISTER_OP(F80CVT, F80CVT);
REGISTER_OP(F80CVTINT, F80CVTINT);
REGISTER_OP(F80CVTTO, F80CVTTO);
REGISTER_OP(F80CVTTOINT, F80CVTTOINT);
REGISTER_OP(F80ROUND, F80ROUND);
REGISTER_OP(F80F2XM1, F80F2XM1);
REGISTER_OP(F80TAN, F80TAN);
REGISTER_OP(F80SQRT, F80SQRT);
REGISTER_OP(F80SIN, F80SIN);
REGISTER_OP(F80COS, F80COS);
REGISTER_OP(F80XTRACT_EXP, F80XTRACT_EXP);
REGISTER_OP(F80XTRACT_SIG, F80XTRACT_SIG);
REGISTER_OP(F80CMP, F80CMP);
REGISTER_OP(F80BCDLOAD, F80BCDLOAD);
REGISTER_OP(F80BCDSTORE, F80BCDSTORE);
return Handlers;
}();
void InterpreterOps::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
LOGMAN_MSG_A_FMT("Unhandled IR Op: {}", FEXCore::IR::GetName(IROp->Op));
}
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node) {
}
template<typename R, typename... Args>
FallbackInfo GetFallbackInfo(R(*fn)(Args...)) {
return {FABI_UNKNOWN, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(float)) {
return {FABI_F80_F32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(double)) {
return {FABI_F80_F64, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int16_t)) {
return {FABI_F80_I16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(void(*fn)(uint16_t)) {
return {FABI_VOID_U16, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(int32_t)) {
return {FABI_F80_I32, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(float(*fn)(X80SoftFloat)) {
return {FABI_F32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(X80SoftFloat)) {
return {FABI_F64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int16_t(*fn)(X80SoftFloat)) {
return {FABI_I16_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int32_t(*fn)(X80SoftFloat)) {
return {FABI_I32_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(int64_t(*fn)(X80SoftFloat)) {
return {FABI_I64_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(uint64_t(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_I64_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat)) {
return {FABI_F80_F80, (void*)fn};
}
template<>
FallbackInfo GetFallbackInfo(X80SoftFloat(*fn)(X80SoftFloat, X80SoftFloat)) {
return {FABI_F80_F80_F80, (void*)fn};
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80LOADFCW>::handle);
return true;
}
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->Size) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4);
return true;
}
case 8: {
*Info = GetFallbackInfo(Op->Truncate ? &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t : &FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
decltype(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>) handlers[] = {
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7> };
*Info = GetFallbackInfo(handlers[Op->Flags]);
return true;
}
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->Size) {
case 2: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
return true;
}
case 4: {
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle); \
return true; \
}
// Unary
COMMON_X87_OP(ROUND)
COMMON_X87_OP(F2XM1)
COMMON_X87_OP(TAN)
COMMON_X87_OP(SQRT)
COMMON_X87_OP(SIN)
COMMON_X87_OP(COS)
COMMON_X87_OP(XTRACT_EXP)
COMMON_X87_OP(XTRACT_SIG)
COMMON_X87_OP(BCDSTORE)
COMMON_X87_OP(BCDLOAD)
// Binary
COMMON_X87_OP(ADD)
COMMON_X87_OP(SUB)
COMMON_X87_OP(MUL)
COMMON_X87_OP(DIV)
COMMON_X87_OP(FYL2X)
COMMON_X87_OP(ATAN)
COMMON_X87_OP(FPREM1)
COMMON_X87_OP(FPREM)
COMMON_X87_OP(SCALE)
default:
break;
}
return false;
void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node) {
}
void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData) {
volatile void *StackEntry = alloca(0);
// Debug data is only passed in debug builds
#ifndef NDEBUG
// TODO: should be moved to an IR Op
Thread->Stats.InstructionsExecuted.fetch_add(DebugData->GuestInstructionCount);
#endif
uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
@@ -280,11 +355,11 @@ void InterpreterOps::InterpretIR(FEXCore::Core::InternalThreadState *Thread, uin
auto CodeLast = CurrentIR->at(BlockIROp->Last);
for (auto [CodeNode, IROp] : CurrentIR->GetCode(BlockNode)) {
uint32_t ID = CurrentIR->GetID(CodeNode);
uint32_t Op = IROp->Op;
const auto ID = CurrentIR->GetID(CodeNode);
const uint32_t Op = IROp->Op;
// Execute handler
OpHandler Handler = InterpreterOps::OpHandlers[Op];
OpHandler Handler = InterpreterOpHandlers[Op];
Handler(IROp, &OpData, ID);
@@ -1,6 +1,7 @@
#pragma once
#include <stdint.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
@@ -38,26 +39,16 @@ namespace FEXCore::CPU {
struct FallbackInfo {
FallbackABI ABI;
void *fn;
FEXCore::Core::FallbackHandlerIndex HandlerIndex;
};
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, uint64_t Entry, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
static void RegisterALUHandlers();
static void RegisterAtomicHandlers();
static void RegisterBranchHandlers();
static void RegisterConversionHandlers();
static void RegisterFlagHandlers();
static void RegisterMemoryHandlers();
static void RegisterMiscHandlers();
static void RegisterMoveHandlers();
static void RegisterVectorHandlers();
static void RegisterEncryptionHandlers();
static void RegisterF80Handlers();
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
uint64_t CurrentEntry{};
@@ -72,10 +63,7 @@ namespace FEXCore::CPU {
IR::NodeIterator BlockIterator{0, 0};
};
using OpHandler = std::function<void(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)>;
static std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers;
#define DEF_OP(x) static void Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) static void Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -111,6 +99,8 @@ namespace FEXCore::CPU {
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -152,13 +142,13 @@ namespace FEXCore::CPU {
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
DEF_OP(CondJump);
DEF_OP(Syscall);
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
@@ -194,6 +184,7 @@ namespace FEXCore::CPU {
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -204,6 +195,8 @@ namespace FEXCore::CPU {
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -213,8 +206,6 @@ namespace FEXCore::CPU {
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
@@ -300,6 +291,7 @@ namespace FEXCore::CPU {
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -308,6 +300,7 @@ namespace FEXCore::CPU {
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
///< F80 ops
DEF_OP(F80LOADFCW);
@@ -407,4 +400,4 @@ namespace FEXCore::CPU {
}
};
};
} // namespace FEXCore::CPU
@@ -22,7 +22,7 @@ static inline void CacheLineFlush(char *Addr) {
#endif
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
@@ -264,26 +264,22 @@ DEF_OP(CacheLineClear) {
CacheLineFlush(MemData);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
uintptr_t MemData = *GetSrc<uintptr_t*>(Data->SSAData, Op->Addr);
// Force cacheline alignment
MemData = MemData & ~(CPUIDEmu::CACHELINE_SIZE - 1);
using DataType = uint64_t;
DataType *MemData64 = reinterpret_cast<DataType*>(MemData);
// 64-byte cache line zero
for (size_t i = 0; i < (CPUIDEmu::CACHELINE_SIZE / sizeof(DataType)); ++i) {
MemData64[i] = 0;
}
}
#undef DEF_OP
void InterpreterOps::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
REGISTER_OP(FILLREGISTER, FillRegister);
REGISTER_OP(LOADFLAG, LoadFlag);
REGISTER_OP(STOREFLAG, StoreFlag);
REGISTER_OP(LOADMEM, LoadMem);
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -8,10 +8,13 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <FEXHeaderUtils/Syscalls.h>
#include <cstdint>
#ifdef _M_X86_64
#include <xmmintrin.h>
#endif
#include <sys/random.h>
namespace FEXCore::CPU {
[[noreturn]]
@@ -22,7 +25,7 @@ static void StopThread(FEXCore::Core::InternalThreadState *Thread) {
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
@@ -42,9 +45,12 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 4: // HLT
case FEXCore::IR::Break_Halt: // HLT
StopThread(Data->State);
break;
case FEXCore::IR::Break_InvalidInstruction:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
break;
default: LOGMAN_MSG_A_FMT("Unknown Break Reason: {}", Op->Reason); break;
}
}
@@ -137,22 +143,20 @@ DEF_OP(Print) {
LOGMAN_MSG_A_FMT("Unknown value size: {}", OpSize);
}
DEF_OP(ProcessorID) {
uint32_t CPU, CPUNode;
FHU::Syscalls::getcpu(&CPU, &CPUNode);
GD = (CPUNode << 12) | CPU;
}
DEF_OP(RDRAND) {
// We are ignoring Op->GetReseeded in the interpreter
uint64_t *DstPtr = GetDest<uint64_t*>(Data->SSAData, Node);
ssize_t Result = ::getrandom(&DstPtr[0], 8, 0);
// Second result is if we managed to read a valid random number or not
DstPtr[1] = Result == 8 ? 1 : 0;
}
#undef DEF_OP
void InterpreterOps::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(DUMMY, NoOp);
REGISTER_OP(IRHEADER, NoOp);
REGISTER_OP(CODEBLOCK, NoOp);
REGISTER_OP(BEGINBLOCK, NoOp);
REGISTER_OP(ENDBLOCK, NoOp);
REGISTER_OP(FENCE, Fence);
REGISTER_OP(BREAK, Break);
REGISTER_OP(PHI, NoOp);
REGISTER_OP(PHIVALUE, NoOp);
REGISTER_OP(PRINT, Print);
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -11,7 +11,7 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
uintptr_t Src = GetSrc<uintptr_t>(Data->SSAData, Op->Header.Args[0]);
@@ -26,8 +26,8 @@ DEF_OP(CreateElementPair) {
uint8_t *Dst = GetDest<uint8_t*>(Data->SSAData, Node);
memcpy(Dst, Src_Lower, Op->Header.Size);
memcpy(Dst + Op->Header.Size, Src_Upper, Op->Header.Size);
memcpy(Dst, Src_Lower, IROp->ElementSize);
memcpy(Dst + IROp->ElementSize, Src_Upper, IROp->ElementSize);
}
DEF_OP(Mov) {
@@ -38,13 +38,5 @@ DEF_OP(Mov) {
}
#undef DEF_OP
void InterpreterOps::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
#undef REGISTER_OP
}
}
} // namespace FEXCore::CPU
@@ -7,11 +7,13 @@ $end_info$
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/Interpreter/InterpreterDefines.h"
#include <FEXCore/Utils/BitUtils.h>
#include <bit>
#include <cstdint>
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(FEXCore::IR::IROp_Header *IROp, IROpData *Data, uint32_t Node)
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(VectorZero) {
uint8_t OpSize = IROp->Size;
memset(GDP, 0, OpSize);
@@ -37,39 +39,6 @@ DEF_OP(VectorImm) {
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(CreateVector2) {
auto Op = IROp->C<IR::IROp_CreateVector2>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize <= 16, "Can't handle a vector of size: {}", OpSize);
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Tmp[16];
uint8_t ElementSize = OpSize / 2;
#define CREATE_VECTOR(elementsize, type) \
case elementsize: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
Dst_d[0] = *Src1_d; \
Dst_d[1] = *Src2_d; \
break; \
}
switch (ElementSize) {
CREATE_VECTOR(1, uint8_t)
CREATE_VECTOR(2, uint16_t)
CREATE_VECTOR(4, uint32_t)
CREATE_VECTOR(8, uint64_t)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize); break;
}
#undef CREATE_VECTOR
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1401,10 +1370,9 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
uint8_t OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
LOGMAN_THROW_A_FMT(OpSize <= 16, "OpSize is too large for VExtractElement: {}", OpSize);
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractElement: {}", IROp->Size);
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
@@ -1930,102 +1898,39 @@ DEF_OP(VTBL1) {
memcpy(GDP, Tmp, OpSize);
}
#undef DEF_OP
void InterpreterOps::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &InterpreterOps::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
REGISTER_OP(VAND, VAnd);
REGISTER_OP(VBIC, VBic);
REGISTER_OP(VOR, VOr);
REGISTER_OP(VXOR, VXor);
REGISTER_OP(VADD, VAdd);
REGISTER_OP(VSUB, VSub);
REGISTER_OP(VUQADD, VUQAdd);
REGISTER_OP(VUQSUB, VUQSub);
REGISTER_OP(VSQADD, VSQAdd);
REGISTER_OP(VSQSUB, VSQSub);
REGISTER_OP(VADDP, VAddP);
REGISTER_OP(VADDV, VAddV);
REGISTER_OP(VUMINV, VUMinV);
REGISTER_OP(VURAVG, VURAvg);
REGISTER_OP(VABS, VAbs);
REGISTER_OP(VPOPCOUNT, VPopcount);
REGISTER_OP(VFADD, VFAdd);
REGISTER_OP(VFADDP, VFAddP);
REGISTER_OP(VFSUB, VFSub);
REGISTER_OP(VFMUL, VFMul);
REGISTER_OP(VFDIV, VFDiv);
REGISTER_OP(VFMIN, VFMin);
REGISTER_OP(VFMAX, VFMax);
REGISTER_OP(VFRECP, VFRecp);
REGISTER_OP(VFSQRT, VFSqrt);
REGISTER_OP(VFRSQRT, VFRSqrt);
REGISTER_OP(VNEG, VNeg);
REGISTER_OP(VFNEG, VFNeg);
REGISTER_OP(VNOT, VNot);
REGISTER_OP(VUMIN, VUMin);
REGISTER_OP(VSMIN, VSMin);
REGISTER_OP(VUMAX, VUMax);
REGISTER_OP(VSMAX, VSMax);
REGISTER_OP(VZIP, VZip);
REGISTER_OP(VZIP2, VZip);
REGISTER_OP(VUNZIP, VUnZip);
REGISTER_OP(VUNZIP2, VUnZip);
REGISTER_OP(VBSL, VBSL);
REGISTER_OP(VCMPEQ, VCMPEQ);
REGISTER_OP(VCMPEQZ, VCMPEQZ);
REGISTER_OP(VCMPGT, VCMPGT);
REGISTER_OP(VCMPGTZ, VCMPGTZ);
REGISTER_OP(VCMPLTZ, VCMPLTZ);
REGISTER_OP(VFCMPEQ, VFCMPEQ);
REGISTER_OP(VFCMPNEQ, VFCMPNEQ);
REGISTER_OP(VFCMPLT, VFCMPLT);
REGISTER_OP(VFCMPGT, VFCMPGT);
REGISTER_OP(VFCMPLE, VFCMPLE);
REGISTER_OP(VFCMPORD, VFCMPORD);
REGISTER_OP(VFCMPUNO, VFCMPUNO);
REGISTER_OP(VUSHL, VUShl);
REGISTER_OP(VUSHR, VUShr);
REGISTER_OP(VSSHR, VSShr);
REGISTER_OP(VUSHLS, VUShlS);
REGISTER_OP(VUSHRS, VUShrS);
REGISTER_OP(VSSHRS, VSShrS);
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VINSSCALARELEMENT, VInsScalarElement);
REGISTER_OP(VEXTRACTELEMENT, VExtractElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
REGISTER_OP(VUSHRNI, VUShrNI);
REGISTER_OP(VUSHRNI2, VUShrNI2);
REGISTER_OP(VBITCAST, VBitcast);
REGISTER_OP(VSXTL, VSXTL);
REGISTER_OP(VSXTL2, VSXTL2);
REGISTER_OP(VUXTL, VUXTL);
REGISTER_OP(VUXTL2, VUXTL2);
REGISTER_OP(VSQXTN, VSQXTN);
REGISTER_OP(VSQXTN2, VSQXTN2);
REGISTER_OP(VSQXTUN, VSQXTUN);
REGISTER_OP(VSQXTUN2, VSQXTUN2);
REGISTER_OP(VUMUL, VUMul);
REGISTER_OP(VSMUL, VSMul);
REGISTER_OP(VUMULL, VUMull);
REGISTER_OP(VSMULL, VSMull);
REGISTER_OP(VUMULL2, VUMull2);
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
#undef REGISTER_OP
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VRev64>();
uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16];
uint8_t Elements = OpSize / 8;
// The element working size is always 64-bit
// The defined element size in the op is the operating size of the element swapping
auto Func8 = [](auto a) { return BSwap64(a); };
auto Func16 = [](auto a) {
return (a >> 48) | // Element[3] -> Element[0]
((a >> 16) & 0xFFFF'0000U) | // Element[2] -> Element[1]
((a << 16) & 0xFFFF'0000'0000ULL) | // Element[1] -> Element[2]
(a << 48); // Element[0] -> Element[3]
};
auto Func32 = [](auto a) {
return (a >> 32) | (a << 32);
};
switch (Op->Header.ElementSize) {
DO_VECTOR_1SRC_OP(1, uint64_t, Func8)
DO_VECTOR_1SRC_OP(2, uint64_t, Func16)
DO_VECTOR_1SRC_OP(4, uint64_t, Func32)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
memcpy(GDP, Tmp, Op->Header.Size);
}
#undef DEF_OP
} // namespace FEXCore::CPU
+258 -85
View File
@@ -8,37 +8,17 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
auto Dst = GetSrcPair<RA_32>(Node);
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
@@ -46,7 +26,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -61,7 +41,13 @@ DEF_OP(EntrypointOffset) {
auto Constant = Entry + Op->Offset;
auto Dst = GetReg<RA_64>(Node);
LoadConstant(Dst, Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
LoadConstant(Dst, Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -80,8 +66,6 @@ DEF_OP(CycleCounter) {
#endif
}
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
DEF_OP(Add) {
auto Op = IROp->C<IR::IROp_Add>();
uint8_t OpSize = IROp->Size;
@@ -511,6 +495,125 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const Register Input = GRS(Op->Args(0).ID());
const Register Mask = GRS(Op->Args(1).ID());
const Register Dest = GRS(Node);
const Register ShiftedBitReg = OpSize <= 4 ? TMP1.W() : TMP1;
const Register BitReg = OpSize <= 4 ? TMP2.W() : TMP2;
const Register SubMaskReg = OpSize <= 4 ? TMP3.W() : TMP3;
const Register IndexReg = OpSize <= 4 ? TMP4.W() : TMP4;
const Register SizedZero = OpSize <= 4 ? Register{wzr} : Register{xzr};
const Register InputReg = OpSize <= 4 ? SRA64[0].W() : SRA64[0];
const Register MaskReg = OpSize <= 4 ? SRA64[1].W() : SRA64[1];
const Register DestReg = OpSize <= 4 ? SRA64[2].W() : SRA64[2];
const auto SpillCode = 1U << InputReg.GetCode() |
1U << MaskReg.GetCode() |
1U << DestReg.GetCode();
aarch64::Label EarlyExit;
aarch64::Label NextBit;
aarch64::Label Done;
cbz(Mask, &EarlyExit);
mov(IndexReg, SizedZero);
// We sadly need to spill regs for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, SpillCode);
mov(InputReg, Input);
mov(MaskReg, Mask);
mov(DestReg, SizedZero);
// Main loop
bind(&NextBit);
rbit(ShiftedBitReg, MaskReg);
clz(ShiftedBitReg, ShiftedBitReg);
lsrv(BitReg, InputReg, IndexReg);
and_(BitReg, BitReg, 1);
sub(SubMaskReg, MaskReg, 1);
add(IndexReg, IndexReg, 1);
ands(MaskReg, MaskReg, SubMaskReg);
lslv(ShiftedBitReg, BitReg, ShiftedBitReg);
orr(DestReg, DestReg, ShiftedBitReg);
b(&NextBit, Condition::ne);
// Store result in a temp so it doesn't get clobbered.
// and restore it after the re-fill below.
mov(IndexReg, DestReg);
// Restore our registers before leaving
// TODO: Also remove along with above TODO.
FillStaticRegs(false, SpillCode);
mov(Dest, IndexReg);
b(&Done);
// Early exit
bind(&EarlyExit);
mov(Dest, SizedZero);
// All done with nothing to do.
bind(&Done);
}
DEF_OP(PExt) {
auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const Register Input = GRS(Op->Args(0).ID());
const Register Mask = GRS(Op->Args(1).ID());
const Register Dest = GRS(Node);
const Register MaskReg = OpSize <= 4 ? TMP1.W() : TMP1;
const Register BitReg = OpSize <= 4 ? TMP2.W() : TMP2;
const Register SubMaskReg = OpSize <= 4 ? TMP3.W() : TMP3;
const Register Offset = OpSize <= 4 ? TMP4.W() : TMP4;
const Register SizedZero = OpSize <= 4 ? Register{wzr} : Register{xzr};
aarch64::Label EarlyExit;
aarch64::Label NextBit;
aarch64::Label Done;
cbz(Mask, &EarlyExit);
mov(MaskReg, Mask);
mov(Offset, SizedZero);
// We sadly need to spill a reg for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(false, 1U << Mask.GetCode());
mov(Mask, SizedZero);
// Main loop
bind(&NextBit);
rbit(BitReg, MaskReg);
clz(BitReg, BitReg);
sub(SubMaskReg, MaskReg, 1);
ands(MaskReg, SubMaskReg, MaskReg);
lsrv(BitReg, Input, BitReg);
and_(BitReg, BitReg, 1);
lslv(BitReg, BitReg, Offset);
add(Offset, Offset, 1);
orr(Mask, BitReg, Mask);
b(&NextBit, Condition::ne);
mov(Dest, Mask);
// Restore our mask register before leaving
// TODO: Also remove along with above TODO.
FillStaticRegs(false, 1U << Mask.GetCode());
b(&Done);
// Early exit
bind(&EarlyExit);
mov(Dest, SizedZero);
// All done with nothing to do.
bind(&Done);
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -533,23 +636,42 @@ DEF_OP(LDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LDIV Size: {}", Size); break;
@@ -576,23 +698,38 @@ DEF_OP(LUDiv) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LUDIV));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
udiv(GetReg<RA_64>(Node), Lower64Bit, Divisor);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUDIV Size: {}", Size); break;
@@ -629,23 +766,42 @@ DEF_OP(LRem) {
break;
}
case 8: {
PushDynamicRegsAndLR();
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
sbfx(TMP1, Lower64Bit, 63, 1);
eor(TMP1, TMP1, Upper64Bit);
LoadConstant(x3, reinterpret_cast<uint64_t>(LREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
// If the sign bit matches then the result is zero
cbz(TMP1, &Only64Bit);
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Skip 64-bit path
b(&LongDIVRet);
}
bind(&Only64Bit);
// 64-Bit only
{
sdiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LREM Size: {}", Size); break;
@@ -678,24 +834,39 @@ DEF_OP(LURem) {
break;
}
case 8: {
auto Upper64Bit = GetReg<RA_64>(Op->Header.Args[1].ID());
auto Lower64Bit = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Divisor = GetReg<RA_64>(Op->Header.Args[2].ID());
Label Only64Bit{};
Label LongDIVRet{};
PushDynamicRegsAndLR();
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
cbz(Upper64Bit, &Only64Bit);
mov(x0, GetReg<RA_64>(Op->Header.Args[1].ID()));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[2].ID()));
// Long divide
{
mov(x0, Upper64Bit);
mov(x1, Lower64Bit);
mov(x2, Divisor);
LoadConstant(x3, reinterpret_cast<uint64_t>(LUREM));
SpillStaticRegs();
blr(x3);
FillStaticRegs();
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREMHandler)));
blr(x3);
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Skip 64-bit path
b(&LongDIVRet);
}
// Result is now in x0
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
bind(&Only64Bit);
// 64-Bit only
{
udiv(TMP1, Lower64Bit, Divisor);
msub(GetReg<RA_64>(Node), TMP1, Divisor, Lower64Bit);
}
bind(&LongDIVRet);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown LUREM Size: {}", OpSize); break;
@@ -916,7 +1087,7 @@ Condition MapSelectCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FLU: return Condition::lt;
case FEXCore::IR::COND_FGE: return Condition::ge;
case FEXCore::IR::COND_FLEU:return Condition::le;
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FGT: return Condition::gt;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:
@@ -966,16 +1137,16 @@ DEF_OP(VExtractToGPR) {
uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Index);
break;
case 2:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V8H(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V8H(), Op->Index);
break;
case 4:
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V4S(), Op->Idx);
umov(GetReg<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()).V4S(), Op->Index);
break;
case 8:
umov(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Idx);
umov(GetReg<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Index);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", OpSize);
}
@@ -1099,6 +1270,8 @@ void Arm64JITCore::RegisterALUHandlers() {
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
+109 -113
View File
@@ -9,21 +9,20 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
uint8_t OpSize = IROp->Size;
// Size is the size of each pair element
auto Dst = GetSrcPair<RA_64>(Node);
auto Expected = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetSrcPair<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetSrcPair<RA_64>(Op->Expected.ID());
auto Desired = GetSrcPair<RA_64>(Op->Desired.ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP3, Expected.first);
mov(TMP4, Expected.second);
switch (OpSize) {
switch (IROp->ElementSize) {
case 4:
caspal(TMP3.W(), TMP4.W(), Desired.first.W(), Desired.second.W(), MemOperand(MemSrc));
mov(Dst.first.W(), TMP3.W());
@@ -34,11 +33,11 @@ DEF_OP(CASPair) {
mov(Dst.first, TMP3);
mov(Dst.second, TMP4);
break;
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
else {
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
aarch64::Label LoopTop;
aarch64::Label LoopNotExpected;
@@ -91,7 +90,7 @@ DEF_OP(CASPair) {
bind(&LoopExpected);
break;
}
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
}
@@ -99,18 +98,15 @@ DEF_OP(CASPair) {
DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Expected
// Args[1]: Desired
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
auto Expected = GetReg<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetReg<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetReg<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetReg<RA_64>(Op->Expected.ID());
auto Desired = GetReg<RA_64>(Op->Desired.ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, Expected);
switch (OpSize) {
case 1: casalb(TMP2.W(), Desired.W(), MemOperand(MemSrc)); break;
@@ -216,14 +212,14 @@ DEF_OP(CAS) {
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: staddlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: staddlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: staddlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: staddl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: staddl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -234,7 +230,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -243,7 +239,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -252,7 +248,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -261,7 +257,7 @@ DEF_OP(AtomicAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
add(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
add(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -274,10 +270,10 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: staddlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: staddlh(TMP2.W(), MemOperand(MemSrc)); break;
@@ -293,7 +289,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -302,7 +298,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -311,7 +307,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -320,7 +316,7 @@ DEF_OP(AtomicSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
sub(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
sub(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -333,10 +329,10 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: stclrlb(TMP2.W(), MemOperand(MemSrc)); break;
case 2: stclrlh(TMP2.W(), MemOperand(MemSrc)); break;
@@ -352,7 +348,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -361,7 +357,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -370,7 +366,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -379,7 +375,7 @@ DEF_OP(AtomicAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
and_(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -392,14 +388,14 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: stsetlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: stsetlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: stsetlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: stsetl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: stsetl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -410,7 +406,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -419,7 +415,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -428,7 +424,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -437,7 +433,7 @@ DEF_OP(AtomicOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
orr(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
orr(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -450,14 +446,14 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: steorlb(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc)); break;
case 1: steorlb(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 2: steorlh(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 4: steorl(GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc)); break;
case 8: steorl(GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -468,7 +464,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -477,7 +473,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -486,7 +482,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP2.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP2.W(), TMP2.W(), MemOperand(MemSrc));
cbnz(TMP2.W(), &LoopTop);
break;
@@ -495,7 +491,7 @@ DEF_OP(AtomicXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
eor(TMP2, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
eor(TMP2, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP2, TMP2, MemOperand(MemSrc));
cbnz(TMP2, &LoopTop);
break;
@@ -508,10 +504,10 @@ DEF_OP(AtomicXor) {
DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
if (CTX->HostFeatures.SupportsAtomics) {
mov(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: swplb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: swplh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -527,7 +523,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
stlxrb(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxrb(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
uxtb(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -536,7 +532,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
stlxrh(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxrh(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
uxtw(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -545,7 +541,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
stlxr(TMP4.W(), GetReg<RA_32>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxr(TMP4.W(), GetReg<RA_32>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
break;
@@ -554,7 +550,7 @@ DEF_OP(AtomicSwap) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
stlxr(TMP4, GetReg<RA_64>(Op->Header.Args[1].ID()), MemOperand(MemSrc));
stlxr(TMP4, GetReg<RA_64>(Op->Value.ID()), MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2.X());
break;
@@ -566,14 +562,14 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldaddalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldaddalb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldaddal(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldaddal(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -584,7 +580,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -594,7 +590,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -604,7 +600,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
add(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -614,7 +610,7 @@ DEF_OP(AtomicFetchAdd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
add(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
add(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -627,10 +623,10 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
if (CTX->HostFeatures.SupportsAtomics) {
neg(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: ldaddalb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldaddalh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -646,7 +642,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -656,7 +652,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -666,7 +662,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
sub(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -676,7 +672,7 @@ DEF_OP(AtomicFetchSub) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
sub(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
sub(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -689,10 +685,10 @@ DEF_OP(AtomicFetchSub) {
DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
if (CTX->HostFeatures.SupportsAtomics) {
mvn(TMP2, GetReg<RA_64>(Op->Value.ID()));
switch (IROp->Size) {
case 1: ldclralb(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldclralh(TMP2.W(), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
@@ -708,7 +704,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -718,7 +714,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -728,7 +724,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
and_(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -738,7 +734,7 @@ DEF_OP(AtomicFetchAnd) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
and_(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
and_(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -751,14 +747,14 @@ DEF_OP(AtomicFetchAnd) {
DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldsetalb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldsetalb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldsetalh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldsetal(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldsetal(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -769,7 +765,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -779,7 +775,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -789,7 +785,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
orr(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -799,7 +795,7 @@ DEF_OP(AtomicFetchOr) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
orr(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
orr(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -812,14 +808,14 @@ DEF_OP(AtomicFetchOr) {
DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
if (SupportsAtomics) {
if (CTX->HostFeatures.SupportsAtomics) {
switch (IROp->Size) {
case 1: ldeoralb(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Header.Args[1].ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Header.Args[1].ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
case 1: ldeoralb(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 2: ldeoralh(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 4: ldeoral(GetReg<RA_32>(Op->Value.ID()), GetReg<RA_32>(Node), MemOperand(MemSrc)); break;
case 8: ldeoral(GetReg<RA_64>(Op->Value.ID()), GetReg<RA_64>(Node), MemOperand(MemSrc)); break;
default: LOGMAN_MSG_A_FMT("Unhandled Atomic size: {}", IROp->Size);
}
}
@@ -830,7 +826,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrb(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrb(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -840,7 +836,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxrh(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxrh(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -850,7 +846,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2.W(), MemOperand(MemSrc));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Header.Args[1].ID()));
eor(TMP3.W(), TMP2.W(), GetReg<RA_32>(Op->Value.ID()));
stlxr(TMP4.W(), TMP3.W(), MemOperand(MemSrc));
cbnz(TMP4.W(), &LoopTop);
mov(GetReg<RA_32>(Node), TMP2.W());
@@ -860,7 +856,7 @@ DEF_OP(AtomicFetchXor) {
aarch64::Label LoopTop;
bind(&LoopTop);
ldaxr(TMP2, MemOperand(MemSrc));
eor(TMP3, TMP2, GetReg<RA_64>(Op->Header.Args[1].ID()));
eor(TMP3, TMP2, GetReg<RA_64>(Op->Value.ID()));
stlxr(TMP4, TMP3, MemOperand(MemSrc));
cbnz(TMP4, &LoopTop);
mov(GetReg<RA_64>(Node), TMP2);
@@ -873,7 +869,7 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
auto MemSrc = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemSrc = GetReg<RA_64>(Op->Addr.ID());
// TMP2-TMP3
switch (IROp->Size) {
+195 -67
View File
@@ -4,6 +4,7 @@ tags: backend|arm64
$end_info$
*/
#include "FEXCore/IR/IR.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -11,12 +12,13 @@ $end_info$
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/MathUtils.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -25,17 +27,13 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// First we must reset the stack
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
LoadConstant(x0, ThreadSharedData.SignalReturnInstruction);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalReturnHandler)));
br(x0);
}
@@ -48,7 +46,7 @@ DEF_OP(CallbackReturn) {
ResetStack();
// We can now lower the ref counter again
LoadConstant(x0, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SignalHandlerRefCountPointer)));
ldr(w2, MemOperand(x0));
sub(w2, w2, 1);
str(w2, MemOperand(x0));
@@ -87,7 +85,7 @@ DEF_OP(ExitFunction) {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// L1 Cache
LoadConstant(x0, ThreadState->LookupCache->GetL1Pointer());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.L1Pointer)));
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
@@ -98,30 +96,23 @@ DEF_OP(ExitFunction) {
br(x1);
bind(&FullLookup);
LoadConstant(TMP1, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.DispatcherLoopTop)));
str(RipReg, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
br(TMP1);
}
}
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
Label *TargetLabel;
auto IsTarget = JumpTargets.find(Op->Header.Args[0].ID());
if (IsTarget == JumpTargets.end()) {
TargetLabel = &JumpTargets.try_emplace(Op->Header.Args[0].ID()).first->second;
}
else {
TargetLabel = &IsTarget->second;
}
PendingTargetLabel = TargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
#define GRFCMP(Node) (Op->CompareSize == 4 ? GetDst(Node).S() : GetDst(Node).D())
Condition MapBranchCC(IR::CondClassType Cond) {
static Condition MapBranchCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return Condition::eq;
case FEXCore::IR::COND_NEQ: return Condition::ne;
@@ -136,7 +127,7 @@ Condition MapBranchCC(IR::CondClassType Cond) {
case FEXCore::IR::COND_FLU: return Condition::lt;
case FEXCore::IR::COND_FGE: return Condition::ge;
case FEXCore::IR::COND_FLEU:return Condition::le;
case FEXCore::IR::COND_FGT: return Condition::hi;
case FEXCore::IR::COND_FGT: return Condition::gt;
case FEXCore::IR::COND_FU: return Condition::vs;
case FEXCore::IR::COND_FNU: return Condition::vc;
case FEXCore::IR::COND_VS:
@@ -153,22 +144,10 @@ Condition MapBranchCC(IR::CondClassType Cond) {
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
Label *TrueTargetLabel;
Label *FalseTargetLabel;
auto TrueIter = JumpTargets.find(Op->TrueBlock.ID());
auto FalseIter = JumpTargets.find(Op->FalseBlock.ID());
if (TrueIter == JumpTargets.end()) {
TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
}
else {
TrueTargetLabel = &TrueIter->second;
}
Label *TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
uint64_t Const;
bool isConst = IsInlineConstant(Op->Cmp2, &Const);
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
if (isConst && Const == 0 && Op->Cond.Val == FEXCore::IR::COND_EQ) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
@@ -178,10 +157,11 @@ DEF_OP(CondJump) {
cbnz(GRCMP(Op->Cmp1.ID()), TrueTargetLabel);
} else {
if (IsGPR(Op->Cmp1.ID())) {
if (isConst)
if (isConst) {
cmp(GRCMP(Op->Cmp1.ID()), Const);
else
} else {
cmp(GRCMP(Op->Cmp1.ID()), GRCMP(Op->Cmp2.ID()));
}
} else if (IsFPR(Op->Cmp1.ID())) {
fcmp(GRFCMP(Op->Cmp1.ID()), GRFCMP(Op->Cmp2.ID()));
} else {
@@ -191,13 +171,7 @@ DEF_OP(CondJump) {
b(TrueTargetLabel, MapBranchCC(Op->Cond));
}
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
else {
FalseTargetLabel = &FalseIter->second;
}
PendingTargetLabel = FalseTargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
DEF_OP(Syscall) {
@@ -207,8 +181,16 @@ DEF_OP(Syscall) {
// X1: ThreadState
// X2: Pointer to SyscallArguments
FEXCore::IR::SyscallFlags Flags = Op->Flags;
PushDynamicRegsAndLR();
SpillStaticRegs();
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) {
SpillStaticRegs();
}
else {
// Need to spill all caller saved registers still
SpillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(sp, sp, SPOffset);
@@ -217,22 +199,178 @@ DEF_OP(Syscall) {
str(GetReg<RA_64>(Op->Header.Args[i].ID()), MemOperand(sp, i * 8));
}
LoadConstant(x0, reinterpret_cast<uint64_t>(CTX->SyscallHandler));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
LoadConstant(x3, reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall));
blr(x3);
add(sp, sp, SPOffset);
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs();
if ((Flags & FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY) != FEXCore::IR::SyscallFlags::NOSYNCSTATEONENTRY &&
(Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
FillStaticRegs();
}
else {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, CALLER_GPR_MASK, CALLER_FPR_MASK);
}
PopDynamicRegsAndLR();
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Move result to its destination register
mov(GetReg<RA_64>(Node), x0);
}
}
DEF_OP(InlineSyscall) {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
// Arguments are passed as follows:
// X8: SyscallNumber - RA INTERSECT
// X0: Arg0 & Return
// X1: Arg1
// X2: Arg2
// X3: Arg3
// X4: Arg4 - RA INTERSECT
// X5: Arg5 - RA INTERSECT
// X6: Arg6 - Doesn't exist in x86-64 land. RA INTERSECT
// One argument is removed from the SyscallArguments::MAX_ARGS since the first argument was syscall number
const static std::array<vixl::aarch64::Register, FEXCore::HLE::SyscallArguments::MAX_ARGS-1> RegArgs = {{
x0, x1, x2, x3, x4, x5
}};
bool Intersects{};
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
std::vector<vixl::aarch64::Register> IntersectRegs(FEXCore::HLE::SyscallArguments::MAX_ARGS);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
if (Reg.GetCode() == x8.GetCode() ||
Reg.GetCode() == x4.GetCode() ||
Reg.GetCode() == x5.GetCode()) {
SpillMask |= (1U << Reg.GetCode());
Intersects = true;
}
}
// XXX: For some reason spilling only the x4, x5, and x8 registers was causing issues
// Come back to this once investigation reveals why it fails the gvisor ioctl test
// For now override to all GPRs
SpillMask = ~0U;
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Now that we have claimed to be a syscall we can set up the arguments
if (Intersects) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
}
}
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
auto Reg = GetReg<RA_64>(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
// for registers RAX, RBX, and RSI. Which have just been spilled
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg.GetCode() == x8.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == x4.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == x5.GetCode()) {
ldr(RegArgs[i], MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
mov(RegArgs[i], Reg);
}
}
else {
auto Reg = GetReg<RA_32>(Op->Header.Args[i].ID());
if (Reg.GetCode() == w8.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSI])));
}
else if (Reg.GetCode() == w4.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX])));
}
else if (Reg.GetCode() == w5.GetCode()) {
ldr(RegArgs[i].W(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RBX])));
}
else {
uxtw(RegArgs[i].W(), Reg);
}
}
}
}
else {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
if (CTX->Config.Is64BitMode()) {
mov(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
else {
uxtw(RegArgs[i], GetReg<RA_64>(Op->Header.Args[i].ID()));
}
}
}
LoadConstant(x8, Op->HostSyscallNumber);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
if ((Op->Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Result is now in x0
// Move result to its destination register
if (CTX->Config.Is64BitMode()) {
mov(GetReg<RA_64>(Node), x0);
}
else {
uxtw(GetReg<RA_64>(Node), x0);
}
}
}
DEF_OP(Thunk) {
@@ -256,7 +394,6 @@ DEF_OP(Thunk) {
FillStaticRegs(); // load from ctx after ra64 refill
}
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
const auto *OldCode = (const uint8_t *)&Op->CodeOriginalLow;
@@ -315,7 +452,7 @@ DEF_OP(RemoveCodeEntry) {
mov(x0, STATE);
LoadConstant(x1, Entry);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.RemoveCodeEntryFromJIT)));
SpillStaticRegs();
blr(x2);
FillStaticRegs();
@@ -332,19 +469,10 @@ DEF_OP(CPUID) {
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
LoadConstant(x0, reinterpret_cast<uint64_t>(&CTX->CPUID));
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDObj)));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Header.Args[0].ID()));
mov(x2, GetReg<RA_64>(Op->Header.Args[1].ID()));
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t, uint32_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::CPUIDEmu::RunFunction;
LoadConstant(x3, Ptr.Data);
SpillStaticRegs();
blr(x3);
FillStaticRegs();
@@ -363,13 +491,13 @@ void Arm64JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
REGISTER_OP(JUMP, Jump);
REGISTER_OP(CONDJUMP, CondJump);
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVECODEENTRY, RemoveCodeEntry);
@@ -10,25 +10,25 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
switch (Op->Header.ElementSize) {
case 1: {
ins(GetDst(Node).V16B(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 2: {
ins(GetDst(Node).V8H(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 4: {
ins(GetDst(Node).V4S(), Op->Index, GetReg<RA_32>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Header.Args[1].ID()));
break;
}
case 8: {
ins(GetDst(Node).V2D(), Op->Index, GetReg<RA_64>(Op->Header.Args[1].ID()));
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Header.Args[1].ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -39,11 +39,11 @@ DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
switch (Op->Header.ElementSize) {
case 1:
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
uxtb(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 2:
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()).W());
uxth(TMP1.W(), GetReg<RA_32>(Op->Header.Args[0].ID()));
fmov(GetDst(Node).S(), TMP1.W());
break;
case 4:
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -54,7 +54,7 @@ DEF_OP(AESDecLast) {
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Label Constant;
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
// Do a "regular" AESE step
@@ -63,8 +63,7 @@ DEF_OP(AESKeyGenAssist) {
aese(VTMP1.V16B(), VTMP2.V16B());
// Do a table shuffle to undo ShiftRows
adr(TMP1.X(), &Constant);
ldr(VTMP3, MemOperand(TMP1.X()));
ldr(VTMP3, &ConstantLiteral);
// Now EOR in the RCON
if (Op->RCON) {
@@ -80,14 +79,29 @@ DEF_OP(AESKeyGenAssist) {
}
b(&PastConstant);
bind(&Constant);
dc32(0x0B0E0104);
dc32(0x040B0E01);
dc32(0x0306090C);
dc32(0x0C030609);
place(&ConstantLiteral);
bind(&PastConstant);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (Op->SrcSize) {
case 1:
crc32cb(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 2:
crc32ch(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 4:
crc32cw(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_32>(Op->Src2.ID()));
break;
case 8:
crc32cx(GetReg<RA_32>(Node), GetReg<RA_32>(Op->Src1.ID()), GetReg<RA_64>(Op->Src2.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
#undef DEF_OP
void Arm64JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -97,7 +111,7 @@ void Arm64JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Header.Args[0].ID()), Op->Flag, 1);
+169 -100
View File
@@ -21,10 +21,13 @@ $end_info$
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <sys/mman.h>
@@ -32,6 +35,40 @@ $end_info$
#include <unistd.h>
#include <string.h>
namespace {
static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
}
namespace FEXCore::CPU {
void Arm64JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
@@ -42,7 +79,7 @@ void Arm64JITCore::CopyNecessaryDataForCompileThread(CPUBackend *Original) {
using namespace vixl;
using namespace vixl::aarch64;
void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -56,8 +93,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
LoadConstant(x1, (uintptr_t)Info.fn);
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -72,8 +108,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
LoadConstant(x0, (uintptr_t)Info.fn);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -92,8 +127,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
LoadConstant(x0, (uintptr_t)Info.fn);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x0);
PopDynamicRegsAndLR();
@@ -118,8 +152,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
else {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
LoadConstant(x1, (uintptr_t)Info.fn);
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x1);
PopDynamicRegsAndLR();
@@ -140,8 +173,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -160,8 +192,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -180,8 +211,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -199,8 +229,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -218,8 +247,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -240,8 +268,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -259,8 +286,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
LoadConstant(x2, (uintptr_t)Info.fn);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x2);
PopDynamicRegsAndLR();
@@ -283,8 +309,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
LoadConstant(x4, (uintptr_t)Info.fn);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.FallbackHandlerPointers[Info.HandlerIndex])));
blr(x4);
PopDynamicRegsAndLR();
@@ -307,7 +332,7 @@ void Arm64JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
}
void Arm64JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
Arm64JITCore::CodeBuffer Arm64JITCore::AllocateNewCodeBuffer(size_t Size) {
@@ -333,27 +358,9 @@ void Arm64JITCore::FreeCodeBuffer(CodeBuffer Buffer) {
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread)
: Arm64Emitter(0)
: Arm64Emitter(ctx, 0)
, CTX {ctx}
, ThreadState {Thread} {
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
CurrentCodeBuffer = &InitialCodeBuffer;
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
#if DEBUG
@@ -393,42 +400,100 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
RegisterVectorHandlers();
RegisterEncryptionHandlers();
{
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<Arm64Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
CallbackPtr = Dispatcher->CallbackPtr;
}
if (!CompileThread) {
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalReturnInstruction = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.AArch64;
// Process specific
Pointers.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
Pointers.LDIV = reinterpret_cast<uint64_t>(LDIV);
Pointers.LUREM = reinterpret_cast<uint64_t>(LUREM);
Pointers.LREM = reinterpret_cast<uint64_t>(LREM);
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
// Can't allocate a code buffer until after dispatcher is created
InitialCodeBuffer = AllocateNewCodeBuffer(Arm64JITCore::INITIAL_CODE_SIZE);
*GetBuffer() = vixl::CodeBuffer(InitialCodeBuffer.Ptr, InitialCodeBuffer.Size);
SetAllowAssembler(true);
EmitDetectionString();
CurrentCodeBuffer = &InitialCodeBuffer;
}
void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
if (!Core->Dispatcher->IsAddressInJITCode(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
return false;
}
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Core->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
Arm64JITCore *Core = reinterpret_cast<Arm64JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
void Arm64JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::Arm64JITCore::";
auto Buffer = GetBuffer();
Buffer->EmitString(JITString);
Buffer->Align();
}
void Arm64JITCore::ClearCache() {
@@ -471,6 +536,7 @@ void Arm64JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
*Buffer = vixl::CodeBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
@@ -482,7 +548,7 @@ Arm64JITCore::~Arm64JITCore() {
FreeCodeBuffer(InitialCodeBuffer);
}
IR::PhysicalRegister Arm64JITCore::GetPhys(uint32_t Node) const {
IR::PhysicalRegister Arm64JITCore::GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
@@ -491,7 +557,7 @@ IR::PhysicalRegister Arm64JITCore::GetPhys(uint32_t Node) const {
}
template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(uint32_t Node) const {
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
@@ -506,7 +572,7 @@ aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(uint32_t Node) const
}
template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(uint32_t Node) const {
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
@@ -521,18 +587,18 @@ aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(uint32_t Node) const
}
template<>
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_32>(uint32_t Node) const {
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_32>(IR::NodeID Node) const {
uint32_t Reg = GetPhys(Node).Reg;
return RA32Pair[Reg];
}
template<>
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_64>(uint32_t Node) const {
std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JITCore::RA_64>(IR::NodeID Node) const {
uint32_t Reg = GetPhys(Node).Reg;
return RA64Pair[Reg];
}
aarch64::VRegister Arm64JITCore::GetSrc(uint32_t Node) const {
aarch64::VRegister Arm64JITCore::GetSrc(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
@@ -546,7 +612,7 @@ aarch64::VRegister Arm64JITCore::GetSrc(uint32_t Node) const {
FEX_UNREACHABLE;
}
aarch64::VRegister Arm64JITCore::GetDst(uint32_t Node) const {
aarch64::VRegister Arm64JITCore::GetDst(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
@@ -580,7 +646,12 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -588,18 +659,18 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
}
}
FEXCore::IR::RegisterClassType Arm64JITCore::GetRegClass(uint32_t Node) const {
FEXCore::IR::RegisterClassType Arm64JITCore::GetRegClass(IR::NodeID Node) const {
return FEXCore::IR::RegisterClassType {GetPhys(Node).Class};
}
bool Arm64JITCore::IsFPR(uint32_t Node) const {
bool Arm64JITCore::IsFPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
}
bool Arm64JITCore::IsGPR(uint32_t Node) const {
bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
auto Class = GetRegClass(Node);
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
@@ -666,13 +737,13 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip)));
// Stop the thread
LoadConstant(x0, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(x0);
}
bind(&RunBlock);
}
//LOGMAN_THROW_A(RAData->HasFullRA(), "Arm64 JIT only works with RA");
//LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
SpillSlots = RAData->SpillSlots();
@@ -694,12 +765,10 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uintptr_t BlockStartHostCode = GetCursorAddress<uintptr_t>();
{
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
@@ -711,12 +780,8 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({GetCursorAddress<uintptr_t>(), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
uint32_t ID = IR->GetID(CodeNode);
const auto ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
@@ -724,7 +789,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = GetCursorAddress<uintptr_t>() - DebugData->Subblocks.back().HostCodeStart;
DebugData->Subblocks.push_back({BlockStartHostCode, static_cast<uint32_t>(GetCursorAddress<uintptr_t>() - BlockStartHostCode)});
}
}
@@ -801,4 +866,8 @@ uint64_t Arm64JITCore::ExitFunctionLink(Arm64JITCore *core, FEXCore::Core::CpuSt
std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<Arm64JITCore>(ctx, Thread, CompileThread);
}
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX) {
Arm64JITCore::InitializeSignalHandlers(CTX);
}
}
+35 -20
View File
@@ -64,6 +64,11 @@ public:
[[nodiscard]] CodeBuffer AllocateNewCodeBuffer(size_t Size);
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
@@ -75,7 +80,7 @@ private:
FEXCore::IR::IRListView const *IR;
uint64_t Entry;
std::map<IR::OrderedNodeWrapper::NodeOffsetType, aarch64::Label> JumpTargets;
std::map<IR::NodeID, aarch64::Label> JumpTargets;
/**
* @name Register Allocation
@@ -99,30 +104,30 @@ private:
constexpr static uint8_t RA_FPR = 2;
template<uint8_t RAType>
[[nodiscard]] aarch64::Register GetReg(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg(IR::NodeID Node) const;
template<>
[[nodiscard]] aarch64::Register GetReg<RA_32>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] aarch64::Register GetReg<RA_64>(uint32_t Node) const;
[[nodiscard]] aarch64::Register GetReg<RA_64>(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair(IR::NodeID Node) const;
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_32>(IR::NodeID Node) const;
template<>
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(uint32_t Node) const;
[[nodiscard]] std::pair<aarch64::Register, aarch64::Register> GetSrcPair<RA_64>(IR::NodeID Node) const;
[[nodiscard]] aarch64::VRegister GetSrc(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetDst(uint32_t Node) const;
[[nodiscard]] aarch64::VRegister GetSrc(IR::NodeID Node) const;
[[nodiscard]] aarch64::VRegister GetDst(IR::NodeID Node) const;
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(uint32_t Node) const;
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
[[nodiscard]] MemOperand GenerateMemOperand(uint8_t AccessSize,
aarch64::Register Base,
@@ -172,17 +177,22 @@ private:
struct CompilerSharedData {
uint64_t SignalReturnInstruction{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
};
CompilerSharedData ThreadSharedData;
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
using OpHandler = void (Arm64JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
void RegisterBranchHandlers();
@@ -193,7 +203,7 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -229,6 +239,8 @@ private:
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -270,13 +282,13 @@ private:
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
DEF_OP(Jump);
DEF_OP(CondJump);
DEF_OP(Syscall);
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveCodeEntry);
@@ -297,7 +309,7 @@ private:
DEF_OP(GetHostFlag);
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
@@ -316,6 +328,7 @@ private:
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -326,6 +339,8 @@ private:
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -335,8 +350,6 @@ private:
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector2);
DEF_OP(SplatVector4);
DEF_OP(VMov);
@@ -424,6 +437,7 @@ private:
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -432,6 +446,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
@@ -11,7 +11,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -623,7 +623,7 @@ DEF_OP(LoadMemTSO) {
LOGMAN_MSG_A_FMT("LoadMemTSO: No offset allowed");
}
if (SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
auto Dst = GetReg<RA_64>(Node);
@@ -698,29 +698,29 @@ DEF_OP(LoadMemTSO) {
DEF_OP(StoreMem) {
auto Op = IROp->C<IR::IROp_StoreMem>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
auto MemReg = GetReg<RA_64>(Op->Addr.ID());
auto MemSrc = GenerateMemOperand(IROp->Size, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
switch (IROp->Size) {
case 1:
strb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
strb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 2:
strh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
strh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
str(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
str(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
str(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
str(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
@@ -744,7 +744,7 @@ DEF_OP(StoreMem) {
DEF_OP(StoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("StoreMemTSO: No offset allowed");
@@ -753,19 +753,19 @@ DEF_OP(StoreMemTSO) {
if (Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
nop();
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", IROp->Size);
}
@@ -774,7 +774,7 @@ DEF_OP(StoreMemTSO) {
}
else {
dmb(InnerShareable, BarrierAll);
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
switch (IROp->Size) {
case 1:
str(Src.B(), MemSrc);
@@ -800,7 +800,7 @@ DEF_OP(StoreMemTSO) {
DEF_OP(ParanoidLoadMemTSO) {
auto Op = IROp->C<IR::IROp_LoadMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("ParanoidLoadMemTSO: No offset allowed");
@@ -857,7 +857,7 @@ DEF_OP(ParanoidLoadMemTSO) {
DEF_OP(ParanoidStoreMemTSO) {
auto Op = IROp->C<IR::IROp_StoreMemTSO>();
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Header.Args[0].ID()));
auto MemSrc = MemOperand(GetReg<RA_64>(Op->Addr.ID()));
if (!Op->Offset.IsInvalid()) {
LOGMAN_MSG_A_FMT("ParanoidStoreMemTSO: No offset allowed");
@@ -866,25 +866,25 @@ DEF_OP(ParanoidStoreMemTSO) {
if (Op->Class == FEXCore::IR::GPRClass) {
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
stlrb(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrb(GetReg<RA_64>(Op->Value.ID()), MemSrc);
}
else {
switch (IROp->Size) {
case 2:
stlrh(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlrh(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
case 4:
stlr(GetReg<RA_32>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_32>(Op->Value.ID()), MemSrc);
break;
case 8:
stlr(GetReg<RA_64>(Op->Header.Args[1].ID()), MemSrc);
stlr(GetReg<RA_64>(Op->Value.ID()), MemSrc);
break;
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", IROp->Size);
}
}
}
else {
auto Src = GetSrc(Op->Header.Args[1].ID());
auto Src = GetSrc(Op->Value.ID());
if (IROp->Size == 1) {
// 8bit load is always aligned to natural alignment
mov(TMP1.W(), Src.V16B(), 0);
@@ -939,13 +939,35 @@ DEF_OP(CacheLineClear) {
// Clear dcache only
// icache doesn't matter here since the guest application shouldn't be calling clflush on JIT code.
mov(TMP1, MemReg);
for (size_t i = 0; i < std::max(1U, DCacheLineSize / 64U); ++i) {
for (size_t i = 0; i < std::max(1U, CTX->HostFeatures.DCacheLineSize / 64U); ++i) {
dc(DataCacheOp::CVAU, TMP1);
add(TMP1, TMP1, DCacheLineSize);
add(TMP1, TMP1, CTX->HostFeatures.DCacheLineSize);
}
dsb(InnerShareable, BarrierAll);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
auto MemReg = GetReg<RA_64>(Op->Header.Args[0].ID());
if (CTX->HostFeatures.SupportsCLZERO) {
// We can use this instruction directly
dc(DataCacheOp::ZVA, MemReg);
}
else {
// We must walk the cacheline ourselves
// Force cacheline alignment
and_(TMP1, MemReg, ~(CPUIDEmu::CACHELINE_SIZE - 1));
// This will end up being four STPs
// Depending on uarch it could be slightly more efficient in instructions emitted
// and uops to use vector pair STP, but we want the non-temporal bit specifically here
for (size_t i = 0; i < CPUIDEmu::CACHELINE_SIZE; i += 16) {
stnp(xzr, xzr, MemOperand(TMP1, i, Offset));
}
}
}
#undef DEF_OP
void Arm64JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -972,6 +994,7 @@ void Arm64JITCore::RegisterMemoryHandlers() {
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
}
}
+94 -25
View File
@@ -7,17 +7,9 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -38,35 +30,39 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 0: // Hard fault
case 5: // Guest ud2
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
hlt(4);
break;
case 1: // Int <imm8>
hlt(4);
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.OverflowExceptionHandler)));
br(TMP1);
break;
case 2: // overflow
hlt(4);
break;
case 3: // int 1
hlt(4);
break;
case 4: { // HLT
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddressSpillSRA);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadStopHandlerSpillSRA)));
br(TMP1);
break;
}
case 6: { // INT3
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.ThreadPauseHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
ResetStack();
LoadConstant(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddressSpillSRA);
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.UnimplementedInstructionHandler)));
br(TMP1);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
@@ -138,13 +134,13 @@ DEF_OP(Print) {
if (IsGPR(Op->Header.Args[0].ID())) {
mov(x0, GetReg<RA_64>(Op->Header.Args[0].ID()));
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintValue));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintValue)));
}
else {
fmov(x0, GetSrc(Op->Header.Args[0].ID()).V1D());
// Bug in vixl that source vector needs to b V1D rather than V2D?
fmov(x1, GetSrc(Op->Header.Args[0].ID()).V1D(), 1);
LoadConstant(x3, reinterpret_cast<uint64_t>(PrintVectorValue));
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.PrintVectorValue)));
}
SpillStaticRegs();
blr(x3);
@@ -153,6 +149,76 @@ DEF_OP(Print) {
PopDynamicRegsAndLR();
}
DEF_OP(ProcessorID) {
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
// Ordering is incredibly important here
// We must spill any overlapping registers first THEN claim we are in a syscall without invalidating state at all
// Only spill the registers that intersect with our usage
SpillStaticRegs(false, SpillMask);
// Now that we are spilled, store in the state that we are in a syscall
// Still without overwriting registers that matter
// 16bit LoadConstant to be a single instruction
// We must always spill at least one register (x8) so this value always has a bit set
// This gives the signal handler a value to check to see if we are in a syscall at all
LoadConstant(x0, SpillMask & 0xFFFF);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Allocate some temporary space for storing the uint32_t CPU and Node IDs
sub(sp, sp, 16);
// Load the getcpu syscall number
LoadConstant(x8, SYS_getcpu);
// CPU pointer in x0
add(x0, sp, 0);
// Node in x1
add(x1, sp, 4);
svc(0);
// On updated signal mask we can receive a signal RIGHT HERE
// Load the values returned by the kernel
ldp(w0, w1, MemOperand(sp));
// Deallocate stack space
sub(sp, sp, 16);
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
str(xzr, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, InSyscallInfo)));
// Now store the result in the destination in the expected format
// uint32_t Res = (node << 12) | cpu;
// CPU is in w0
// Node is in w1
orr(GetReg<RA_64>(Node), x0, Operand(x1, LSL, 12));
}
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetSrcPair<RA_64>(Node);
if (Op->GetReseeded) {
mrs(Dst.first, RNDRRS);
}
else {
mrs(Dst.first, RNDR);
}
// If the rng number is valid then NZCV is 0b0000, otherwise NZCV is 0b0100
cset(Dst.second, Condition::ne);
}
#undef DEF_OP
void Arm64JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
@@ -169,6 +235,9 @@ void Arm64JITCore::RegisterMiscHandlers() {
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
#undef REGISTER_OP
}
}
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -37,7 +37,7 @@ DEF_OP(CreateElementPair) {
aarch64::Register RegSecond;
aarch64::Register RegTmp;
switch (Op->Header.Size) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VectorZero) {
uint8_t OpSize = IROp->Size;
switch (OpSize) {
@@ -42,14 +42,6 @@ DEF_OP(VectorImm) {
}
}
DEF_OP(CreateVector2) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector2) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1758,8 +1750,7 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
switch (Op->Header.Size) {
case 1:
mov(GetDst(Node).B(), GetSrc(Op->Header.Args[0].ID()).V16B(), Op->Index);
break;
@@ -1772,7 +1763,7 @@ DEF_OP(VExtractElement) {
case 8:
mov(GetDst(Node).D(), GetSrc(Op->Header.Args[0].ID()).V2D(), Op->Index);
break;
default: LOGMAN_MSG_A_FMT("Unhandled VExtractElement element size: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unhandled VExtractElement element size: {}", Op->Header.Size);
}
}
@@ -2318,13 +2309,27 @@ DEF_OP(VTBL1) {
}
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VRev64>();
uint8_t OpSize = IROp->Size;
uint8_t Elements = OpSize / Op->Header.ElementSize;
// Vector
switch (Op->Header.ElementSize) {
case 1:
case 2:
case 4:
rev64(GetDst(Node).VCast(OpSize * 8, Elements), GetSrc(Op->Header.Args[0].ID()).VCast(OpSize * 8, Elements));
break;
case 8:
default: LOGMAN_MSG_A_FMT("Invalid Element Size: {}", Op->Header.ElementSize); break;
}
}
#undef DEF_OP
void Arm64JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector2);
REGISTER_OP(SPLATVECTOR4, SplatVector4);
REGISTER_OP(VMOV, VMov);
@@ -2413,6 +2418,7 @@ void Arm64JITCore::RegisterVectorHandlers() {
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
#undef REGISTER_OP
}
}
+3
View File
@@ -16,8 +16,11 @@ class CPUBackend;
[[nodiscard]] std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX);
[[nodiscard]] std::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::Context *ctx,
FEXCore::Core::InternalThreadState *Thread,
bool CompileThread);
void InitializeArm64JITSignalHandlers(FEXCore::Context::Context *CTX);
} // namespace FEXCore::CPU
@@ -20,11 +20,11 @@ namespace FEXCore::CPU {
#define GRD(Node) (IROp->Size <= 4 ? GetDst<RA_32>(Node) : GetDst<RA_64>(Node))
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
switch (Op->Size) {
switch (IROp->Size) {
case 4: {
auto Dst = GetSrcPair<RA_32>(Node);
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
@@ -32,7 +32,7 @@ DEF_OP(TruncElementPair) {
mov(Dst.second, Src.second);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", Op->Size); break;
default: LOGMAN_MSG_A_FMT("Unhandled Truncation size: {}", IROp->Size); break;
}
}
@@ -45,7 +45,13 @@ DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = Entry + Op->Offset;
mov(GetDst<RA_64>(Node), Constant);
uint64_t Mask = ~0ULL;
uint8_t OpSize = IROp->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
mov(GetDst<RA_64>(Node), Constant & Mask);
}
DEF_OP(InlineConstant) {
@@ -671,6 +677,36 @@ DEF_OP(Extr) {
}
}
DEF_OP(PDep) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const auto Input = GRS(Op->Args(0).ID());
const auto Mask = GRS(Op->Args(1).ID());
const auto Dest = GRD(Node);
if (OpSize == 4) {
pdep(Dest.cvt32(), Input.cvt32(), Mask.cvt32());
} else {
pdep(Dest.cvt64(), Input.cvt64(), Mask.cvt64());
}
}
DEF_OP(PExt) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto OpSize = IROp->Size;
const auto Input = GRS(Op->Args(0).ID());
const auto Mask = GRS(Op->Args(1).ID());
const auto Dest = GRD(Node);
if (OpSize == 4) {
pext(Dest.cvt32(), Input.cvt32(), Mask.cvt32());
} else {
pext(Dest.cvt64(), Input.cvt64(), Mask.cvt64());
}
}
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
@@ -1114,19 +1150,19 @@ DEF_OP(VExtractToGPR) {
switch (Op->Header.ElementSize) {
case 1: {
pextrb(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrb(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 2: {
pextrw(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrw(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 4: {
pextrd(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrd(GetDst<RA_32>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
case 8: {
pextrq(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Idx);
pextrq(GetDst<RA_64>(Node), GetSrc(Op->Header.Args[0].ID()), Op->Index);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -1248,6 +1284,8 @@ void X86JITCore::RegisterALUHandlers() {
REGISTER_OP(ASHR, Ashr);
REGISTER_OP(ROR, Ror);
REGISTER_OP(EXTR, Extr);
REGISTER_OP(PDEP, PDep);
REGISTER_OP(PEXT, PExt);
REGISTER_OP(LDIV, LDiv);
REGISTER_OP(LUDIV, LUDiv);
REGISTER_OP(LREM, LRem);
@@ -15,23 +15,19 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Desired
// Args[1]: Expected
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
// Third operand must be a calculated guest memory address
//OrderedNode *CASResult = _CAS(Src3, Src2, Src1);
auto Dst = GetSrcPair<RA_64>(Node);
auto Expected = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Desired = GetSrcPair<RA_64>(Op->Header.Args[1].ID());
auto MemSrc = GetSrc<RA_64>(Op->Header.Args[2].ID());
auto Expected = GetSrcPair<RA_64>(Op->Expected.ID());
auto Desired = GetSrcPair<RA_64>(Op->Desired.ID());
auto MemSrc = GetSrc<RA_64>(Op->Addr.ID());
Xbyak::Reg MemReg = MemSrc;
@@ -47,7 +43,7 @@ DEF_OP(CASPair) {
lock();
switch (OpSize) {
switch (IROp->ElementSize) {
case 4: {
cmpxchg8b(dword [MemReg]);
// EDX:EAX now contains the result
@@ -62,7 +58,7 @@ DEF_OP(CASPair) {
mov(Dst.second, rdx);
break;
}
default: LOGMAN_MSG_A_FMT("Unsupported: {}", OpSize);
default: LOGMAN_MSG_A_FMT("Unsupported: {}", IROp->ElementSize);
}
}
@@ -70,18 +66,15 @@ DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
// Args[0]: Desired
// Args[1]: Expected
// Args[2]: Pointer
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
// Third operand must be a calculated guest memory address
//OrderedNode *CASResult = _CAS(Src3, Src2, Src1);
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[2].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, GetSrc<RA_64>(Op->Expected.ID()));
// RCX now contains pointer
// RAX contains our expected value
@@ -90,23 +83,23 @@ DEF_OP(CAS) {
lock();
switch (OpSize) {
case 1: {
cmpxchg(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
cmpxchg(byte [MemReg], GetSrc<RA_8>(Op->Desired.ID()));
movzx(GetDst<RA_64>(Node), al);
break;
}
case 2: {
cmpxchg(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
cmpxchg(word [MemReg], GetSrc<RA_16>(Op->Desired.ID()));
movzx(GetDst<RA_64>(Node), ax);
break;
}
case 4: {
cmpxchg(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
cmpxchg(dword [MemReg], GetSrc<RA_32>(Op->Desired.ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), eax);
break;
}
case 8: {
cmpxchg(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
cmpxchg(qword [MemReg], GetSrc<RA_64>(Op->Desired.ID()));
// RAX now contains the result
mov (GetDst<RA_64>(Node), rax);
break;
@@ -118,21 +111,21 @@ DEF_OP(CAS) {
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
add(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
add(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
add(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
add(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
add(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
add(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
add(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
add(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -141,20 +134,20 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
sub(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
sub(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
sub(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
sub(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
sub(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
sub(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
sub(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
sub(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -163,20 +156,20 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
and_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
and_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
and_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
and_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -185,20 +178,20 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
or_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
or_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
or_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
or_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -207,20 +200,20 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
lock();
switch (IROp->Size) {
case 1:
xor_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(byte [MemReg], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
xor_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(word [MemReg], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
xor_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(dword [MemReg], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
xor_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(qword [MemReg], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled AtomicAdd size: {}", IROp->Size);
}
@@ -230,26 +223,26 @@ DEF_OP(AtomicSwap) {
auto Op = IROp->C<IR::IROp_AtomicSwap>();
Xbyak::Reg MemReg = rax;
mov(MemReg, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(MemReg, GetSrc<RA_64>(Op->Addr.ID()));
switch (IROp->Size) {
case 1:
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_8>(Op->Value.ID()));
lock();
xchg(byte [MemReg], GetDst<RA_8>(Node));
break;
case 2:
movzx(GetDst<RA_64>(Node), GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(GetDst<RA_64>(Node), GetSrc<RA_16>(Op->Value.ID()));
lock();
xchg(word [MemReg], GetDst<RA_16>(Node));
break;
case 4:
mov(GetDst<RA_64>(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_32>(Op->Value.ID()));
lock();
xchg(dword [MemReg], GetDst<RA_32>(Node));
break;
case 8:
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Value.ID()));
lock();
xchg(qword [MemReg], GetDst<RA_64>(Node));
break;
@@ -260,28 +253,28 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1:
movzx(rcx, GetSrc<RA_8>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_8>(Op->Value.ID()));
lock();
xadd(byte [MemReg], cl);
movzx(GetDst<RA_32>(Node), cl);
break;
case 2:
movzx(rcx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
movzx(rcx, GetSrc<RA_16>(Op->Value.ID()));
lock();
xadd(word [MemReg], cx);
movzx(GetDst<RA_32>(Node), cx);
break;
case 4:
mov(ecx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(ecx, GetSrc<RA_32>(Op->Value.ID()));
lock();
xadd(dword [MemReg], ecx);
mov(GetDst<RA_64>(Node), ecx);
break;
case 8:
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(rcx, GetSrc<RA_64>(Op->Value.ID()));
lock();
xadd(qword [MemReg], rcx);
mov(GetDst<RA_64>(Node), rcx);
@@ -293,31 +286,31 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1:
mov(cl, GetSrc<RA_8>(Op->Header.Args[1].ID()));
mov(cl, GetSrc<RA_8>(Op->Value.ID()));
neg(cl);
lock();
xadd(byte [MemReg], cl);
movzx(GetDst<RA_32>(Node), cl);
break;
case 2:
mov(cx, GetSrc<RA_16>(Op->Header.Args[1].ID()));
mov(cx, GetSrc<RA_16>(Op->Value.ID()));
neg(cx);
lock();
xadd(word [MemReg], cx);
movzx(GetDst<RA_32>(Node), cx);
break;
case 4:
mov(ecx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(ecx, GetSrc<RA_32>(Op->Value.ID()));
neg(ecx);
lock();
xadd(dword [MemReg], ecx);
mov(GetDst<RA_32>(Node), ecx);
break;
case 8:
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(rcx, GetSrc<RA_64>(Op->Value.ID()));
neg(rcx);
lock();
xadd(qword [MemReg], rcx);
@@ -331,7 +324,7 @@ DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
@@ -341,7 +334,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -357,7 +350,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -374,7 +367,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -391,7 +384,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -409,7 +402,7 @@ DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -418,7 +411,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -434,7 +427,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -451,7 +444,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -468,7 +461,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -486,7 +479,7 @@ DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
// TMP1 = rax
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -495,7 +488,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -511,7 +504,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -528,7 +521,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -545,7 +538,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Value.ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -562,7 +555,7 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Header.Args[0].ID());
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
switch (IROp->Size) {
case 1: {
mov(TMP1.cvt8(), byte [MemReg]);
@@ -28,7 +28,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GuestCallDirect) {
LogMan::Msg::DFmt("Unimplemented");
}
@@ -37,18 +37,13 @@ DEF_OP(GuestCallIndirect) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(GuestReturn) {
LogMan::Msg::DFmt("Unimplemented");
}
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
}
mov(TMP1, ThreadSharedData.SignalHandlerReturnAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalReturnHandler)]);
}
DEF_OP(CallbackReturn) {
@@ -58,8 +53,7 @@ DEF_OP(CallbackReturn) {
}
// Make sure to adjust the refcounter so we don't clear the cache now
mov(rax, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
sub(dword [rax], 1);
sub(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SignalHandlerRefCountPointer)], 1);
// We need to adjust an additional 8 bytes to get back to the original "misaligned" RSP state
add(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP])], 8);
@@ -104,7 +98,8 @@ DEF_OP(ExitFunction) {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, ThreadState->LookupCache->GetL1Pointer());
mov(rcx, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.L1Pointer)]);
mov(rax, RipReg);
and_(rax, LookupCache::L1_ENTRIES_MASK);
@@ -117,9 +112,8 @@ DEF_OP(ExitFunction) {
jmp(qword[LookupBase + 0]);
L(FullLookup);
mov(rax, ThreadSharedData.Dispatcher->AbsoluteLoopTopAddress);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, State.rip)], RipReg);
jmp(rax);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.DispatcherLoopTop)]);
}
#ifdef BLOCKSTATS
@@ -128,18 +122,10 @@ DEF_OP(ExitFunction) {
}
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
Label *TargetLabel;
auto IsTarget = JumpTargets.find(Op->Header.Args[0].ID());
if (IsTarget == JumpTargets.end()) {
TargetLabel = &JumpTargets.try_emplace(Op->Header.Args[0].ID()).first->second;
}
else {
TargetLabel = &IsTarget->second;
}
PendingTargetLabel = TargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
@@ -147,18 +133,7 @@ DEF_OP(Jump) {
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
Label *TrueTargetLabel;
Label *FalseTargetLabel;
auto TrueIter = JumpTargets.find(Op->TrueBlock.ID());
auto FalseIter = JumpTargets.find(Op->FalseBlock.ID());
if (TrueIter == JumpTargets.end()) {
TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
}
else {
TrueTargetLabel = &TrueIter->second;
}
Label *TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
if (IsGPR(Op->Cmp1.ID())) {
uint64_t Const;
@@ -168,24 +143,18 @@ DEF_OP(CondJump) {
cmp(GRCMP(Op->Cmp1.ID()), GRCMP(Op->Cmp2.ID()));
}
} else if (IsFPR(Op->Cmp1.ID())) {
if (Op->CompareSize == 4)
if (Op->CompareSize == 4) {
ucomiss(GetSrc(Op->Cmp1.ID()), GetSrc(Op->Cmp2.ID()));
else
} else {
ucomisd(GetSrc(Op->Cmp1.ID()), GetSrc(Op->Cmp2.ID()));
}
}
auto [_, __, JCC] = GetCC(Op->Cond);
(this->*JCC)(*TrueTargetLabel, T_NEAR);
if (FalseIter == JumpTargets.end()) {
FalseTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
else {
FalseTargetLabel = &FalseIter->second;
}
PendingTargetLabel = FalseTargetLabel;
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
}
DEF_OP(Syscall) {
@@ -212,15 +181,13 @@ DEF_OP(Syscall) {
}
mov(rsi, STATE); // Move thread in to rsi
mov(rdi, reinterpret_cast<uint64_t>(CTX->SyscallHandler));
mov(rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerObj)]);
mov(rdx, rsp);
mov(rax, reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall));
if (NumPush & 1)
sub(rsp, 8); // Align
// {rdi, rsi, rdx}
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.SyscallHandlerFunc)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -305,9 +272,7 @@ DEF_OP(RemoveCodeEntry) {
mov(rax, Entry); // imm64 move
mov(rsi, rax);
mov(rax, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit));
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.RemoveCodeEntryFromJIT)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -319,13 +284,6 @@ DEF_OP(RemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
using ClassPtrType = FEXCore::CPUID::FunctionResults (FEXCore::CPUIDEmu::*)(uint32_t Function, uint32_t Leaf);
union {
ClassPtrType ClassPtr;
uint64_t Raw;
} Ptr;
Ptr.ClassPtr = &CPUIDEmu::RunFunction;
for (auto &Reg : RA64)
push(Reg);
@@ -338,18 +296,15 @@ DEF_OP(CPUID) {
// rsi can be in the source registers, so copy argument to edx first
mov (edx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov (esi, GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov (rdi, reinterpret_cast<uint64_t>(&CTX->CPUID));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDObj)]);
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
mov(rax, Ptr.Raw);
// {rdi, rsi, rdx}
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.CPUIDFunction)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -367,7 +322,6 @@ void X86JITCore::RegisterBranchHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(GUESTCALLDIRECT, GuestCallDirect);
REGISTER_OP(GUESTCALLINDIRECT, GuestCallIndirect);
REGISTER_OP(GUESTRETURN, GuestReturn);
REGISTER_OP(SIGNALRETURN, SignalReturn);
REGISTER_OP(CALLBACKRETURN, CallbackReturn);
REGISTER_OP(EXITFUNCTION, ExitFunction);
@@ -15,26 +15,26 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
movapd(GetDst(Node), GetSrc(Op->DestVector.ID()));
switch (Op->Header.ElementSize) {
case 1: {
pinsrb(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrb(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 2: {
pinsrw(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrw(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 4: {
pinsrd(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[1].ID()), Op->Index);
pinsrd(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 8: {
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[1].ID()), Op->Index);
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()), Op->DestIdx);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
@@ -13,7 +13,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -45,6 +45,36 @@ DEF_OP(AESKeyGenAssist) {
vaeskeygenassist(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), Op->RCON);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (IROp->Size) {
case 4:
mov(TMP1, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
break;
case 8:
mov(TMP1, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", IROp->Size);
}
switch (Op->SrcSize) {
case 1:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt8());
break;
case 2:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt16());
break;
case 4:
crc32(GetDst<RA_32>(Node).cvt32(), TMP1.cvt32());
break;
case 8:
crc32(GetDst<RA_64>(Node).cvt64(), TMP1.cvt64());
break;
}
}
#undef DEF_OP
void X86JITCore::RegisterEncryptionHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -54,7 +84,7 @@ void X86JITCore::RegisterEncryptionHandlers() {
REGISTER_OP(VAESDEC, AESDec);
REGISTER_OP(VAESDECLAST, AESDecLast);
REGISTER_OP(VAESKEYGENASSIST, AESKeyGenAssist);
REGISTER_OP(CRC32, CRC32);
#undef REGISTER_OP
}
}
@@ -14,7 +14,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
+118 -86
View File
@@ -15,6 +15,8 @@ $end_info$
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
@@ -27,7 +29,6 @@ $end_info$
#include <algorithm>
#include <array>
#include <bits/types/stack_t.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
@@ -42,6 +43,16 @@ $end_info$
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
namespace {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
}
namespace FEXCore::CPU {
CodeBuffer AllocateNewCodeBuffer(FEXCore::Context::Context *CTX, size_t Size) {
@@ -98,7 +109,7 @@ void X86JITCore::PopRegs() {
add(rsp, 16 * RAXMM_x.size());
}
void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -109,9 +120,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
case FABI_VOID_U16: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
break;
@@ -120,9 +129,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushRegs();
movss(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -136,9 +143,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -153,9 +158,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -171,9 +174,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -187,9 +188,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -203,9 +202,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -218,9 +215,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -233,9 +228,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -251,9 +244,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -266,9 +257,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -286,9 +275,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.FallbackHandlerPointers[Info.HandlerIndex])]);
PopRegs();
@@ -308,7 +295,7 @@ void X86JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
}
}
void X86JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
void X86JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
}
X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread)
@@ -318,6 +305,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
, InitialCodeBuffer {Buffer}
{
CurrentCodeBuffer = &InitialCodeBuffer;
EmitDetectionString();
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -350,6 +338,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
DispatcherConfig config;
config.ExitFunctionLink = reinterpret_cast<uintptr_t>(&ExitFunctionLink);
config.ExitFunctionLinkThis = reinterpret_cast<uintptr_t>(this);
config.StaticRegisterAssignment = ctx->Config.StaticRegisterAllocation;
Dispatcher = std::make_unique<X86Dispatcher>(CTX, ThreadState, config);
DispatchPtr = Dispatcher->DispatchPtr;
@@ -357,27 +346,55 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
ThreadSharedData.SignalHandlerRefCounterPtr = &Dispatcher->SignalHandlerRefCounter;
ThreadSharedData.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress;
ThreadSharedData.UnimplementedInstructionAddress = Dispatcher->UnimplementedInstructionAddress;
ThreadSharedData.OverflowExceptionInstructionAddress = Dispatcher->OverflowExceptionInstructionAddress;
ThreadSharedData.Dispatcher = Dispatcher.get();
}
// This will register the host signal handler per thread, which is fine
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
{
// Set up pointers that the JIT needs to load
auto &Pointers = ThreadState->CurrentFrame->Pointers.X86;
// Process specific
Pointers.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Pointers.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Pointers.RemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntryFromJit);
Pointers.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::CPUIDEmu::RunFunction);
Pointers.CPUIDFunction = PMF.GetConvertedPointer();
}
Pointers.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Pointers.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Pointers.FallbackHandlerPointers);
// Thread Specific
Pointers.SignalHandlerRefCountPointer = reinterpret_cast<uint64_t>(&Dispatcher->SignalHandlerRefCounter);
}
}
void X86JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
CTX->SignalDelegation->RegisterHostSignalHandler(SIGILL, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSIGILL(Signal, info, ucontext);
}, true);
CTX->SignalDelegation->RegisterHostSignalHandler(SignalDelegator::SIGNAL_FOR_PAUSE, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleSignalPause(Signal, info, ucontext);
}, true);
auto GuestSignalHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) -> bool {
X86JITCore *Core = reinterpret_cast<X86JITCore*>(Thread->CPUBackend.get());
return Core->Dispatcher->HandleGuestSignal(Signal, info, ucontext, GuestAction, GuestStack);
};
for (uint32_t Signal = 0; Signal <= SignalDelegator::MAX_SIGNALS; ++Signal) {
CTX->SignalDelegation->RegisterHostSignalHandlerForGuest(Signal, GuestSignalHandler);
}
}
@@ -391,6 +408,13 @@ X86JITCore::~X86JITCore() {
FreeCodeBuffer(InitialCodeBuffer);
}
void X86JITCore::EmitDetectionString() {
const char JITString[] = "FEXJIT::X86JITCore::";
for (char c : JITString) {
db(c);
}
}
void X86JITCore::ClearCache() {
if (*ThreadSharedData.SignalHandlerRefCounterPtr == 0) {
if (!CodeBuffers.empty()) {
@@ -429,9 +453,11 @@ void X86JITCore::ClearCache() {
EmplaceNewCodeBuffer(NewCodeBuffer);
setNewBuffer(NewCodeBuffer.Ptr, NewCodeBuffer.Size);
}
EmitDetectionString();
}
IR::PhysicalRegister X86JITCore::GetPhys(uint32_t Node) const {
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
@@ -439,16 +465,16 @@ IR::PhysicalRegister X86JITCore::GetPhys(uint32_t Node) const {
return PhyReg;
}
bool X86JITCore::IsFPR(uint32_t Node) const {
bool X86JITCore::IsFPR(IR::NodeID Node) const {
return RAData->GetNodeRegister(Node).Class == IR::FPRClass.Val;
}
bool X86JITCore::IsGPR(uint32_t Node) const {
bool X86JITCore::IsGPR(IR::NodeID Node) const {
return RAData->GetNodeRegister(Node).Class == IR::GPRClass.Val;
}
template<uint8_t RAType>
Xbyak::Reg X86JITCore::GetSrc(uint32_t Node) const {
Xbyak::Reg X86JITCore::GetSrc(IR::NodeID Node) const {
// rax, rcx, rdx, rsi, r8, r9,
// r10
// Callee Saved
@@ -467,24 +493,24 @@ Xbyak::Reg X86JITCore::GetSrc(uint32_t Node) const {
}
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_64>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_64>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_32>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_32>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_16>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_16>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_8>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetSrc<X86JITCore::RA_8>(IR::NodeID Node) const;
Xbyak::Xmm X86JITCore::GetSrc(uint32_t Node) const {
Xbyak::Xmm X86JITCore::GetSrc(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
template<uint8_t RAType>
Xbyak::Reg X86JITCore::GetDst(uint32_t Node) const {
Xbyak::Reg X86JITCore::GetDst(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
if constexpr (RAType == RA_64)
return RA64[PhyReg.Reg].cvt64();
@@ -499,19 +525,19 @@ Xbyak::Reg X86JITCore::GetDst(uint32_t Node) const {
}
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_64>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_64>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_32>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_32>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_16>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_16>(IR::NodeID Node) const;
template
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_8>(uint32_t Node) const;
Xbyak::Reg X86JITCore::GetDst<X86JITCore::RA_8>(IR::NodeID Node) const;
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(uint32_t Node) const {
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
if constexpr (RAType == RA_64)
return RA64Pair[PhyReg.Reg];
@@ -520,12 +546,12 @@ std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair(uint32_t Node) const {
}
template
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_64>(uint32_t Node) const;
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_64>(IR::NodeID Node) const;
template
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_32>(uint32_t Node) const;
std::pair<Xbyak::Reg, Xbyak::Reg> X86JITCore::GetSrcPair<X86JITCore::RA_32>(IR::NodeID Node) const;
Xbyak::Xmm X86JITCore::GetDst(uint32_t Node) const {
Xbyak::Xmm X86JITCore::GetDst(IR::NodeID Node) const {
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
@@ -550,7 +576,12 @@ bool X86JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, u
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = Entry + Op->Offset;
uint64_t Mask = ~0ULL;
uint8_t OpSize = OpHeader->Size;
if (OpSize == 4) {
Mask = 0xFFFF'FFFFULL;
}
*Value = (Entry + Op->Offset) & Mask;
}
return true;
} else {
@@ -687,15 +718,11 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
// if there is a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
@@ -724,10 +751,10 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
Inst << "\t" << Name << " ";
}
uint8_t NumArgs = IR::GetArgs(IROp->Op);
const uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t ArgNode = IROp->Args[i].ID();
uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
const auto ArgNode = IROp->Args[i].ID();
const uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
if (PhysReg >= GPRPairBase)
Inst << "Pair" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else if (PhysReg >= XMMBase)
@@ -739,7 +766,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
LogMan::Msg::DFmt("{}", Inst.str());
}
#endif
uint32_t ID = IR->GetID(CodeNode);
const auto ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
@@ -789,4 +816,9 @@ uint64_t X86JITCore::ExitFunctionLink(X86JITCore *core, FEXCore::Core::CpuStateF
std::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return std::make_unique<X86JITCore>(ctx, Thread, AllocateNewCodeBuffer(ctx, CompileThread ? X86JITCore::MAX_CODE_SIZE : X86JITCore::INITIAL_CODE_SIZE), CompileThread);
}
void InitializeX86JITSignalHandlers(FEXCore::Context::Context *CTX) {
X86JITCore::InitializeSignalHandlers(CTX);
}
}
+31 -17
View File
@@ -9,8 +9,6 @@ $end_info$
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Common/MathUtils.h"
#define XBYAK64
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
@@ -20,6 +18,7 @@ using namespace Xbyak;
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/MathUtils.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <tuple>
@@ -80,6 +79,12 @@ public:
static constexpr size_t MAX_CODE_SIZE = 1024 * 1024 * 256;
void CopyNecessaryDataForCompileThread(CPUBackend *Original) override;
bool IsAddressInJITCode(uint64_t Address, bool IncludeDispatcher = true, bool IncludeCompileService = true) const override {
return Dispatcher->IsAddressInJITCode(Address, IncludeDispatcher, IncludeCompileService);
}
static void InitializeSignalHandlers(FEXCore::Context::Context *CTX);
private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
@@ -88,7 +93,7 @@ private:
std::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
uint64_t Entry;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, Label> JumpTargets;
std::unordered_map<IR::NodeID, Label> JumpTargets;
Xbyak::util::Cpu Features{};
bool MemoryDebug = false;
@@ -114,21 +119,21 @@ private:
constexpr static uint8_t RA_64 = 3;
constexpr static uint8_t RA_XMM = 4;
[[nodiscard]] IR::PhysicalRegister GetPhys(uint32_t Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const;
[[nodiscard]] bool IsFPR(uint32_t Node) const;
[[nodiscard]] bool IsGPR(uint32_t Node) const;
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] Xbyak::Reg GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetSrc(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(uint32_t Node) const;
[[nodiscard]] std::pair<Xbyak::Reg, Xbyak::Reg> GetSrcPair(IR::NodeID Node) const;
template<uint8_t RAType>
[[nodiscard]] Xbyak::Reg GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Reg GetDst(IR::NodeID Node) const;
[[nodiscard]] Xbyak::Xmm GetSrc(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(uint32_t Node) const;
[[nodiscard]] Xbyak::Xmm GetSrc(IR::NodeID Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(IR::NodeID Node) const;
[[nodiscard]] Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
@@ -149,6 +154,9 @@ private:
static uint64_t ExitFunctionLink(X86JITCore* code, FEXCore::Core::CpuStateFrame *Frame, uint64_t *record);
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
// This is the initial code buffer that we will fall back to
// In a program without signals and code clearing, we will typically
// only have this code buffer
@@ -163,6 +171,8 @@ private:
struct CompilerSharedData {
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
FEXCore::CPU::Dispatcher *Dispatcher{};
@@ -177,8 +187,8 @@ private:
std::tuple<SetCC, CMovCC, JCC> GetCC(IR::CondClassType cond);
using OpHandler = void (X86JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
using OpHandler = void (X86JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
void RegisterBranchHandlers();
@@ -192,7 +202,7 @@ private:
void PushRegs();
void PopRegs();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -228,6 +238,8 @@ private:
DEF_OP(Rol);
DEF_OP(Ror);
DEF_OP(Extr);
DEF_OP(PDep);
DEF_OP(PExt);
DEF_OP(LDiv);
DEF_OP(LUDiv);
DEF_OP(LRem);
@@ -269,7 +281,6 @@ private:
///< Branch ops
DEF_OP(GuestCallDirect);
DEF_OP(GuestCallIndirect);
DEF_OP(GuestReturn);
DEF_OP(SignalReturn);
DEF_OP(CallbackReturn);
DEF_OP(ExitFunction);
@@ -310,6 +321,7 @@ private:
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
///< Misc ops
DEF_OP(EndBlock);
@@ -320,6 +332,8 @@ private:
DEF_OP(Print);
DEF_OP(GetRoundingMode);
DEF_OP(SetRoundingMode);
DEF_OP(ProcessorID);
DEF_OP(RDRAND);
///< Move ops
DEF_OP(ExtractElementPair);
@@ -329,8 +343,6 @@ private:
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(CreateVector2);
DEF_OP(CreateVector4);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
@@ -417,6 +429,7 @@ private:
DEF_OP(VSMull2);
DEF_OP(VUABDL);
DEF_OP(VTBL1);
DEF_OP(VRev64);
///< Encryption ops
DEF_OP(AESImc);
@@ -425,6 +438,7 @@ private:
DEF_OP(AESDec);
DEF_OP(AESDecLast);
DEF_OP(AESKeyGenAssist);
DEF_OP(CRC32);
#undef DEF_OP
};
@@ -17,13 +17,13 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
movzx(GetDst<RA_32>(Node), byte [STATE + Op->Offset]);
@@ -84,23 +84,23 @@ DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
mov(byte [STATE + Op->Offset], GetSrc<RA_8>(Op->Header.Args[0].ID()));
mov(byte [STATE + Op->Offset], GetSrc<RA_8>(Op->Value.ID()));
}
break;
case 2: {
mov(word [STATE + Op->Offset], GetSrc<RA_16>(Op->Header.Args[0].ID()));
mov(word [STATE + Op->Offset], GetSrc<RA_16>(Op->Value.ID()));
}
break;
case 4: {
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Value.ID()));
}
break;
case 8: {
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Value.ID()));
}
break;
case 16:
@@ -112,27 +112,27 @@ DEF_OP(StoreContext) {
else {
switch (OpSize) {
case 1: {
pextrb(byte [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()), 0);
pextrb(byte [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
}
break;
case 2: {
pextrw(word [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()), 0);
pextrw(word [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
}
break;
case 4: {
vmovd(dword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
vmovd(dword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
case 8: {
vmovq(qword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
vmovq(qword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
case 16: {
if (Op->Offset % 16 == 0)
movaps(xword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
movaps(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
else
movups(xword [STATE + Op->Offset], GetSrc(Op->Header.Args[0].ID()));
movups(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
@@ -143,9 +143,9 @@ DEF_OP(StoreContext) {
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = IROp->Size;
Reg index = GetSrc<RA_64>(Op->Header.Args[0].ID());
Reg index = GetSrc<RA_64>(Op->Index.ID());
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (Op->Stride) {
case 1:
case 2:
@@ -245,11 +245,11 @@ DEF_OP(LoadContextIndexed) {
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
Reg index = GetSrc<RA_64>(Op->Header.Args[1].ID());
Reg index = GetSrc<RA_64>(Op->Index.ID());
size_t size = IROp->Size;
if (Op->Class.Val == 0) {
auto value = GetSrc<RA_64>(Op->Header.Args[0].ID());
if (Op->Class == IR::GPRClass) {
auto value = GetSrc<RA_64>(Op->Value.ID());
lea(rax, dword [STATE + Op->BaseOffset]);
switch (Op->Stride) {
@@ -269,7 +269,7 @@ DEF_OP(StoreContextIndexed) {
}
}
else {
auto value = GetSrc(Op->Header.Args[0].ID());
auto value = GetSrc(Op->Value.ID());
switch (Op->Stride) {
case 1:
case 2:
@@ -339,19 +339,19 @@ DEF_OP(SpillRegister) {
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
mov(byte [rsp + SlotOffset], GetSrc<RA_8>(Op->Header.Args[0].ID()));
mov(byte [rsp + SlotOffset], GetSrc<RA_8>(Op->Value.ID()));
break;
}
case 2: {
mov(word [rsp + SlotOffset], GetSrc<RA_16>(Op->Header.Args[0].ID()));
mov(word [rsp + SlotOffset], GetSrc<RA_16>(Op->Value.ID()));
break;
}
case 4: {
mov(dword [rsp + SlotOffset], GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov(dword [rsp + SlotOffset], GetSrc<RA_32>(Op->Value.ID()));
break;
}
case 8: {
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Value.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -359,15 +359,15 @@ DEF_OP(SpillRegister) {
} else if (Op->Class == FEXCore::IR::FPRClass) {
switch (OpSize) {
case 4: {
movss(dword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movss(dword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
case 8: {
movsd(qword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movsd(qword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
case 16: {
movaps(xword [rsp + SlotOffset], GetSrc(Op->Header.Args[0].ID()));
movaps(xword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
@@ -435,7 +435,7 @@ DEF_OP(LoadFlag) {
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
mov (rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov (rax, GetSrc<RA_64>(Op->Value.ID()));
mov(byte [STATE + (offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag)], al);
}
@@ -469,7 +469,7 @@ DEF_OP(LoadMem) {
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
auto Dst = GetDst<RA_64>(Node);
switch (IROp->Size) {
@@ -537,19 +537,19 @@ DEF_OP(StoreMem) {
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class.Val == 0) {
if (Op->Class == IR::GPRClass) {
switch (IROp->Size) {
case 1:
mov(byte [MemPtr], GetSrc<RA_8>(Op->Header.Args[1].ID()));
mov(byte [MemPtr], GetSrc<RA_8>(Op->Value.ID()));
break;
case 2:
mov(word [MemPtr], GetSrc<RA_16>(Op->Header.Args[1].ID()));
mov(word [MemPtr], GetSrc<RA_16>(Op->Value.ID()));
break;
case 4:
mov(dword [MemPtr], GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(dword [MemPtr], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
mov(qword [MemPtr], GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(qword [MemPtr], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
@@ -557,22 +557,22 @@ DEF_OP(StoreMem) {
else {
switch (IROp->Size) {
case 1:
pextrb(byte [MemPtr], GetSrc(Op->Header.Args[1].ID()), 0);
pextrb(byte [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
case 2:
pextrw(word [MemPtr], GetSrc(Op->Header.Args[1].ID()), 0);
pextrw(word [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
case 4:
vmovd(dword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
vmovd(dword [MemPtr], GetSrc(Op->Value.ID()));
break;
case 8:
vmovq(qword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
vmovq(qword [MemPtr], GetSrc(Op->Value.ID()));
break;
case 16:
if (IROp->Size == Op->Align)
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
else
movups(xword [MemPtr], GetSrc(Op->Header.Args[1].ID()));
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
}
@@ -595,6 +595,23 @@ DEF_OP(CacheLineClear) {
clflush(ptr [MemReg]);
}
DEF_OP(CacheLineZero) {
auto Op = IROp->C<IR::IROp_CacheLineZero>();
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
// Align by cacheline
mov (TMP1, CPUIDEmu::CACHELINE_SIZE - 1);
andn(TMP1, TMP1, MemReg.cvt64());
xor_(TMP2, TMP2);
using DataType = uint64_t;
// 64-byte cache line zero
for (size_t i = 0; i < CPUIDEmu::CACHELINE_SIZE; i += sizeof(DataType)) {
mov (qword [TMP1 + i], TMP2);
}
}
#undef DEF_OP
void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -615,6 +632,7 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
}
}
+49 -32
View File
@@ -18,15 +18,7 @@ $end_info$
#include <xbyak/xbyak.h>
namespace FEXCore::CPU {
static void PrintValue(uint64_t Value) {
LogMan::Msg::DFmt("Value: 0x{:x}", Value);
}
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
@@ -47,30 +39,24 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case 0: // Hard fault
case 5: // Guest ud2
ud2();
break;
case 1: // Int <imm8>
ud2();
break;
case 2: // overflow
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
ud2();
break;
case 3: // int 1
ud2();
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.OverflowExceptionHandler)]);
break;
case 4: { // HLT
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
break;
}
case 6: // INT3
case FEXCore::IR::Break_Interrupt3: // INT3
{
if (CTX->GetGdbServerStatus()) {
// Adjust the stack first for a regular return
@@ -79,8 +65,7 @@ DEF_OP(Break) {
}
// This jump target needs to be a constant offset here
mov(TMP1, ThreadSharedData.Dispatcher->ThreadPauseHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadPauseHandler)]);
}
else {
// If we don't have a gdb server attached then....crash?
@@ -88,11 +73,20 @@ DEF_OP(Break) {
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
mov(TMP1, ThreadSharedData.Dispatcher->ThreadStopHandlerAddress);
jmp(TMP1);
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.ThreadStopHandler)]);
}
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.UnimplementedInstructionHandler)]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
}
}
@@ -136,21 +130,42 @@ DEF_OP(Print) {
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, reinterpret_cast<uintptr_t>(PrintValue));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintValue)]);
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
mov(rax, reinterpret_cast<uintptr_t>(PrintVectorValue));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.X86.PrintVectorValue)]);
}
call(rax);
PopRegs();
}
DEF_OP(ProcessorID) {
// Cyclecounter in EDX:EAX
// IA32_TSC_AUX in ECX
rdtscp();
mov (GetDst<RA_32>(Node), ecx);
}
DEF_OP(RDRAND) {
auto Op = IROp->C<IR::IROp_RDRAND>();
auto Dst = GetSrcPair<RA_64>(Node);
if (Op->GetReseeded) {
rdrand(Dst.first);
}
else {
rdseed(Dst.first);
}
// In the case of RDRAND or RDSEED returning a valid number then CF = 1, else 0
mov (Dst.second, 0);
setc(Dst.second.cvt8());
}
#undef DEF_OP
void X86JITCore::RegisterMiscHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
@@ -167,6 +182,8 @@ void X86JITCore::RegisterMiscHandlers() {
REGISTER_OP(GETROUNDINGMODE, GetRoundingMode);
REGISTER_OP(SETROUNDINGMODE, SetRoundingMode);
REGISTER_OP(INVALIDATEFLAGS, NoOp);
REGISTER_OP(PROCESSORID, ProcessorID);
REGISTER_OP(RDRAND, RDRAND);
#undef REGISTER_OP
}
}
@@ -15,7 +15,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -42,7 +42,7 @@ DEF_OP(CreateElementPair) {
Xbyak::Reg RegSecond;
Xbyak::Reg RegTmp;
switch (Op->Header.Size) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetSrc<RA_32>(Op->Header.Args[0].ID());
@@ -16,7 +16,7 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VectorZero) {
auto Dst = GetDst(Node);
vpxor(Dst, Dst, Dst);
@@ -67,14 +67,6 @@ DEF_OP(VectorImm) {
}
}
DEF_OP(CreateVector2) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CreateVector4) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(SplatVector) {
auto Op = IROp->C<IR::IROp_SplatVector2>();
uint8_t OpSize = IROp->Size;
@@ -1545,7 +1537,7 @@ DEF_OP(VInsScalarElement) {
DEF_OP(VExtractElement) {
auto Op = IROp->C<IR::IROp_VExtractElement>();
switch (Op->Header.ElementSize) {
switch (Op->Header.Size) {
case 1: {
pextrb(eax, GetSrc(Op->Header.Args[0].ID()), Op->Index);
pinsrb(GetDst(Node), eax, 0);
@@ -1566,7 +1558,7 @@ DEF_OP(VExtractElement) {
pinsrq(GetDst(Node), rax, 0);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.Size); break;
}
}
@@ -2151,13 +2143,79 @@ DEF_OP(VTBL1) {
}
}
DEF_OP(VRev64) {
auto Op = IROp->C<IR::IROp_VDupElement>();
switch (Op->Header.ElementSize) {
case 1: {
mov(rax, 0x00'01'02'03'04'05'06'07); // Lower
vmovq(xmm15, rax);
if (IROp->Size == 16) {
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x08'09'0A'0B'0C'0D'0E'0F); // Upper
pinsrq(xmm15, rcx, 1);
}
else {
// 8byte, upper bits get zero
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
pinsrq(xmm15, rcx, 1);
}
vpshufb(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), xmm15);
break;
}
case 2: {
// Full 16-bit byteswap in each 64-bit element
mov(rax, 0x01'00'03'02'05'04'07'06); // Lower
vmovq(xmm15, rax);
if (IROp->Size == 16) {
mov(rcx, 0x09'08'0B'0A'0D'0C'0F'0E); // Upper
pinsrq(xmm15, rcx, 1);
}
else {
// 8byte, upper bits get zero
// Full 8bit byteswap in each 64-bit element
mov(rcx, 0x80'80'80'80'80'80'80'80); // Upper
pinsrq(xmm15, rcx, 1);
}
vpshufb(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), xmm15);
break;
}
case 4: {
if (IROp->Size == 16) {
vpshufd(GetDst(Node),
GetSrc(Op->Header.Args[0].ID()),
(0b11 << 0) |
(0b10 << 2) |
(0b01 << 4) |
(0b00 << 6));
}
else {
vpshufd(GetDst(Node),
GetSrc(Op->Header.Args[0].ID()),
(0b01 << 0) |
(0b00 << 2) |
(0b11 << 4) | // Last two don't matter, will be overwritten with zero
(0b11 << 6));
// Zero upper 64-bits
mov(rcx, 0);
pinsrq(GetDst(Node), rcx, 1);
}
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
}
#undef DEF_OP
void X86JITCore::RegisterVectorHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(CREATEVECTOR2, CreateVector2);
REGISTER_OP(CREATEVECTOR4, CreateVector4);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
@@ -2246,6 +2304,7 @@ void X86JITCore::RegisterVectorHandlers() {
REGISTER_OP(VSMULL2, VSMull2);
REGISTER_OP(VUABDL, VUABDL);
REGISTER_OP(VTBL1, VTBL1);
REGISTER_OP(VREV64, VRev64);
#undef REGISTER_OP
}
}
+2 -2
View File
@@ -37,11 +37,11 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(PageMemory != -1ULL, "Failed to allocate page memory");
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
+1 -1
View File
@@ -50,7 +50,7 @@ public:
auto InsertPoint =
#endif
BlockList.emplace(Address, (uintptr_t)HostCode);
LOGMAN_THROW_A(InsertPoint.second == true, "Dupplicate block mapping added");
LOGMAN_THROW_A_FMT(InsertPoint.second == true, "Dupplicate block mapping added");
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
CodePages[CurrentPage].push_back(Address);
File diff suppressed because it is too large. Load diff
+653 -49
View File
@@ -5,6 +5,7 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
@@ -13,6 +14,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <cstdint>
#include <fmt/format.h>
#include <map>
#include <stddef.h>
#include <utility>
@@ -26,18 +28,52 @@ class OpDispatchBuilder final : public IREmitter {
friend class FEXCore::IR::Pass;
friend class FEXCore::IR::PassManager;
enum {
FLAGS_OP_NONE, // must rely on x86 flags
FLAGS_OP_CMP, // flags were set by a CMP between flagsOpDest/flagsOpDestSigned and flagsOpSrc/flagsOpSrcSigned with flagsOpSize size
FLAGS_OP_AND, // flags were set by an AND/TEST, flagsOpDest contains the resulting value of flagsOpSize size
FLAGS_OP_FCMP, // flags were set by a ucomis* / comis*
enum class SelectionFlag {
Nothing, // must rely on x86 flags
CMP, // flags were set by a CMP between flagsOpDest/flagsOpDestSigned and flagsOpSrc/flagsOpSrcSigned with flagsOpSize size
AND, // flags were set by an AND/TEST, flagsOpDest contains the resulting value of flagsOpSize size
FCMP, // flags were set by a ucomis* / comis*
};
public:
int flagsOp;
uint8_t flagsOpSize;
OrderedNode* flagsOpDest, *flagsOpSrc;
OrderedNode* flagsOpDestSigned, *flagsOpSrcSigned;
enum class FlagsGenerationType : uint8_t {
TYPE_NONE,
TYPE_ADC,
TYPE_SBB,
TYPE_SUB,
TYPE_ADD,
TYPE_MUL,
TYPE_UMUL,
TYPE_LOGICAL,
TYPE_LSHL,
TYPE_LSHLI,
TYPE_LSHR,
TYPE_LSHRI,
TYPE_ASHR,
TYPE_ASHRI,
TYPE_ROR,
TYPE_RORI,
TYPE_ROL,
TYPE_ROLI,
TYPE_FCMP,
TYPE_BEXTR,
TYPE_BLSI,
TYPE_BLSMSK,
TYPE_BLSR,
TYPE_POPCOUNT,
TYPE_BZHI,
TYPE_TZCNT,
TYPE_LZCNT,
TYPE_BITSELECT,
TYPE_RDRAND,
};
SelectionFlag flagsOp{};
uint8_t flagsOpSize{};
OrderedNode* flagsOpDest{};
OrderedNode* flagsOpSrc{};
OrderedNode* flagsOpDestSigned{};
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
bool ShouldDump {false};
@@ -51,7 +87,7 @@ public:
OrderedNode* GetNewJumpBlock(uint64_t RIP) {
auto it = JumpTargets.find(RIP);
LOGMAN_THROW_A(it != JumpTargets.end(), "Couldn't find block generated for 0x%lx", RIP);
LOGMAN_THROW_A_FMT(it != JumpTargets.end(), "Couldn't find block generated for 0x{:x}", RIP);
return it->second.BlockEntry;
}
@@ -69,7 +105,7 @@ public:
}
void StartNewBlock() {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
}
bool FinishOp(uint64_t NextRIP, bool LastOp) {
@@ -84,6 +120,9 @@ public:
// cmp qword [rdi-8], 0
// jne .label
if (LastOp && !BlockSetRIP) {
// Calculate flags first
CalculateDeferredFlags();
auto it = JumpTargets.find(NextRIP);
if (it == JumpTargets.end()) {
@@ -98,6 +137,11 @@ public:
return true;
}
}
if (LastOp) {
LOGMAN_THROW_A_FMT(IsDeferredFlagsStored(), "FinishOp: Deferred flags weren't generated at end of block");
}
BlockSetRIP = false;
return false;
@@ -230,10 +274,12 @@ public:
void XADDOp(OpcodeArgs);
void PopcountOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
enum Segment {
Segment_FS,
Segment_GS,
enum class Segment {
FS,
GS,
};
template<Segment Seg>
void ReadSegmentReg(OpcodeArgs);
@@ -253,9 +299,14 @@ public:
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorALUROp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorScalarALUOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize, bool Scalar>
void VectorUnaryOp(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void VectorUnaryDuplicateOp(OpcodeArgs);
void MOVQOp(OpcodeArgs);
template<size_t ElementSize>
void PADDQOp(OpcodeArgs);
@@ -324,13 +375,24 @@ public:
template<size_t ElementSize>
void PSIGN(OpcodeArgs);
// BMI Ops
// BMI1 Ops
void ANDNBMIOp(OpcodeArgs);
void BEXTRBMIOp(OpcodeArgs);
void BLSIBMIOp(OpcodeArgs);
void BLSMSKBMIOp(OpcodeArgs);
void BLSRBMIOp(OpcodeArgs);
// BMI2 Ops
void BMI2Shift(OpcodeArgs);
void BZHI(OpcodeArgs);
void MULX(OpcodeArgs);
void PDEP(OpcodeArgs);
void PEXT(OpcodeArgs);
void RORX(OpcodeArgs);
// ADX Ops
void ADXOp(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -431,6 +493,17 @@ public:
template<size_t ElementSize>
void ADDSUBPOp(OpcodeArgs);
void PFNACCOp(OpcodeArgs);
void PFPNACCOp(OpcodeArgs);
void PSWAPDOp(OpcodeArgs);
template<uint8_t CompType>
void VPFCMPOp(OpcodeArgs);
void PI2FWOp(OpcodeArgs);
void PF2IWOp(OpcodeArgs);
void PMULHRWOp(OpcodeArgs);
void PMADDWD(OpcodeArgs);
void PMADDUBSW(OpcodeArgs);
@@ -457,6 +530,8 @@ public:
void FenceOp(OpcodeArgs);
void StoreFenceOrCLFlush(OpcodeArgs);
void CLZeroOp(OpcodeArgs);
void RDTSCPOp(OpcodeArgs);
void PSADBW(OpcodeArgs);
@@ -484,8 +559,12 @@ public:
void MPSADBWOp(OpcodeArgs);
void CRC32(OpcodeArgs);
void UnimplementedOp(OpcodeArgs);
void InvalidOp(OpcodeArgs);
#undef OpcodeArgs
void SetPackedRFLAG(bool Lower8, OrderedNode *Src);
@@ -501,24 +580,35 @@ private:
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, int8_t Align);
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align);
uint8_t GetDstSize(FEXCore::X86Tables::DecodedOp Op) const;
uint8_t GetSrcSize(FEXCore::X86Tables::DecodedOp Op) const;
[[nodiscard]] static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
}
[[nodiscard]] static uint32_t MMBaseOffset() {
return static_cast<uint32_t>(offsetof(Core::CPUState, mm[0][0]));
}
[[nodiscard]] uint8_t GetDstSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint8_t GetSrcSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint32_t GetDstBitSize(X86Tables::DecodedOp Op) const;
[[nodiscard]] uint32_t GetSrcBitSize(X86Tables::DecodedOp Op) const;
template<unsigned BitOffset>
void SetRFLAG(OrderedNode *Value) {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
_StoreFlag(_Bfe(1, 0, Value), BitOffset);
}
void SetRFLAG(OrderedNode *Value, unsigned BitOffset) {
flagsOp = FLAGS_OP_NONE;
flagsOp = SelectionFlag::Nothing;
_StoreFlag(_Bfe(1, 0, Value), BitOffset);
}
@@ -528,32 +618,533 @@ private:
OrderedNode *SelectCC(uint8_t OP, OrderedNode *TrueValue, OrderedNode *FalseValue);
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High);
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High);
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
/**
* @name Deferred RFLAG calculation and generation.
*
* Only handles the six flags that ALU ops typically generate.
* Specifically: CF, PF, AF, ZF, SF, OF
* These six flags are heavily generated through basic ALU ops and balloon the IR if not early eliminated.
* This tracking structure only tracks single blocks and requires RFLAGS calculation at block-ending ops.
* Some flags generating ALU ops only touch part of the registers, In these cases it will do calculation up front.
* This means we still need our IR passes to eliminate all redundant flags accesses but this light OpcodeDispatcher optimization
* doesn't take it to that level.
* @{ */
// Deferred flag generation tracking structure.
// This structure is used to track RFlags from ALU ops for invalidation.
//
// Future ideas: Use an invalidation mask to do partial generation of flags.
// Particularly for the instructions that don't do the full set of flags calculations.
// These instructions currently calculate the deferred RFLAGS immediately then overwrite rflags state.
// RCLSE IR pass will catch and remove redundant rflags stores like this currently.
struct DeferredFlagData {
// What type of flags to generate
FlagsGenerationType Type {FlagsGenerationType::TYPE_NONE};
// Source size of the op
uint8_t SrcSize;
// Every flag generation type has a result
OrderedNode *Res{};
union {
// UMUL, BEXTR, BLSI, BLSMSK, POPCOUNT, TZCNT, LZCNT, BITSELECT, RDRAND
struct {
} NoSource;
// MUL, BLSR, BZHI
struct {
OrderedNode *Src1;
} OneSource;
// Logical, LSHL, LSHR, ASHR, ROR, ROL
struct {
OrderedNode *Src1;
OrderedNode *Src2;
} TwoSource;
// ADC, SBB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
OrderedNode *Src3;
} ThreeSource;
// LSHLI, LSHRI, ASHRI, RORI, ROLI
struct {
OrderedNode *Src1;
uint64_t Imm;
} OneSrcImmediate;
// ADD, SUB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
bool UpdateCF;
} TwoSrcImmediate;
} Sources{};
};
DeferredFlagData CurrentDeferredFlags{};
/**
* @brief Takes the current deferred flag state and stores the result in to RFLAGS.
*
* Once executed there will no longer be any deferred flag state and RFLAGS will have the correct flags in it.
* Necessary to do when leaving a IR block, or if an instruction is doing a partial overwrite of the flags.
*/
void CalculateDeferredFlags(uint32_t FlagsToCalculateMask = ~0U);
/**
* @brief Invalidates the current deferred flags structure.
*
* If the emulated instruction is going to overwrite all of the flags but isn't tracked using the deferred flag system
* then use this function to stop tracking the current active deferred flags.
*/
void InvalidateDeferredFlags() {
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
/**
* @brief Checks if there is any deferred flag state active.
*
* @return True if RFLAGs contains the flags. False if deferred flags is tracking the data.
*/
bool IsDeferredFlagsStored() const {
return CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE;
}
/**
* @name These functions are used by the deferred flag handling while it is calculating and storing flags in to RFLAGs.
* @{ */
void CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High);
void CalculcateFlags_UMUL(OrderedNode *High);
void CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculcateFlags_BEXTR(OrderedNode *Src);
void CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BLSMSK(OrderedNode *Src);
void CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src);
void CalculcateFlags_POPCOUNT(OrderedNode *Src);
void CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src);
void CalculcateFlags_TZCNT(OrderedNode *Src);
void CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src);
void CalculcateFlags_BITSELECT(OrderedNode *Src);
void CalculcateFlags_RDRAND(OrderedNode *Src);
/** @} */
/**
* @name These functions generated deferred RFLAGs tracking.
*
* Depending on the operation it may force a RFLAGs calculation before storing the new deferred state.
* @{ */
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADC,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SBB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SUB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADD,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_MUL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = High,
},
},
};
}
void GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_UMUL,
.SrcSize = GetSrcSize(Op),
.Res = High,
};
}
void GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LOGICAL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Flags need to be used, generate incoming flags first.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ASHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero.
if (Shift == 0) return;
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LSHRI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROL,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
},
};
}
void GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_RORI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
},
};
}
void GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
// Doesn't set all the flags, needs to calculate.
CalculateDeferredFlags();
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ROLI,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSrcImmediate = {
.Src1 = Src1,
.Imm = Shift,
},
}
};
}
void GenerateFlags_FCMP(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_FCMP,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSource = {
.Src1 = Src1,
.Src2 = Src2,
},
}
};
}
void GenerateFlags_BEXTR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BEXTR,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSI,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSMSK(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSMSK,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BLSR(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSR,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_POPCOUNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_POPCOUNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BZHI(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Result, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BZHI,
.SrcSize = GetSrcSize(Op),
.Res = Result,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
void GenerateFlags_TZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_TZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_LZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_BITSELECT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BITSELECT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_RDRAND(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_RDRAND,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
/** @} */
/** @} */
OrderedNode * GetX87Top();
enum X87Tag {
TAG_VALID = 0b00,
TAG_ZERO = 0b01,
TAG_SPECIAL = 0b10,
TAG_EMPTY = 0b11
enum class X87Tag {
Valid = 0b00,
Zero = 0b01,
Special = 0b10,
Empty = 0b11
};
void SetX87TopTag(OrderedNode *Value, uint32_t Tag);
void SetX87TopTag(OrderedNode *Value, X87Tag Tag);
OrderedNode *GetX87FTW(OrderedNode *Value);
void SetX87Top(OrderedNode *Value);
@@ -571,24 +1162,37 @@ private:
bool Multiblock{};
uint64_t Entry;
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, OrderedNode *ssa1, uint8_t Align = 1) {
OrderedNode* _StoreMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *Addr, OrderedNode *Value, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _StoreMemTSO(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _StoreMem(ssa0, ssa1, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
OrderedNode* _LoadMemAutoTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode *ssa0, uint8_t Align = 1) {
if (CTX->Config.TSOEnabled)
return _LoadMemTSO(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
else
return _LoadMem(ssa0, Invalid(), Align, Class, MEM_OFFSET_SXTX, 1, Size);
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
}
void InstallHostSpecificOpcodeHandlers();
};
void InstallOpcodeHandlers(Context::OperatingMode Mode);
}
template <>
struct fmt::formatter<FEXCore::IR::OpDispatchBuilder::FlagsGenerationType> : fmt::formatter<int> {
using Base = fmt::formatter<int>;
// Pass-through the underlying value, so IDs can
// be formatted like any integral value.
template <typename FormatContext>
auto format(const FEXCore::IR::OpDispatchBuilder::FlagsGenerationType& ID, FormatContext& ctx) {
return Base::format(static_cast<int>(ID), ctx);
}
};
@@ -53,7 +53,7 @@ void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t RCON = Op->Src[1].Data.Literal.Value;
auto Res = _VAESKeyGenAssist(Src, RCON);
@@ -39,35 +39,229 @@ constexpr std::array<uint32_t, 17> FlagOffsets = {
};
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode *Src) {
uint8_t NumFlags = FlagOffsets.size();
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
// Calculate flags early.
// Could use InvalidateDeferredFlags() if we had masked invalidation.
// This is only a partial overwrite of flags since OF isn't stored here.
CalculateDeferredFlags();
NumFlags = 5;
}
else {
// We are overwriting all RFLAGS. Invalidate the deferred flag state.
InvalidateDeferredFlags();
}
auto OneConst = _Constant(1);
for (int i = 0; i < NumFlags; ++i) {
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffsets[i])), OneConst);
SetRFLAG(Tmp, FlagOffsets[i]);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
auto Tmp = _And(_Lshr(Src, _Constant(FlagOffset)), OneConst);
SetRFLAG(Tmp, FlagOffset);
}
}
OrderedNode *OpDispatchBuilder::GetPackedRFLAG(bool Lower8) {
// Calculate flags early.
CalculateDeferredFlags();
OrderedNode *Original = _Constant(2);
uint8_t NumFlags = FlagOffsets.size();
size_t NumFlags = FlagOffsets.size();
if (Lower8) {
NumFlags = 5;
}
for (int i = 0; i < NumFlags; ++i) {
OrderedNode *Flag = _LoadFlag(FlagOffsets[i]);
for (size_t i = 0; i < NumFlags; ++i) {
const auto FlagOffset = FlagOffsets[i];
OrderedNode *Flag = _LoadFlag(FlagOffset);
Flag = _Bfe(4, 32, 0, Flag);
Flag = _Lshl(Flag, _Constant(FlagOffsets[i]));
Flag = _Lshl(Flag, _Constant(FlagOffset));
Original = _Or(Original, Flag);
}
return Original;
}
void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
if (CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE) {
// Nothing to do
return;
}
switch (CurrentDeferredFlags.Type) {
case FlagsGenerationType::TYPE_ADC:
CalculcateFlags_ADC(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SBB:
CalculcateFlags_SBB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SUB:
CalculcateFlags_SUB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_ADD:
CalculcateFlags_ADD(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_MUL:
CalculcateFlags_MUL(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_UMUL:
CalculcateFlags_UMUL(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LOGICAL:
CalculcateFlags_Logical(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHL:
CalculcateFlags_ShiftLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHLI:
CalculcateFlags_ShiftLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_LSHR:
CalculcateFlags_ShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHRI:
CalculcateFlags_ShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ASHR:
CalculcateFlags_SignShiftRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ASHRI:
CalculcateFlags_SignShiftRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROR:
CalculcateFlags_RotateRight(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_RORI:
CalculcateFlags_RotateRightImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ROL:
CalculcateFlags_RotateLeft(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_ROLI:
CalculcateFlags_RotateLeftImmediate(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1,
CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_FCMP:
CalculcateFlags_FCMP(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_BEXTR:
CalculcateFlags_BEXTR(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSI:
CalculcateFlags_BLSI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSMSK:
CalculcateFlags_BLSMSK(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSR:
CalculcateFlags_BLSR(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_POPCOUNT:
CalculcateFlags_POPCOUNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BZHI:
CalculcateFlags_BZHI(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_TZCNT:
CalculcateFlags_TZCNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LZCNT:
CalculcateFlags_LZCNT(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BITSELECT:
CalculcateFlags_BITSELECT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_RDRAND:
CalculcateFlags_RDRAND(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_NONE:
default: ERROR_AND_DIE_FMT("Unhandled flags type {}", CurrentDeferredFlags.Type);
}
// Done calculating
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
}
void OpDispatchBuilder::CalculcateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
auto Size = SrcSize * 8;
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -77,7 +271,7 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(Size - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -130,13 +324,15 @@ void OpDispatchBuilder::GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, Orde
case 64:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", Size); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", Size);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
void OpDispatchBuilder::CalculcateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -146,7 +342,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -185,7 +381,7 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (GetSrcSize(Op)) {
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
@@ -198,13 +394,15 @@ void OpDispatchBuilder::GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, Orde
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", GetSrcSize(Op)); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
void OpDispatchBuilder::CalculcateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -214,7 +412,7 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -255,13 +453,13 @@ void OpDispatchBuilder::GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, Orde
auto XorOp2 = _Xor(Res, Src1);
OrderedNode *FinalAnd = _And(XorOp1, XorOp2);
FinalAnd = _Bfe(1, GetSrcSize(Op) * 8 - 1, FinalAnd);
FinalAnd = _Bfe(1, SrcSize * 8 - 1, FinalAnd);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(FinalAnd);
}
}
void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
void OpDispatchBuilder::CalculcateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// AF
{
OrderedNode *AFRes = _Xor(_Xor(Src1, Src2), Res);
@@ -271,7 +469,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -308,7 +506,7 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
OrderedNode *AndOp1 = _And(XorOp1, XorOp2);
switch (GetSrcSize(Op)) {
switch (SrcSize) {
case 1:
AndOp1 = _Bfe(1, 7, AndOp1);
break;
@@ -321,13 +519,15 @@ void OpDispatchBuilder::GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, Orde
case 8:
AndOp1 = _Bfe(1, 63, AndOp1);
break;
default: LOGMAN_MSG_A("Unknown BFESize: %d", GetSrcSize(Op)); break;
default:
LOGMAN_MSG_A_FMT("Unknown BFE size: {}", SrcSize);
break;
}
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(AndOp1);
}
}
void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High) {
// PF/AF/ZF/SF
// Undefined
{
@@ -342,7 +542,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
// CF and OF are set if the result of the operation can't be fit in to the destination register
// If the value can fit then the top bits will be zero
auto SignBit = _Sbfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto SignBit = _Sbfe(1, SrcSize * 8 - 1, Res);
auto SelectOp = _Select(FEXCore::IR::COND_EQ, High, SignBit, _Constant(0), _Constant(1));
@@ -351,7 +551,7 @@ void OpDispatchBuilder::GenerateFlags_MUL(FEXCore::X86Tables::DecodedOp Op, Orde
}
}
void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, OrderedNode *High) {
void OpDispatchBuilder::CalculcateFlags_UMUL(OrderedNode *High) {
// AF/SF/PF/ZF
// Undefined
{
@@ -373,7 +573,7 @@ void OpDispatchBuilder::GenerateFlags_UMUL(FEXCore::X86Tables::DecodedOp Op, Ord
}
}
void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// AF
{
// Undefined
@@ -383,7 +583,7 @@ void OpDispatchBuilder::GenerateFlags_Logical(FEXCore::X86Tables::DecodedOp Op,
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -418,11 +618,11 @@ auto oldflag = GetRFLAG(FEXCore::X86State::flag);\
auto newval = _Select(FEXCore::IR::COND_EQ, cond, _Constant(0), oldflag, newflag);\
SetRFLAG<FEXCore::X86State::flag>(newval);
void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
auto Size = _Constant(GetSrcSize(Op) * 8);
auto Size = _Constant(SrcSize * 8);
auto ShiftAmt = _Sub(Size, Src2);
auto LastBit = _And(_Lshr(Src1, ShiftAmt), _Constant(1));
COND_FLAG_SET(Src2, RFLAG_CF_LOC, LastBit);
@@ -454,7 +654,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
// SF
{
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val = _Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -462,12 +662,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeft(FEXCore::X86Tables::DecodedOp Op
{
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
// When Shift > 1 then OF is undefined
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_ShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -502,7 +702,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
// SF
{
auto val =_Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto val =_Bfe(1, SrcSize * 8 - 1, Res);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, val);
}
@@ -510,12 +710,12 @@ void OpDispatchBuilder::GenerateFlags_ShiftRight(FEXCore::X86Tables::DecodedOp O
{
// Only defined when Shift is 1 else undefined
// OF flag is set if a sign change occurred
auto val = _Bfe(1, GetSrcSize(Op) * 8 - 1, _Xor(Src1, Res));
auto val = _Bfe(1, SrcSize * 8 - 1, _Xor(Src1, Res));
COND_FLAG_SET(Src2, RFLAG_OF_LOC, val);
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
// CF
{
// Extract the last bit shifted in to CF
@@ -550,7 +750,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
COND_FLAG_SET(Src2, RFLAG_SF_LOC, LshrOp);
@@ -562,14 +762,14 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRight(FEXCore::X86Tables::Decoded
}
}
void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
// CF
{
// Extract the last bit shifted in to CF
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - Shift, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, SrcSize * 8 - Shift, Src1));
}
// PF
@@ -598,20 +798,20 @@ void OpDispatchBuilder::GenerateFlags_ShiftLeftImmediate(FEXCore::X86Tables::Dec
// SF
{
auto LshrOp = _Bfe(1, GetSrcSize(Op) * 8 - 1, Res);
auto LshrOp = _Bfe(1, SrcSize * 8 - 1, Res);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
// OF
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
if (Shift == 1) {
auto SourceBit = _Bfe(1, GetSrcSize(Op) * 8 - 1, Src1);
auto SourceBit = _Bfe(1, SrcSize * 8 - 1, Src1);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(SourceBit, LshrOp));
}
}
}
void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -647,7 +847,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -662,7 +862,7 @@ void OpDispatchBuilder::GenerateFlags_SignShiftRightImmediate(FEXCore::X86Tables
}
}
void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// No flags changed if shift is zero
if (Shift == 0) return;
@@ -698,7 +898,7 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// SF
{
auto SignBitConst = _Constant(GetSrcSize(Op) * 8 - 1);
auto SignBitConst = _Constant(SrcSize * 8 - 1);
auto LshrOp = _Lshr(Res, SignBitConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(LshrOp);
@@ -709,13 +909,13 @@ void OpDispatchBuilder::GenerateFlags_ShiftRightImmediate(FEXCore::X86Tables::De
// Only defined when Shift is 1 else undefined
// Is set to the MSB of the original value
if (Shift == 1) {
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, GetSrcSize(Op) * 8 - 1, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src1));
}
}
}
void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
auto NewCF = _Bfe(1, OpSize - 1, Res);
@@ -743,8 +943,8 @@ void OpDispatchBuilder::GenerateFlags_RotateRight(FEXCore::X86Tables::DecodedOp
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = GetSrcSize(Op) * 8;
void OpDispatchBuilder::CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
auto OpSize = SrcSize * 8;
// Extract the last bit shifted in to CF
//auto Size = _Constant(GetSrcSize(Res) * 8);
@@ -773,10 +973,10 @@ void OpDispatchBuilder::GenerateFlags_RotateLeft(FEXCore::X86Tables::DecodedOp O
}
}
void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
auto NewCF = _Bfe(1, OpSize - Shift, Src1);
@@ -795,10 +995,10 @@ void OpDispatchBuilder::GenerateFlags_RotateRightImmediate(FEXCore::X86Tables::D
}
}
void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
void OpDispatchBuilder::CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
if (Shift == 0) return;
auto OpSize = GetSrcSize(Op) * 8;
auto OpSize = SrcSize * 8;
// CF
{
@@ -815,4 +1015,264 @@ void OpDispatchBuilder::GenerateFlags_RotateLeftImmediate(FEXCore::X86Tables::De
}
}
void OpDispatchBuilder::CalculcateFlags_FCMP(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
}
void OpDispatchBuilder::CalculcateFlags_BEXTR(OrderedNode *Src) {
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// CF and OF are defined as being set to zero
//
SetRFLAG<X86State::RFLAG_CF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_OF_LOC>(_Constant(0));
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
SetRFLAG<X86State::RFLAG_AF_LOC>(_Constant(0));
SetRFLAG<X86State::RFLAG_SF_LOC>(_Constant(0));
// PF
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(_Constant(0));
}
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Src, _Constant(0),
_Constant(1), _Constant(0));
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src) {
// Now for the flags:
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
//
// AF and PF are documented as being in an undefined state after
// a BLSI operation, however, we choose to reliably clear them.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Src, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Src, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_BLSMSK(OrderedNode *Src) {
// Now for the flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_ZF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
void OpDispatchBuilder::CalculcateFlags_BLSR(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_EQ,
Src, Zero,
Zero, One);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SignBit = _Constant(SrcSize * 8 - 1);
auto SFOp = _Lshr(Result, SignBit);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_POPCOUNT(OrderedNode *Src) {
// Set ZF
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(Zero);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(Zero);
}
void OpDispatchBuilder::CalculcateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for the flags
auto Bounds = _Constant(SrcSize * 8- 1);
auto Zero = _Constant(0);
auto One = _Constant(1);
SetRFLAG<X86State::RFLAG_OF_LOC>(Zero);
SetRFLAG<X86State::RFLAG_AF_LOC>(Zero);
if (CTX->Config.ABINoPF) {
_InvalidateFlags(1UL << X86State::RFLAG_PF_LOC);
} else {
SetRFLAG<X86State::RFLAG_PF_LOC>(Zero);
}
// ZF
{
auto ZFOp = _Select(IR::COND_EQ,
Result, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZFOp);
}
// CF
{
auto CFOp = _Select(IR::COND_UGT,
Src, Bounds,
One, Zero);
SetRFLAG<X86State::RFLAG_CF_LOC>(CFOp);
}
// SF
{
auto SFOp = _Lshr(Result, Bounds);
SetRFLAG<X86State::RFLAG_SF_LOC>(SFOp);
}
}
void OpDispatchBuilder::CalculcateFlags_TZCNT(OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, 0, Src));
}
void OpDispatchBuilder::CalculcateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src) {
// OF, SF, AF, PF all undefined
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(_Bfe(1, SrcSize * 8 - 1, Src));
}
void OpDispatchBuilder::CalculcateFlags_BITSELECT(OrderedNode *Src) {
// OF, SF, AF, PF, CF all undefined
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
// ZF is set to 1 if the source was zero
auto ZFSelectOp = _Select(FEXCore::IR::COND_EQ,
Src, ZeroConst,
OneConst, ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(ZFSelectOp);
}
void OpDispatchBuilder::CalculcateFlags_RDRAND(OrderedNode *Src) {
// OF, SF, ZF, AF, PF all zero
// CF is set to the incoming source
auto ZeroConst = _Constant(0);
SetRFLAG<X86State::RFLAG_OF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_ZF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<X86State::RFLAG_PF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Src);
}
}
@@ -241,6 +241,8 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 2>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPGT, 8>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VCMPEQ, 2>(OpcodeArgs);
@@ -295,6 +297,24 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 2>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorALUROp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
auto ALUOp = _VAdd(Size, ElementSize, Src, Dest);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
StoreResult(FPRClass, Op, ALUOp, -1);
}
template
void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
@@ -390,14 +410,35 @@ void OpDispatchBuilder::VectorUnaryOp<IR::OP_VABS, 2, false>(OpcodeArgs);
template
void OpDispatchBuilder::VectorUnaryOp<IR::OP_VABS, 4, false>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorUnaryDuplicateOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto ALUOp = _VFSqrt(ElementSize, ElementSize, Src);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
// Duplicate the lower bits
auto Result = _VDupElement(Size, ElementSize, ALUOp, 0);
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRSQRT, 4>(OpcodeArgs);
template
void OpDispatchBuilder::VectorUnaryDuplicateOp<IR::OP_VFRECP, 4>(OpcodeArgs);
void OpDispatchBuilder::MOVQOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// This instruction is a bit special that if the destination is a register then it'll ZEXT the 64bit source to 128bit
if (Op->Dest.IsGPR()) {
const auto gpr = Op->Dest.Data.GPR.GPR;
_StoreContext(FPRClass, 8, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][0]), Src);
_StoreContext(8, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][0]));
auto Const = _Constant(0);
_StoreContext(GPRClass, 8, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][1]), Const);
_StoreContext(8, GPRClass, Const, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][1]));
}
else {
// This is simple, just store the result
@@ -438,14 +479,15 @@ void OpDispatchBuilder::MOVMSKOpOne(OpcodeArgs) {
//TODO: We could remove this VCastFromGOR + VInsGPR pair if we had a VDUPFromGPR instruction that maps directly to AArch64.
auto M = _Constant(0x80'40'20'10'08'04'02'01ULL);
OrderedNode *VMask = _VCastFromGPR(16, 8, M);
VMask = _VInsGPR(16, 8, VMask, M, 1);
auto VCMP = _VCMPLTZ(Src, 16, 1);
auto VAnd = _VAnd(VCMP, VMask, 16, 1);
VMask = _VInsGPR(16, 8, 1, VMask, M);
auto VAdd1 = _VAddP(VAnd, VAnd, 16, 1);
auto VAdd2 = _VAddP(VAdd1, VAdd1, 8, 1);
auto VAdd3 = _VAddP(VAdd2, VAdd2, 8, 1);
auto VCMP = _VCMPLTZ(16, 1, Src);
auto VAnd = _VAnd(16, 1, VCMP, VMask);
auto VAdd1 = _VAddP(16, 1, VAnd, VAnd);
auto VAdd2 = _VAddP(8, 1, VAdd1, VAdd1);
auto VAdd3 = _VAddP(8, 1, VAdd2, VAdd2);
StoreResult(GPRClass, Op, _VExtractToGPR(16, 2, VAdd3, 0), -1);
}
@@ -502,11 +544,11 @@ void OpDispatchBuilder::PSHUFBOp(OpcodeArgs) {
// Bits [6:4] is reserved for 128bit
// Bits [6:3] is reserved for 64bit
if (Size == 8) {
auto MaskVector = _VectorImm(0b1000'0111, Size, 1);
auto MaskVector = _VectorImm(Size, 1, 0b1000'0111);
Src = _VAnd(Size, Size, Src, MaskVector);
}
else {
auto MaskVector = _VectorImm(0b1000'1111, Size, 1);
auto MaskVector = _VectorImm(Size, 1, 0b1000'1111);
Src = _VAnd(Size, Size, Src, MaskVector);
}
auto Res = _VTBL1(Size, Dest, Src);
@@ -515,8 +557,8 @@ void OpDispatchBuilder::PSHUFBOp(OpcodeArgs) {
template<size_t ElementSize, bool HalfSize, bool Low>
void OpDispatchBuilder::PSHUFDOp(OpcodeArgs) {
LOGMAN_THROW_A(ElementSize != 0, "What. No element size?");
auto Size = GetSrcSize(Op);
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
@@ -552,8 +594,8 @@ void OpDispatchBuilder::PSHUFDOp<4, false, true>(OpcodeArgs);
template<size_t ElementSize>
void OpDispatchBuilder::SHUFOp(OpcodeArgs) {
LOGMAN_THROW_A(ElementSize != 0, "What. No element size?");
auto Size = GetSrcSize(Op);
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
@@ -620,14 +662,14 @@ void OpDispatchBuilder::PINSROp(OpcodeArgs) {
Src = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ElementSize, Op->Flags, -1);
}
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Index = Op->Src[1].Data.Literal.Value;
uint8_t NumElements = Size / ElementSize;
Index &= NumElements - 1;
// This maps 1:1 to an AArch64 NEON Op
auto ALUOp = _VInsGPR(Size, ElementSize, Dest, Src, Index);
auto ALUOp = _VInsGPR(Size, ElementSize, Index, Dest, Src);
StoreResult(FPRClass, Op, ALUOp, -1);
}
@@ -641,7 +683,7 @@ template
void OpDispatchBuilder::PINSROp<8>(OpcodeArgs);
void OpDispatchBuilder::InsertPSOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Imm = Op->Src[1].Data.Literal.Value;
uint8_t CountS = (Imm >> 6);
uint8_t CountD = (Imm >> 4) & 0b11;
@@ -670,10 +712,10 @@ void OpDispatchBuilder::InsertPSOp(OpcodeArgs) {
// ZMask happens after insert
if (ZMask == 0xF) {
Dest = _VectorImm(0, 16, 4);
Dest = _VectorImm(16, 4, 0);
}
else if (ZMask) {
auto Zero = _VectorImm(0, 16, 4);
auto Zero = _VectorImm(16, 4, 0);
for (size_t i = 0; i < 4; ++i) {
if (ZMask & (1 << i)) {
Dest = _VInsElement(GetDstSize(Op), 4, i, 0, Dest, Zero);
@@ -689,7 +731,7 @@ void OpDispatchBuilder::PExtrOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Index = Op->Src[1].Data.Literal.Value;
const uint8_t NumElements = Size / ElementSize;
@@ -766,7 +808,7 @@ void OpDispatchBuilder::PSRLDOp(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VUShrS(Size, ElementSize, Dest, Src);
@@ -784,7 +826,7 @@ template<size_t ElementSize>
void OpDispatchBuilder::PSRLI(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t ShiftConstant = Op->Src[1].Data.Literal.Value;
auto Size = GetSrcSize(Op);
@@ -804,7 +846,7 @@ template<size_t ElementSize>
void OpDispatchBuilder::PSLLI(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t ShiftConstant = Op->Src[1].Data.Literal.Value;
auto Size = GetSrcSize(Op);
@@ -830,7 +872,7 @@ void OpDispatchBuilder::PSLL(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VUShlS(Size, ElementSize, Dest, Src);
@@ -854,7 +896,7 @@ void OpDispatchBuilder::PSRAOp(OpcodeArgs) {
OrderedNode *Result{};
// Incoming element size for the shift source is always 8
auto MaxShift = _VectorImm(ElementSize * 8, 8, 8);
auto MaxShift = _VectorImm(8, 8, ElementSize * 8);
Src = _VUMin(8, 8, MaxShift, Src);
Result = _VSShrS(Size, ElementSize, Dest, Src);
@@ -867,7 +909,7 @@ template
void OpDispatchBuilder::PSRAOp<4>(OpcodeArgs);
void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -879,7 +921,7 @@ void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
}
void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -892,7 +934,7 @@ void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
template<size_t ElementSize>
void OpDispatchBuilder::PSRAIOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Shift = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -957,12 +999,11 @@ void OpDispatchBuilder::CVTFPR_To_GPR(OpcodeArgs) {
// Source Element size is determined by instruction
size_t GPRSize = GetDstSize(Op);
size_t ElementSize = SrcElementSize;
if constexpr (HostRoundingMode) {
Src = _Float_ToGPR_S(Src, ElementSize, GPRSize);
Src = _Float_ToGPR_S(GPRSize, SrcElementSize, Src);
}
else {
Src = _Float_ToGPR_ZS(Src, ElementSize, GPRSize);
Src = _Float_ToGPR_ZS(GPRSize, SrcElementSize, Src);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Src, GPRSize, -1);
@@ -985,11 +1026,11 @@ void OpDispatchBuilder::Vector_CVT_Int_To_Float(OpcodeArgs) {
size_t ElementSize = SrcElementSize;
size_t Size = GetDstSize(Op);
if constexpr (Widen) {
Src = _VSXTL(Src, Size, ElementSize);
Src = _VSXTL(Size, ElementSize, Src);
ElementSize <<= 1;
}
Src = _Vector_SToF(Src, Size, ElementSize);
Src = _Vector_SToF(Size, ElementSize, Src);
StoreResult(FPRClass, Op, Src, -1);
}
@@ -1007,15 +1048,15 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Int(OpcodeArgs) {
size_t Size = GetDstSize(Op);
if constexpr (Narrow) {
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
ElementSize >>= 1;
}
if constexpr (HostRoundingMode) {
Src = _Vector_FToS(Src, Size, ElementSize);
Src = _Vector_FToS(Size, ElementSize, Src);
}
else {
Src = _Vector_FToZS(Src, Size, ElementSize);
Src = _Vector_FToZS(Size, ElementSize, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
@@ -1025,6 +1066,8 @@ template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, false>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, false, true>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<4, true, false>(OpcodeArgs);
template
void OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, true>(OpcodeArgs);
@@ -1053,10 +1096,10 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Float(OpcodeArgs) {
size_t Size = GetDstSize(Op);
if constexpr (DstElementSize > SrcElementSize) {
Src = _Vector_FToF(Size, SrcElementSize << 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize << 1, Src, SrcElementSize);
}
else {
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
}
StoreResult(FPRClass, Op, Src, -1);
@@ -1074,12 +1117,12 @@ void OpDispatchBuilder::MMX_To_XMM_Vector_CVT_Int_To_Float(OpcodeArgs) {
size_t ElementSize = SrcElementSize;
size_t DstSize = GetDstSize(Op);
if constexpr (Widen) {
Src = _VSXTL(Src, DstSize, ElementSize);
Src = _VSXTL(DstSize, ElementSize, Src);
ElementSize <<= 1;
}
// Always signed
Src = _Vector_SToF(Src, DstSize, ElementSize);
Src = _Vector_SToF(DstSize, ElementSize, Src);
OrderedNode *Dest{};
if constexpr (Widen) {
@@ -1107,14 +1150,14 @@ void OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int(OpcodeArgs) {
size_t Size = GetDstSize(Op);
// Always narrows
Src = _Vector_FToF(Size, SrcElementSize >> 1, SrcElementSize, Src);
Src = _Vector_FToF(Size, SrcElementSize >> 1, Src, SrcElementSize);
ElementSize >>= 1;
if constexpr (HostRoundingMode) {
Src = _Vector_FToS(Src, Size, ElementSize);
Src = _Vector_FToS(Size, ElementSize, Src);
}
else {
Src = _Vector_FToZS(Src, Size, ElementSize);
Src = _Vector_FToZS(Size, ElementSize, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
@@ -1133,7 +1176,7 @@ void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *MemDest = _LoadContext(GPRSize, offsetof(FEXCore::Core::CPUState, gregs[FEXCore::X86State::REG_RDI]), GPRClass);
OrderedNode *MemDest = _LoadContext(GPRSize, GPRClass, GPROffset(X86State::REG_RDI));
const size_t NumElements = Size / 64;
for (size_t Element = 0; Element < NumElements; ++Element) {
@@ -1187,13 +1230,6 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
OrderedNode *Src2{};
if constexpr (Scalar) {
Src2 = _VExtractElement(GetDstSize(Op), Size, Dest, 0);
}
else {
Src2 = Dest;
}
uint8_t CompType = Op->Src[1].Data.Literal.Value;
OrderedNode *Result{};
@@ -1201,32 +1237,34 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
//auto ALUOp = _VCMPGT(Size, ElementSize, Dest, Src);
switch (CompType) {
case 0x00: case 0x08: case 0x10: case 0x18: // EQ
Result = _VFCMPEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPEQ(Size, ElementSize, Dest, Src);
break;
case 0x01: case 0x09: case 0x11: case 0x19: // LT, GT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
break;
case 0x02: case 0x0A: case 0x12: case 0x1A: // LE, GE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
break;
case 0x03: case 0x0B: case 0x13: case 0x1B: // Unordered
Result = _VFCMPUNO(Size, ElementSize, Src2, Src);
Result = _VFCMPUNO(Size, ElementSize, Dest, Src);
break;
case 0x04: case 0x0C: case 0x14: case 0x1C: // NEQ
Result = _VFCMPNEQ(Size, ElementSize, Src2, Src);
Result = _VFCMPNEQ(Size, ElementSize, Dest, Src);
break;
case 0x05: case 0x0D: case 0x15: case 0x1D: // NLT, NGT(Swapped operand)
Result = _VFCMPLT(Size, ElementSize, Src2, Src);
Result = _VFCMPLT(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x06: case 0x0E: case 0x16: case 0x1E: // NLE, NGE(Swapped operand)
Result = _VFCMPLE(Size, ElementSize, Src2, Src);
Result = _VFCMPLE(Size, ElementSize, Dest, Src);
Result = _VNot(Size, ElementSize, Result);
break;
case 0x07: case 0x0F: case 0x17: case 0x1F: // Ordered
Result = _VFCMPORD(Size, ElementSize, Src2, Src);
Result = _VFCMPORD(Size, ElementSize, Dest, Src);
break;
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
break;
default: LOGMAN_MSG_A("Unknown Comparison type: %d", CompType);
}
if constexpr (Scalar) {
@@ -1266,7 +1304,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
}
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, 2, Mem, FCW, 2);
}
@@ -1292,7 +1330,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, 2, MemLocation, FTW, 2);
}
@@ -1341,7 +1379,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
// If OSFXSR bit in CR4 is not set than FXSAVE /may/ not save the XMM registers
// This is implementation dependent
for (unsigned i = 0; i < 8; ++i) {
OrderedNode *MMReg = _LoadContext(16, offsetof(FEXCore::Core::CPUState, mm[i]), FPRClass);
OrderedNode *MMReg = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, mm[i]));
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 32));
_StoreMem(FPRClass, 16, MemLocation, MMReg, 16);
@@ -1349,7 +1387,7 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *XMMReg = _LoadContext(16, offsetof(FEXCore::Core::CPUState, xmm[i]), FPRClass);
OrderedNode *XMMReg = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[i]));
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
_StoreMem(FPRClass, 16, MemLocation, XMMReg, 16);
@@ -1362,7 +1400,7 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
{
OrderedNode *MemLocation = _Add(Mem, _Constant(2));
@@ -1387,20 +1425,20 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(4));
auto NewFTW = _LoadMem(GPRClass, 2, MemLocation, 2);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
for (unsigned i = 0; i < 8; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 32));
auto MMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, mm[i]), MMReg);
_StoreContext(16, FPRClass, MMReg, offsetof(FEXCore::Core::CPUState, mm[i]));
}
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
auto XMMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, xmm[i]), XMMReg);
_StoreContext(16, FPRClass, XMMReg, offsetof(FEXCore::Core::CPUState, xmm[i]));
}
}
@@ -1425,25 +1463,14 @@ template<size_t ElementSize>
void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Res = _FCmp(Src1, Src2, ElementSize,
OrderedNode *Res = _FCmp(ElementSize, Src1, Src2,
(1 << FCMP_FLAG_EQ) |
(1 << FCMP_FLAG_LT) |
(1 << FCMP_FLAG_UNORDERED));
OrderedNode *HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
OrderedNode *HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
OrderedNode *HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
GenerateFlags_FCMP(Op, Res, Src1, Src2);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
auto ZeroConst = _Constant(0);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
flagsOp = FLAGS_OP_FCMP;
flagsOp = SelectionFlag::FCMP;
flagsOpDest = Src1;
flagsOpSrc = Src2;
flagsOpSize = GetSrcSize(Op);
@@ -1557,8 +1584,8 @@ void OpDispatchBuilder::MOVQ2DQ(OpcodeArgs) {
// This instruction is a bit special in that if the source is MMX then it zexts to 128bit
if constexpr (ToXMM) {
Src = _VMov(Src, 16);
_StoreContext(FPRClass, 16, offsetof(FEXCore::Core::CPUState, xmm[Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0][0]), Src);
Src = _VMov(16, Src);
_StoreContext(16, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0][0]));
}
else {
// This is simple, just store the result
@@ -1647,11 +1674,151 @@ void OpDispatchBuilder::ADDSUBPOp(OpcodeArgs) {
StoreResult(FPRClass, Op, ResAdd, -1);
}
void OpDispatchBuilder::PFNACCOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *ResSubSrc{};
OrderedNode *ResSubDest{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
auto UpperSubSrc = _VExtractElement(Size, 4, Src, 1);
ResSubDest = _VFSub(4, 4, Dest, UpperSubDest);
ResSubSrc = _VFSub(4, 4, Src, UpperSubSrc);
auto Result = _VInsElement(8, 4, 1, 0, ResSubDest, ResSubSrc);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PFPNACCOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *ResAdd{};
OrderedNode *ResSub{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
ResSub = _VFSub(4, 4, Dest, UpperSubDest);
ResAdd = _VFAddP(Size, 4, Src, Src);
auto Result = _VInsElement(8, 4, 1, 0, ResSub, ResAdd);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PSWAPDOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _VRev64(Size, 4, Src);
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::ADDSUBPOp<4>(OpcodeArgs);
template
void OpDispatchBuilder::ADDSUBPOp<8>(OpcodeArgs);
void OpDispatchBuilder::PI2FWOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
// We now need to transpose the lower 16-bits of each element together
// Only needing to move the upper element down in this case
Src = _VInsElement(Size, 2, 1, 2, Src, Src);
// Now we need to sign extend the 16bit value to 32-bit
Src = _VSXTL(Size, 2, Src);
// int32_t to float
Src = _Vector_SToF(Size, 4, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
}
void OpDispatchBuilder::PF2IWOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
// Float to int32_t
Src = _Vector_FToZS(Size, 4, Src);
// We now need to transpose the lower 16-bits of each element together
// Only needing to move the upper element down in this case
Src = _VInsElement(Size, 2, 1, 2, Src, Src);
// Now we need to sign extend the 16bit value to 32-bit
Src = _VSXTL(Size, 2, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, Size, -1);
}
void OpDispatchBuilder::PMULHRWOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Res{};
// Implementation is more efficient for 8byte registers
// Multiplies 4 16bit values in to 4 32bit values
Res = _VSMull(Size * 2, 2, Dest, Src);
//TODO: We could remove this VCastFromGOR + VInsGPR pair if we had a VDUPFromGPR instruction that maps directly to AArch64.
auto M = _Constant(0x0000'8000'0000'8000ULL);
OrderedNode *VConstant = _VCastFromGPR(16, 8, M);
VConstant = _VInsGPR(16, 8, 1, VConstant, M);
Res = _VAdd(Size * 2, 4, Res, VConstant);
// Now shift and narrow to convert 32-bit values to 16bit, storing the top 16bits
Res = _VUShrNI(Size * 2, 4, Res, 16);
StoreResult(FPRClass, Op, Res, -1);
}
template<uint8_t CompType>
void OpDispatchBuilder::VPFCMPOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetDstSize(Op), Op->Flags, -1);
OrderedNode *Result{};
// This maps 1:1 to an AArch64 NEON Op
//auto ALUOp = _VCMPGT(Size, 4, Dest, Src);
LogMan::Msg::DFmt("CompType: {} Size: {}", CompType, Size);
switch (CompType) {
case 0x00: // EQ
Result = _VFCMPEQ(Size, 4, Dest, Src);
break;
case 0x01: // GE(Swapped operand)
Result = _VFCMPLE(Size, 4, Src, Dest);
break;
case 0x02: // GT
Result = _VFCMPGT(Size, 4, Dest, Src);
break;
default:
LOGMAN_MSG_A_FMT("Unknown Comparison type: {}", CompType);
break;
}
StoreResult(FPRClass, Op, Result, -1);
ShouldDump = true;
}
template
void OpDispatchBuilder::VPFCMPOp<0>(OpcodeArgs);
template
void OpDispatchBuilder::VPFCMPOp<1>(OpcodeArgs);
template
void OpDispatchBuilder::VPFCMPOp<2>(OpcodeArgs);
void OpDispatchBuilder::PMADDWD(OpcodeArgs) {
// This is a pretty curious operation
// Does two MADD operations across 4 16bit signed integers and accumulates to 32bit integers in the destination
@@ -1804,7 +1971,7 @@ void OpDispatchBuilder::PMULHRSW(OpcodeArgs) {
// Implementation is more efficient for 8byte registers
Res = _VSMull(Size * 2, 2, Dest, Src);
Res = _VSShrI(Size * 2, 4, Res, 14);
auto OneVector = _VectorImm(1, Size * 2, 4);
auto OneVector = _VectorImm(Size * 2, 4, 1);
Res = _VAdd(Size * 2, 4, Res, OneVector);
Res = _VUShrNI(Size * 2, 4, Res, 1);
}
@@ -1818,7 +1985,7 @@ void OpDispatchBuilder::PMULHRSW(OpcodeArgs) {
ResultLow = _VSShrI(Size, 4, ResultLow, 14);
ResultHigh = _VSShrI(Size, 4, ResultHigh, 14);
auto OneVector = _VectorImm(1, Size, 4);
auto OneVector = _VectorImm(Size, 4, 1);
ResultLow = _VAdd(Size, 4, ResultLow, OneVector);
ResultHigh = _VAdd(Size, 4, ResultHigh, OneVector);
@@ -2073,10 +2240,10 @@ void OpDispatchBuilder::ExtendVectorElements(OpcodeArgs) {
CurrentElementSize != DstElementSize;
CurrentElementSize <<= 1) {
if constexpr (Signed) {
Result = _VSXTL(Result, Size, CurrentElementSize);
Result = _VSXTL(Size, CurrentElementSize, Result);
}
else {
Result = _VUXTL(Result, Size, CurrentElementSize);
Result = _VUXTL(Size, CurrentElementSize, Result);
}
}
StoreResult(FPRClass, Op, Result, -1);
@@ -2110,7 +2277,7 @@ void OpDispatchBuilder::ExtendVectorElements<4, 8, true>(OpcodeArgs);
template<size_t ElementSize, bool Scalar>
void OpDispatchBuilder::VectorRound(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint64_t Mode = Op->Src[1].Data.Literal.Value;
uint64_t RoundControlSource = (Mode >> 2) & 1;
uint64_t RoundControl = Mode & 0b11;
@@ -2130,7 +2297,7 @@ void OpDispatchBuilder::VectorRound(OpcodeArgs) {
FEXCore::IR::Round_Host,
};
Src = _Vector_FToI(Src, SourceModes[(RoundControlSource << 2) | RoundControl], Size, ElementSize);
Src = _Vector_FToI(Size, ElementSize, Src, SourceModes[(RoundControlSource << 2) | RoundControl]);
if constexpr (Scalar) {
// Insert the lower bits
@@ -2155,7 +2322,7 @@ void OpDispatchBuilder::VectorRound<8, true>(OpcodeArgs);
template<size_t ElementSize>
void OpDispatchBuilder::VectorBlend(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Select = Op->Src[1].Data.Literal.Value;
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -2184,13 +2351,14 @@ void OpDispatchBuilder::VectorVariableBlend(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// The mask is hardcoded to be xmm0 in this instruction
OrderedNode *Mask = _LoadContext(16, offsetof(FEXCore::Core::CPUState, xmm[0]), FPRClass);
OrderedNode *Mask = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
// Each element is selected by the high bit of that element size
// Dest[ElementIdx] = Xmm0[ElementIndex][HighBit] ? Src : Dest;
//
// To emulate this on AArch64
// Arithmetic shift right by the element size, then use BSL to select the registers
Mask = _VSShrI(Size, ElementSize, Mask, (ElementSize * 8) - 1);
auto Result = _VBSL(Mask, Src, Dest);
StoreResult(FPRClass, Op, Result, -1);
@@ -2203,13 +2371,16 @@ template
void OpDispatchBuilder::VectorVariableBlend<8>(OpcodeArgs);
void OpDispatchBuilder::PTestOp(OpcodeArgs) {
// Invalidate deferred flags early
InvalidateDeferredFlags();
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Test1 = _VAnd(Dest, Src, Size, 1);
OrderedNode *Test2 = _VBic(Src, Dest, Size, 1);
OrderedNode *Test1 = _VAnd(Size, 1, Dest, Src);
OrderedNode *Test2 = _VBic(Size, 1, Src, Dest);
Test1 = _VPopcount(Size, 1, Test1);
Test2 = _VPopcount(Size, 1, Test2);
@@ -2231,8 +2402,14 @@ void OpDispatchBuilder::PTestOp(OpcodeArgs) {
Test2 = _Select(FEXCore::IR::COND_EQ,
Test2, ZeroConst, OneConst, ZeroConst);
// Careful, these flags are different between {V,}PTEST and VTESTP{S,D}
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(Test1);
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(Test2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_SF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(ZeroConst);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(ZeroConst);
}
void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
@@ -2265,17 +2442,17 @@ void OpDispatchBuilder::PHMINPOSUWOp(OpcodeArgs) {
}
// Insert the minimum in to bits [15:0]
OrderedNode *Result = _VMov(Min, 2);
OrderedNode *Result = _VMov(2, Min);
// Insert position in to bits [18:16]
Result = _VInsGPR(16, 2, Result, Pos, 1);
Result = _VInsGPR(16, 2, 1, Result, Pos);
StoreResult(FPRClass, Op, Result, -1);
}
template<size_t ElementSize>
void OpDispatchBuilder::DPPOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Mask = Op->Src[1].Data.Literal.Value;
uint8_t SrcMask = Mask >> 4;
uint8_t DstMask = Mask & 0xF;
@@ -2323,7 +2500,7 @@ template
void OpDispatchBuilder::DPPOp<8>(OpcodeArgs);
void OpDispatchBuilder::MPSADBWOp(OpcodeArgs) {
LOGMAN_THROW_A(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
uint8_t Select = Op->Src[1].Data.Literal.Value;
// Src1 needs to be in byte offset
+124 -120
View File
@@ -10,6 +10,7 @@ $end_info$
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/X86Tables.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
@@ -24,26 +25,26 @@ class OrderedNode;
OrderedNode *OpDispatchBuilder::GetX87Top() {
// Yes, we are storing 3 bits in a single flag register.
// Deal with it
return _LoadContext(1, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC, GPRClass);
return _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, uint32_t Tag) {
void OpDispatchBuilder::SetX87TopTag(OrderedNode *Value, X87Tag Tag) {
// if we are popping then we must first mark this location as empty
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
Mask = _Lshl(Mask, TopOffset);
OrderedNode *NewFTW = _Andn(FTW, Mask);
if (Tag != 0) {
auto TagVal = _Lshl(_Constant(Tag), TopOffset);
if (Tag != X87Tag::Valid) {
auto TagVal = _Lshl(_Constant(ToUnderlying(Tag)), TopOffset);
NewFTW = _Or(NewFTW, TagVal);
}
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
OrderedNode *OpDispatchBuilder::GetX87FTW(OrderedNode *Value) {
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
OrderedNode *Mask = _Constant(0b11);
auto TopOffset = _Lshl(Value, _Constant(1));
auto NewFTW = _Lshr(FTW, TopOffset);
@@ -51,7 +52,7 @@ OrderedNode *OpDispatchBuilder::GetX87FTW(OrderedNode *Value) {
}
void OpDispatchBuilder::SetX87Top(OrderedNode *Value) {
_StoreContext(GPRClass, 1, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC, Value);
_StoreContext(1, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
template<size_t width>
@@ -72,7 +73,7 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
// Implicit arg
auto offset = _Constant(Op->OP & 7);
data = _And(_Add(orig_top, offset), mask);
data = _LoadContextIndexed(data, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
data = _LoadContextIndexed(data, 16, MMBaseOffset(), 16, FPRClass);
}
OrderedNode *converted = data;
@@ -82,10 +83,10 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
}
auto top = _And(_Sub(orig_top, _Constant(1)), mask);
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
// Write to ST[TOP]
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
//_StoreContext(converted, 16, offsetof(FEXCore::Core::CPUState, mm[7][0]));
}
@@ -101,25 +102,25 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
auto orig_top = GetX87Top();
auto mask = _Constant(7);
auto top = _And(_Sub(orig_top, _Constant(1)), mask);
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
// Read from memory
OrderedNode *data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags, -1);
OrderedNode *converted = _F80BCDLoad(data);
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *converted = _F80BCDStore(data);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
}
@@ -129,15 +130,15 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
// Update TOP
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto low = _Constant(Lower);
auto high = _Constant(Upper);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
// Write to ST[TOP]
_StoreContextIndexed(data, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -159,7 +160,7 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
// Update TOP
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
size_t read_width = GetSrcSize(Op);
@@ -190,24 +191,24 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
converted = _VInsElement(16, 8, 1, 0, converted, _VCastFromGPR(16, 8, upper));
// Write to ST[TOP]
_StoreContextIndexed(converted, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
}
template<size_t width>
void OpDispatchBuilder::FST(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
if constexpr (width == 80) {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, data, 10, 1);
}
else if constexpr (width == 32 || width == 64) {
auto result = _F80CVT(data, width / 8);
auto result = _F80CVT(width / 8, data);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, result, width / 8, 1);
}
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
// Set the new top now
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
@@ -226,14 +227,14 @@ void OpDispatchBuilder::FIST(OpcodeArgs) {
auto Size = GetSrcSize(Op);
auto orig_top = GetX87Top();
OrderedNode *data = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
data = _F80CVTInt(data, Truncate, Size);
OrderedNode *data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
data = _F80CVTInt(Size, data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, 1);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
// Set the new top now
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
@@ -274,22 +275,22 @@ void OpDispatchBuilder::FADD(OpcodeArgs) {
if constexpr (ResInST0 == OpResult::RES_STI) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Add(a, b);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -336,23 +337,23 @@ void OpDispatchBuilder::FMUL(OpcodeArgs) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Mul(a, b);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -399,10 +400,10 @@ void OpDispatchBuilder::FDIV(OpcodeArgs) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
if constexpr (reverse) {
@@ -414,14 +415,14 @@ void OpDispatchBuilder::FDIV(OpcodeArgs) {
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -483,10 +484,10 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
if constexpr (ResInST0 == OpResult::RES_STI) {
StackLocation = arg;
}
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
if constexpr (reverse) {
@@ -498,7 +499,7 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
@@ -506,7 +507,7 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
}
// Write to ST[TOP]
_StoreContextIndexed(result, StackLocation, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, StackLocation, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -541,37 +542,37 @@ void OpDispatchBuilder::FSUB<32, true, true, OpDispatchBuilder::OpResult::RES_ST
void OpDispatchBuilder::FCHS(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(0);
auto high = _Constant(0b1'000'0000'0000'0000ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
auto result = _VXor(a, data, 16, 1);
auto result = _VXor(16, 1, a, data);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FABS(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(~0ULL);
auto high = _Constant(0b0'111'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
auto result = _VAnd(a, data, 16, 1);
auto result = _VAnd(16, 1, a, data);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FTST(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(0);
OrderedNode *data = _VCastFromGPR(16, 8, low);
@@ -595,35 +596,35 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
void OpDispatchBuilder::FRNDINT(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Round(a);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FXTRACT(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto exp = _F80XTRACT_EXP(a);
auto sig = _F80XTRACT_SIG(a);
// Write to ST[TOP]
_StoreContextIndexed(exp, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(sig, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(exp, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(sig, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037
auto NewFCW = _Constant(16, 0x037);
// Init FCW to 0x037F
auto NewFCW = _Constant(16, 0x037F);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Init FSW to 0
SetX87Top(_Constant(0));
@@ -634,7 +635,7 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(_Constant(0));
// Tags all get set to 0b11
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), _Constant(0xFFFF));
_StoreContext(2, GPRClass, _Constant(0xFFFF), offsetof(FEXCore::Core::CPUState, FTW));
}
template<size_t width, bool Integer, OpDispatchBuilder::FCOMIFlags whichflags, bool poptwice>
@@ -661,10 +662,10 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
// Implicit arg
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
}
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *Res = _F80Cmp(a, b,
(1 << FCMP_FLAG_EQ) |
@@ -684,24 +685,27 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
}
else {
// Invalidate deferred flags early
// OF, SF, AF, PF all undefined
InvalidateDeferredFlags();
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(HostFlag_CF);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_LOC>(HostFlag_ZF);
SetRFLAG<FEXCore::X86State::RFLAG_PF_LOC>(HostFlag_Unordered);
}
if constexpr (poptwice) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
top = _And(_Add(top, _Constant(1)), mask);
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
}
else if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
// Set the new top now
top = _And(_Add(top, _Constant(1)), mask);
SetX87Top(top);
@@ -738,12 +742,12 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
// Write to ST[TOP]
_StoreContextIndexed(b, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(b, top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::FST(OpcodeArgs) {
@@ -756,14 +760,14 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
// Write to ST[TOP]
_StoreContextIndexed(a, arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
top = _And(_Add(top, _Constant(1)), _Constant(7));
SetX87Top(top);
}
@@ -772,14 +776,14 @@ void OpDispatchBuilder::FST(OpcodeArgs) {
template<FEXCore::IR::IROps IROp>
void OpDispatchBuilder::X87UnaryOp(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Round(a);
// Overwrite the op
result.first->Header.Op = IROp;
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -798,8 +802,8 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
auto mask = _Constant(7);
OrderedNode *st1 = _And(_Add(top, _Constant(1)), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
st1 = _LoadContextIndexed(st1, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
st1 = _LoadContextIndexed(st1, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80Add(a, st1);
// Overwrite the op
@@ -811,7 +815,7 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
}
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
template
@@ -842,17 +846,17 @@ void OpDispatchBuilder::X87ModifySTP<true>(OpcodeArgs);
void OpDispatchBuilder::X87SinCos(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto sin = _F80SIN(a);
auto cos = _F80COS(a);
// Write to ST[TOP]
_StoreContextIndexed(sin, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(cos, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(sin, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(cos, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
@@ -860,61 +864,61 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
auto orig_top = GetX87Top();
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
OrderedNode *st0 = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st0 = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
if (Plus1) {
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
st0 = _F80Add(st0, data);
}
auto result = _F80FYL2X(st0, st1);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87TAN(OpcodeArgs) {
auto orig_top = GetX87Top();
auto top = _And(_Sub(orig_top, _Constant(1)), _Constant(7));
SetX87TopTag(top, TAG_VALID);
SetX87TopTag(top, X87Tag::Valid);
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80TAN(a);
auto low = _Constant(0x8000'0000'0000'0000ULL);
auto high = _Constant(0b0'011'1111'1111'1111ULL);
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, data, high, 1);
data = _VInsGPR(16, 8, 1, data, high);
// Write to ST[TOP]
_StoreContextIndexed(result, orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(data, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87ATAN(OpcodeArgs) {
auto orig_top = GetX87Top();
// if we are popping then we must first mark this location as empty
SetX87TopTag(orig_top, TAG_EMPTY);
SetX87TopTag(orig_top, X87Tag::Empty);
auto top = _And(_Add(orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto result = _F80ATAN(st1, a);
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
@@ -924,7 +928,7 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
@@ -947,7 +951,7 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto NewFTW = _LoadMem(GPRClass, Size, MemLocation, Size);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
}
@@ -976,7 +980,7 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
@@ -1004,7 +1008,7 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, Size, MemLocation, FTW, Size);
}
@@ -1036,11 +1040,11 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
OrderedNode *NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, -1);
}
@@ -1107,7 +1111,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
OrderedNode *Top = GetX87Top();
{
auto FCW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FCW), GPRClass);
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
@@ -1134,7 +1138,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
{
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto FTW = _LoadContext(2, offsetof(FEXCore::Core::CPUState, FTW), GPRClass);
auto FTW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FTW));
_StoreMem(GPRClass, Size, MemLocation, FTW, Size);
}
@@ -1168,14 +1172,14 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
auto SevenConst = _Constant(7);
auto TenConst = _Constant(10);
for (int i = 0; i < 7; ++i) {
auto data = _LoadContextIndexed(Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
_StoreMem(FPRClass, 16, ST0Location, data, 1);
ST0Location = _Add(ST0Location, TenConst);
Top = _And(_Add(Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
auto data = _LoadContextIndexed(Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
@@ -1195,7 +1199,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
_F80LoadFCW(NewFCW);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FCW), NewFCW);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
@@ -1218,7 +1222,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
// FTW
OrderedNode *MemLocation = _Add(Mem, _Constant(Size * 2));
auto NewFTW = _LoadMem(GPRClass, Size, MemLocation, Size);
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), NewFTW);
_StoreContext(2, GPRClass, NewFTW, offsetof(FEXCore::Core::CPUState, FTW));
}
OrderedNode *ST0Location = _Add(Mem, _Constant(Size * 7));
@@ -1230,14 +1234,14 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
OrderedNode *Mask = _VCastFromGPR(16, 8, low);
Mask = _VInsGPR(16, 8, Mask, high, 1);
Mask = _VInsGPR(16, 8, 1, Mask, high);
for (int i = 0; i < 7; ++i) {
OrderedNode *Reg = _LoadMem(FPRClass, 16, ST0Location, 1);
// Mask off the top bits
Reg = _VAnd(16, 16, Reg, Mask);
_StoreContextIndexed(Reg, Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
ST0Location = _Add(ST0Location, TenConst);
Top = _And(_Add(Top, OneConst), SevenConst);
@@ -1252,12 +1256,12 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
ST0Location = _Add(ST0Location, _Constant(8));
OrderedNode *RegHigh = _LoadMem(FPRClass, 2, ST0Location, 1);
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
_StoreContextIndexed(Reg, Top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
OrderedNode *Result = _VExtractToGPR(16, 8, a, 1);
// Extract the sign bit
@@ -1333,7 +1337,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
Type = COMPARE_ZERO;
break;
default:
LOGMAN_MSG_A("Unhandled FCMOV op: 0x%x", Opcode);
LOGMAN_MSG_A_FMT("Unhandled FCMOV op: 0x{:x}", Opcode);
break;
}
@@ -1358,7 +1362,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
SrcCond = _Sbfe(1, 0, SrcCond);
OrderedNode *VecCond = _VCastFromGPR(16, 8, SrcCond);
VecCond = _VInsGPR(16, 8, VecCond, SrcCond, 1);
VecCond = _VInsGPR(16, 8, 1, VecCond, SrcCond);
auto top = GetX87Top();
OrderedNode* arg;
@@ -1369,17 +1373,17 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
auto offset = _Constant(Op->OP & 7);
arg = _And(_Add(top, offset), mask);
auto a = _LoadContextIndexed(top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
auto Result = _VBSL(VecCond, b, a);
// Write to ST[TOP]
_StoreContextIndexed(Result, top, 16, offsetof(FEXCore::Core::CPUState, mm[0][0]), 16, FPRClass);
_StoreContextIndexed(Result, top, 16, MMBaseOffset(), 16, FPRClass);
}
void OpDispatchBuilder::X87EMMS(OpcodeArgs) {
// Tags all get set to 0b11
_StoreContext(GPRClass, 2, offsetof(FEXCore::Core::CPUState, FTW), _Constant(0xFFFF));
_StoreContext(2, GPRClass, _Constant(0xFFFF), offsetof(FEXCore::Core::CPUState, FTW));
}
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
@@ -1391,7 +1395,7 @@ void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
top = _And(_Add(top, offset), _Constant(7));
// Set this argument's tag as empty now
SetX87TopTag(top, TAG_EMPTY);
SetX87TopTag(top, X87Tag::Empty);
}
}
Loaded 100 of 538 files, more files were not shown because too many files have changed in this diff. Show more