Compare commits

..
238 Commits
Author SHA1 Message Date
Stefanos Kornilios Mitsis Poiitidis 3978097f15 Merge pull request #705 from FEX-Emu/skmp/mmap-cache-flushing
Flush IR/Code cache on mmap, mmunmap & mprot
2021-03-03 12:59:29 +02:00
Stefanos Kornilios Mitsis Poiitidis 613e5b66ba SMC: Expand --smc-checks options to 'none', 'mman', 'full' 2021-03-03 12:49:23 +02:00
Ryan Houdek e95d1bd761 Merge pull request #811 from phire/remove-operator-names
X64_64 JIt: Remove -fno-operator-names
2021-03-02 18:55:53 -08:00
Scott Mansell ec95610587 X64_64 JIt: Remove -fno-operator-names
Some IDEs didn't recognize this flag and threw
phantom errors when these files were open.

Lets lean towards a smoother developer experience
2021-03-03 15:47:58 +13:00
Ryan Houdek b57cf78cc7 Merge pull request #810 from phire/fix_include
Add missing unordered_map include
2021-03-02 18:00:39 -08:00
Scott Mansell f5ac5d1667 Add missing unordered_map include
It's really unhelpful that libstdc++ includes this by default
2021-03-03 14:06:16 +13:00
Stefanos Kornilios Mitsis Poiitidis e92578817d SMC: Flush IR/Code cache on mmap, mmunmap & mprot 2021-02-26 13:22:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 5b4696976d Merge pull request #805 from FEX-Emu/skmp/asm-thunks
Thunks: Convert guest thunk to asm, to avoid issues with older gcc versions
2021-02-26 12:57:14 +02:00
Stefanos Kornilios Mitsis Poiitidis aecea294d5 Merge pull request #792 from Sonicadvance1/implemented_unaligned_memory_ops
Implements unaligned atomic memory ops for ARMv8.1+
2021-02-26 12:50:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 70a98cd465 Thunks: Convert guest thunk to asm, to avoid issues with older gcc verions 2021-02-26 12:43:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 2071bab238 Merge pull request #803 from Sonicadvance1/faccessat2
Implements support for faccessat2 in syscallhandler
2021-02-25 23:58:11 +02:00
Ryan Houdek 9746e0175e Change unsupported faccessat2 to use safe syscall unsupported 2021-02-25 13:45:24 -08:00
Ryan Houdek d78a3784fb Guards SYS_faccessat2 define for new enough glibc defines 2021-02-25 13:44:40 -08:00
Ryan Houdek 7dd5252d68 Implements support for faccessat2 in syscallhandler
Only exists if the host kernel is >= 5.8.0.
2021-02-24 09:52:30 -08:00
Ryan Houdek ac92a103df Adds support for faccessat2 to FileManager 2021-02-24 09:51:29 -08:00
Ryan Houdek 74a00fee0f Updates syscall definitions enums 2021-02-24 09:51:03 -08:00
Ryan Houdek ebe447f86a Merge pull request #801 from Sonicadvance1/fix_cmpxchg_flags
Fixes CMPXCHG flags being incorrect aside from ZF
2021-02-24 09:01:01 -08:00
Ryan Houdek ebac071ae1 Adds cmpxchg to register zext and flags unit test 2021-02-23 23:22:24 -08:00
Ryan Houdek 68db163efc Adds cmpxchg to memory zext and flag unit test 2021-02-23 23:22:00 -08:00
Ryan Houdek 6c469c21db Fixes Zext and flags behaviour of CMPXCHG to register 2021-02-23 23:21:28 -08:00
Ryan Houdek 9798522bc5 Fixes Zext behaviour of CMPXCHG to memory 2021-02-23 23:21:05 -08:00
Ryan Houdek 55ddc5b6e4 Adds unit tests to ensure cmpxchg flag correctness 2021-02-23 22:02:39 -08:00
Ryan Houdek 373f4275f6 Fixes CMPXCHG flags being incorrect aside from ZF
Almost everything only checks ZF but we had the arguments reversed for
the rest of the comparison flag results.
2021-02-23 22:01:52 -08:00
Scott Mansell 6fee5b9ca2 Merge pull request #789 from FEX-Emu/skmp/add-cacheinfo-cpuid
CPUID: Add cache information, function 0x2
2021-02-24 03:06:20 +13:00
Stefanos Kornilios Mitsis Poiitidis aaa63a2409 Merge pull request #730 from FEX-Emu/skmp/ir-cache
AOTIR: Initial implementation
2021-02-23 13:10:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 6d0d3eaf04 AOTIR: Review feedback 2021-02-23 12:23:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 075cd423ed AOTIR: Review feedback 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 8e06966ddc AOTIR: Make OP_REMOVECODEENTRY relocatable 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis bdcf3ff606 AOTIR: Merge IRLists, RALists and DebugDataLists to LocalIRCache; cleanups and fixups 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b29121c71a AOTIR: Rename AOTCache to AOTIRCache 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis a5ee4bdf3c AOTIR: Fix double free issues 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis abb2e09c20 AOTIR: Add support for 32-bit process 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 56d10dcbbb IR: Make OP_THUNK use an inline sha256 hash of the thunk name, update thunk scripts 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 1f919949cc JIT: Fix arm64 build 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 63507f5ece IR: Fix int64_t parsing 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis dbc8b8e8e9 AOTIR: Append optimization flags to fileid 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 681deca4eb IR: Make entrypoint implicit, Add InlineEntrypointOffset, make ValidateCode Entrypoint-relative 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b3d12fbb7b AOTIR: Cleanup interface 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fafc987eab AOTIR: Rename Generate to Capture 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis d825ed2b03 AOTIR: Reduce map lookup 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis f2f62c5f2c AOTIR: Implement EntrypointOffset for aarch64 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 3e1e093b24 AOTIR: Store in ~/.fex-emu/aotir, per so 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 71dceb5339 Fix relocation support for FinishOp 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis be40d1604a Improve aotir format 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 47d4481a75 AOTIR: Support relocations via new ir op 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis fc30efe040 AOTIR: Add hashing of code bytes 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 57680a8614 AOTIR: Add --aotir-generate and --aotir-load to FEXLoader & Config 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 66bdface49 AOTIR: Make copies for insertions, only cache RA'd blocks 2021-02-23 12:08:57 +02:00
Stefanos Kornilios Mitsis Poiitidis 816c4656df IR: Add AOT Cache 2021-02-23 12:08:57 +02:00
Scott Mansell 7b97c3fa23 Merge pull request #755 from Sonicadvance1/host_uname
Pulls uname nodename from host system
2021-02-23 23:07:29 +13:00
Stefanos Kornilios Mitsis Poiitidis 8642406b8f CPUID: Add cache information, function 0x2 2021-02-23 11:01:15 +02:00
Stefanos Kornilios Mitsis Poiitidis e2599db3ed Merge pull request #793 from Sonicadvance1/assert_on_missing_lock
Adds assert checks on missing LOCK support
2021-02-23 10:51:30 +02:00
Ryan Houdek bb0850eab8 Adds assert checks on missing LOCK support
These usually don't get hit, but Geekbench4 DOES manage to hit LOCK on
BTS.

Which we just don't support right now.
2021-02-22 23:20:18 -08:00
Stefanos Kornilios Mitsis Poiitidis 1af541475a Merge pull request #795 from Sonicadvance1/atomic_bittest_ops
Implements BTC, BTR, BTS atomic variants
2021-02-23 09:08:39 +02:00
Stefanos Kornilios Mitsis Poiitidis 6669bd2d4a Merge pull request #794 from Sonicadvance1/fix_log_move_fail
Pass log moves on buildbot stage failure
2021-02-23 08:59:21 +02:00
Ryan Houdek d3e98fedb4 Pulls uname nodename from host system
Hardcoding FEXCore as the nodename is an annoyance.
Pull the host's nodename instead.

Fixes #600
2021-02-22 22:39:05 -08:00
Ryan Houdek 8a39f4b25c Implements BTC, BTR, BTS atomic unit tests
This just takes the regular non-atomic unit tests and changes them to
have lock prefixes.
These are all handled as byte sized atomics so there aren't any
alignment problems.
2021-02-22 22:29:31 -08:00
Ryan Houdek 232eeff483 Implements BTC, BTR, BTS atomic variants
BTS specifically is being used for threading related tasks. Which could
be why our threading has been unstable.
2021-02-22 22:27:10 -08:00
Ryan Houdek 12ef69650f Pass log moves on buildbot stage failure
Passes the log movement stage to clean up the output.
Really it is only the first unit test stage that is failing, but each
subsequent log run will have failed to move since the previous stage
didn't run.

Makes it easier to scan over a failure as only the first unit test
failure step.
2021-02-22 18:42:28 -08:00
Ryan Houdek f1349becb4 Disables the unaligned atomic memory op tests
These don't run on the armv8.0 runner
2021-02-22 18:33:17 -08:00
Ryan Houdek f848957e17 Implements unit tests for the new unaligned atomics
Tests all the unaligned atomic ops we support now
2021-02-22 18:25:25 -08:00
Ryan Houdek 55fc47e0b7 Passes SIGBUS handler to unaligned atomic op handler for ARMv8.1
ARMv8.0 still unsupported here.

This allows us to support unaligned atomic ops for every atomic x86 op
that we support.
2021-02-22 18:22:10 -08:00
Ryan Houdek 4f40166902 Implements unaligned atomic memory ops for ARMv8.1+
This takes the existing unaligned CAS handling code and makes it more
robust to handle both true unaligned CAS and unaligned atomic memory
ops.

Specifically it inserts some lambda helpers to calculate the Desired and
Expected values inside of the CAS loops to account for CAS and memory
operation.
2021-02-22 18:17:53 -08:00
Stefanos Kornilios Mitsis Poiitidis fba626c547 Merge pull request #790 from FEX-Emu/skmp/workaround-exit-group
Threading: Workaround exit_group bug
2021-02-22 17:50:59 +02:00
Stefanos Kornilios Mitsis Poiitidis 9b360e66dc Threading: Workaround exit_group bug 2021-02-22 14:24:42 +02:00
Stefanos Kornilios Mitsis Poiitidis 3a6fd00154 Merge pull request #785 from FEX-Emu/skmp/fix-sar8-16
OpDisp: Imm SAR OpSize < 32 needs sign extension
2021-02-18 22:33:46 +02:00
Stefanos Kornilios Mitsis Poiitidis 3bc12de017 OpDisp: Imm SAR OpSize < 32 needs sign extension 2021-02-17 08:17:20 +02:00
Stefanos Kornilios Mitsis Poiitidis 5ae6a64800 Merge pull request #778 from FEX-Emu/skmp/x87-round-trunc
x87: Round, Truncate & Precision control
2021-02-17 05:27:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 6615d5972f x87: Support for FCW, rounding, truncation, precision control 2021-02-17 05:16:53 +02:00
Stefanos Kornilios Mitsis Poiitidis d075b689a5 Merge pull request #784 from FEX-Emu/skmp/fix-cvt-cvtt
OpDisp: Fix cvt/cvtt mapping to dispatcher handlers
2021-02-17 05:14:44 +02:00
Stefanos Kornilios Mitsis Poiitidis eb60bda1d3 OpDisp: Fix cvt/cvtt mapping to dispatcher handlers 2021-02-16 11:19:28 +02:00
Stefanos Kornilios Mitsis Poiitidis 8c5d88816f Merge pull request #781 from FEX-Emu/skmp/hackfix-frem
X87: FPREM needs to set C2 to 0 to indicate finished iteration
2021-02-16 09:54:37 +02:00
Ryan Houdek 73e9a35e36 Merge pull request #783 from FEX-Emu/skmp/fix-selects
Syscalls/x86-32: Fix selects to remarshal the fd_sets
2021-02-15 23:26:33 -08:00
Stefanos Kornilios Mitsis Poiitidis 568fb3643d Syscalls/x86-32: Fix selects to actually remarshal the sets after running the host syscall 2021-02-16 02:35:38 +02:00
Stefanos Kornilios Mitsis Poiitidis 82367f49e2 X87: FPERM needs to set C2 to 0 to indicate finished iteration 2021-02-16 02:34:52 +02:00
Stefanos Kornilios Mitsis Poiitidis 9f2d49026e Merge pull request #768 from FEX-Emu/skmp/fix-csgo
Fixes for Counter Strike Global Offensive
2021-02-15 03:47:14 +02:00
Stefanos Kornilios Mitsis Poiitidis bed00d11f1 Tests: Disable pr57275.c because it uses VMOVAPS 2021-02-14 21:36:56 +02:00
Stefanos Kornilios Mitsis Poiitidis 9f0a3dab87 Merge pull request #769 from FEX-Emu/skmp/jit-fallbacks
x87: Call C++ handlers instead of forcing interpreter
2021-02-14 21:27:32 +02:00
Stefanos Kornilios Mitsis Poiitidis 07bd2244c7 IR: Remove ShouldInterpret & related logic 2021-02-12 14:21:47 +02:00
Stefanos Kornilios Mitsis Poiitidis 077093eacc JIT/x64: Use xmm0 as temp instead of xmm11, add xmm11 to RA list 2021-02-12 14:15:22 +02:00
Stefanos Kornilios Mitsis Poiitidis 428111b44b JIT/arm64: fallback support 2021-02-12 01:23:15 +02:00
Stefanos Kornilios Mitsis Poiitidis b564927825 IprOps: Add BCDLOAD, BCDSTORE fallbacks 2021-02-11 19:24:57 +02:00
Stefanos Kornilios Mitsis Poiitidis b0a046dfcb JIT/IPR: Add fallbacks for almost all x87 ops 2021-02-11 18:52:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 8f5bbde4dd JIT/x86: Implement VBSL 2021-02-11 18:51:05 +02:00
Stefanos Kornilios Mitsis Poiitidis 7dc0c275cf Add some basic infrastructure + a few fpr ops 2021-02-11 17:20:29 +02:00
Stefanos Kornilios Mitsis Poiitidis 4e31c1d1ce RA: Remove erratic assert 2021-02-11 14:31:53 +02:00
Stefanos Kornilios Mitsis Poiitidis df1c38908c JIT: Support edge case in CreateElementPair 2021-02-11 14:31:32 +02:00
Stefanos Kornilios Mitsis Poiitidis fc4b875ff1 Frontend: Only return error if invalid op in first block 2021-02-11 14:31:02 +02:00
Ryan Houdek 40b891ca69 Merge pull request #766 from phire/brk_fix
BRK: don't munmap if mmap failed
2021-02-08 21:32:02 -08:00
Scott Mansell 3d6461f0ca BRK: don't munmap if mmap failed 2021-02-09 17:29:44 +13:00
Ryan Houdek c11b629189 Merge pull request #764 from phire/fix_cpuinfo
Fix the number of physical cpus reported
2021-02-08 16:13:09 -08:00
Scott Mansell 46ec7c406f emulated-cpuid: add some comments 2021-02-09 13:03:49 +13:00
Ryan Houdek 2a0b9ae681 Merge pull request #763 from phire/bigger_spills
arm64: Fix blocks with over 256 spill slots.
2021-02-08 15:57:15 -08:00
Scott Mansell a28acf5505 Fix the number of physical cpus reported. 2021-02-09 12:52:27 +13:00
Scott Mansell c4dda4a3f3 arm64: Fix blocks with over 256 spill slots.
With multiblock and -n4000, at least one block in geekbench5's camera
benchmark was spilling 324 values. This meant the required stack
adjustment was 5184 bytes, which is larger than can fit into a 12bit
immediate.

Now, it's probally a bug that our RA is spilling that many values, and
we probally should be packing our spill slots closer together, but we
still shouldn't fail to run in this edge case.

So this commit adds a fallback path which uses a temp register to load
the stack adjustment.
2021-02-09 12:08:32 +13:00
Stefanos Kornilios Mitsis Poiitidis b40131ede3 Merge pull request #759 from phire/fix_mul_op
Fix power of 2 OP_MUL optimisation (Fixes Geekbench result upload)
2021-02-07 08:08:37 +02:00
Scott Mansell 5b0ade0bd8 Fix power of 2 OP_MUL optimisation
The new left-shift that replaced the multiply takes
it's size from arg[0]. If arg[0] was 64bit and the original
OP_MUL was 32bit, then the original code would truncate
upper bits, and the replacement code wouldn't.

This bug broke the SSL cert checking code in geekbench,
causing it to fail to upload results.
Fixes #647
2021-02-07 18:49:47 +13:00
Stefanos Kornilios Mitsis Poiitidis fcf1c2572b Merge pull request #756 from Sonicadvance1/python_check
Adds python version check
2021-02-06 14:20:49 +02:00
Ryan Houdek 69891867b1 Adds python version check 2021-02-05 10:36:52 -08:00
Ryan Houdek abefae2b58 Merge pull request #754 from Sonicadvance1/fix_artifacts_on_failure
Always upload artifact results even on failure
2021-02-03 16:08:51 -08:00
Ryan Houdek 143cb6d0a5 Always upload artifact results even on failure
This was an oversight of the original commit
2021-02-03 15:59:55 -08:00
Ryan Houdek e0042cc3fc Merge pull request #752 from Sonicadvance1/upload_artifacts
Uploads testing artifacts to github
2021-02-03 09:40:57 -08:00
Ryan Houdek 6ccdfd11a3 Uploads testing artifacts to github 2021-02-02 06:56:06 -08:00
Ryan Houdek 0aaed41f72 Merge pull request #751 from FEX-Emu/skmp/fix-fchmodat-faccessat
Syscalls: chmodat, faccessat glibc interface differs from kernel
2021-02-02 05:52:38 -08:00
Stefanos Kornilios Mitsis Poiitidis 3eed439728 gvisor: chmod passes now 2021-02-02 14:58:51 +02:00
Stefanos Kornilios Mitsis Poiitidis 4d05438796 Merge pull request #747 from phire/android_build_fixes
Android build fixes
2021-02-02 14:40:34 +02:00
Stefanos Kornilios Mitsis Poiitidis f9aa457e3f Syscalls: chmodat, faccessat glibc interface differs from kernel 2021-02-02 14:35:27 +02:00
Ryan Houdek eea14690c4 Merge pull request #750 from FEX-Emu/skmp/fix-empty-arg-forwarding
HarnessHelpers: Fix empty arg forwarding
2021-02-02 04:30:47 -08:00
Stefanos Kornilios Mitsis Poiitidis a3b5234d1c HarnessHelpers: Fix empty arg forwarding 2021-02-02 14:19:23 +02:00
Scott Mansell 4fad9a8c0b Assert on unreachable 2021-02-02 17:12:43 +13:00
Scott Mansell 9fa108ff33 Move -lnuma to correct place 2021-02-02 16:47:55 +13:00
Scott Mansell 45962098ef Refactor LinuxSyscalls to it's own static library 2021-02-02 16:47:55 +13:00
Scott Mansell 622c437829 undef PAGE_SIZE if it's defined
The android NDK defines this.
2021-02-02 16:47:55 +13:00
Scott Mansell 89bfacd845 Don't build thunks by default
Allow people to build fex without a cross-compiler installed
2021-02-02 16:47:55 +13:00
Scott Mansell 16e7ba6fa8 Move nasm check inside unit tests
Allow building of fex without nasm if unit tests are disabled
2021-02-02 16:47:55 +13:00
Scott Mansell c3ab5bca5e Fix <elf.h> incompatibilities
ELF32_ST_VISIBILITY is a pretty simple macro, just reimplement it.
2021-02-02 16:47:44 +13:00
Scott Mansell ef043f10ed Remove std::unexpected
It's deprecated in c++17, and android's bionic has removed it
2021-02-01 21:08:45 +13:00
Scott Mansell 0ff39ccad2 Remove dependance on system sigset_t
Android's sigset_t is the wrong size. Just define our own one.
2021-02-01 21:08:45 +13:00
Scott Mansell 93d5cdde1f Fix missing libc++ includes
gnu's libc++ is lenient and allows you to miss includes at times.

Android's binoic is more strict and needs these fixes to build.
2021-02-01 21:08:45 +13:00
Ryan Houdek 68beb28a56 Merge pull request #748 from phire/disable-more-tests
Disable more flaky gvisor tests
2021-01-31 21:51:07 -08:00
Scott Mansell ba3bacb65b Disable more flaky gvisor tests
They assume fair scheduling
2021-02-01 18:43:35 +13:00
Scott Mansell ebb05b56e4 Merge pull request #739 from Sonicadvance1/remove_warnings_10
Removes warning from Static Register Allocation Pass
2021-02-01 17:31:39 +13:00
Scott Mansell f5dd0fea5d Merge pull request #742 from Sonicadvance1/remove_warnings_8
Removes warning from IRCompaction
2021-02-01 17:31:28 +13:00
Ryan Houdek 4f4a8f7440 Merge pull request #746 from FEX-Emu/phire-disable-mcount-pic
Disable `mcount_pic.c.gcc-target-test-64` test
2021-01-31 19:50:45 -08:00
Scott Mansell 2e5f59cdc1 Merge pull request #734 from Sonicadvance1/remove_warnings_4
Removes warnings from OpcodeDispatcher
2021-02-01 16:45:37 +13:00
Scott Mansell 767813eb97 Merge pull request #735 from Sonicadvance1/remove_warnings_5
Removes warning from X86Tables
2021-02-01 16:44:24 +13:00
Scott Mansell 9c81fbb2fd Merge pull request #736 from Sonicadvance1/remove_warnings_6
Removes warning from IRDumper
2021-02-01 16:42:52 +13:00
Scott Mansell 5f624bec05 Merge pull request #733 from Sonicadvance1/remove_warnings_3
Removes warnings from x86-64 JIT.cpp
2021-02-01 16:42:10 +13:00
Scott Mansell a156210064 Merge pull request #738 from Sonicadvance1/remove_warnings_9
Removes warnings from Register Allocation Pass
2021-02-01 16:32:05 +13:00
Scott Mansell 03db1b2040 Disable mcount_pic.c.gcc-target-test-64 test
It's flaky on the ARMv8.0 runner
2021-02-01 16:26:39 +13:00
Scott Mansell d4fbca3cbf Merge pull request #732 from Sonicadvance1/remove_warnings_2
Removes warnings from Frontend
2021-02-01 16:23:29 +13:00
Scott Mansell 687176eca6 Merge pull request #741 from Sonicadvance1/remove_warnings_12
Removes warnings from HostFactory
2021-02-01 16:23:11 +13:00
Scott Mansell c14dfb9d52 Merge pull request #737 from Sonicadvance1/remove_warnings_7
Removes warning from DeadStoreElimination
2021-02-01 16:15:12 +13:00
Scott Mansell 27adabef83 Merge pull request #740 from Sonicadvance1/remove_warnings_11
Removes warnings from ELFLoader
2021-02-01 16:14:43 +13:00
Scott Mansell 1d102b9e59 Merge pull request #731 from Sonicadvance1/remove_warnings_1
Removes warnings from SoftFloat-3e
2021-02-01 16:08:42 +13:00
Ryan Houdek 1b1e922830 Removes warnings from HostFactory 2021-01-31 08:01:53 -08:00
Ryan Houdek 436e6e96ce Removes warnings from ELFLoader 2021-01-31 08:01:25 -08:00
Ryan Houdek d1161de4e4 Removes warning from Static Register Allocation Pass 2021-01-31 08:00:58 -08:00
Ryan Houdek 0c3ddd3016 Removes warnings from Register Allocation Pass 2021-01-31 08:00:28 -08:00
Ryan Houdek 5c9848650c Removes warning from IRCompaction 2021-01-31 07:59:59 -08:00
Ryan Houdek a90e2268f9 Removes warning from DeadStoreElimination 2021-01-31 07:59:26 -08:00
Ryan Houdek 175078b739 Removes warning from IRDumper 2021-01-31 07:58:58 -08:00
Ryan Houdek 9e55f1b177 Removes warning from X86Tables 2021-01-31 07:58:08 -08:00
Ryan Houdek 6012fcfee6 Removes warnings from OpcodeDispatcher 2021-01-31 07:57:28 -08:00
Ryan Houdek c2f4ea4646 Removes warnings from x86-64 JIT.cpp 2021-01-31 07:56:19 -08:00
Ryan Houdek 1f6017c37c Removes warnings from Frontend 2021-01-31 07:55:26 -08:00
Ryan Houdek 47a6b49f7f Removes warnings from SoftFloat-3e 2021-01-31 07:54:47 -08:00
Ryan Houdek fc7e0ede03 Merge pull request #728 from FEX-Emu/skmp/lookupcache-hotfix
LookupCache: Fix crash on L2 insertion
2021-01-29 17:14:29 -08:00
Stefanos Kornilios Mitsis Poiitidis 1f7d815f3a LookupCache: Fix crash on L2 insertion 2021-01-29 22:00:07 +02:00
Ryan Houdek 2edb14538f Merge pull request #725 from FEX-Emu/skmp/single-compiler-backend
Core: Single compiler backend
2021-01-29 05:30:47 -08:00
Stefanos Kornilios Mitsis Poiitidis 9ecde7a0b8 Interpreter: Move execution to IntepreterOps::ExecuteIR, single core refactor 2021-01-29 15:17:37 +02:00
Ryan Houdek f23fe030a1 Merge pull request #722 from FEX-Emu/skmp/ra-data
RA: Split RAData from RAPass, Optimize data packing, Cleanups
2021-01-29 05:16:18 -08:00
Stefanos Kornilios Mitsis Poiitidis 972219043d Core: Refactor compilation driver, lifetime of objects 2021-01-29 10:42:55 +02:00
Stefanos Kornilios Mitsis Poiitidis 0e00d3dc15 RA: Optimize RAData 2021-01-29 10:42:48 +02:00
Stefanos Kornilios Mitsis Poiitidis a45c2c34e9 RA: Use RAData directly 2021-01-29 10:42:40 +02:00
Stefanos Kornilios Mitsis Poiitidis 513112919b RA: Move allocation data to RAData 2021-01-29 10:42:40 +02:00
Ryan Houdek 810143d304 Merge pull request #723 from FEX-Emu/skmp/lookup-cache-refactor
BlockCache -> LookupCache, L3C
2021-01-29 00:22:27 -08:00
Ryan Houdek 3973449245 Merge pull request #721 from FEX-Emu/skmp/compile-fix
Syscalls: Define MREMAP_DONTUNMAP if missing
2021-01-29 00:18:15 -08:00
Stefanos Kornilios Mitsis Poiitidis 791aa4c32c LookupCache: Add a third level cache that will always resolve all known blocks 2021-01-29 01:53:12 +02:00
Stefanos Kornilios Mitsis Poiitidis 0c9e05ed8f Rename BlockCache to LookupCache 2021-01-29 01:31:48 +02:00
Stefanos Kornilios Mitsis Poiitidis d6ee3caad4 Syscalls: Define MREMAP_DONTUNMAP if missing 2021-01-28 17:09:08 +02:00
Ryan Houdek ded559139c Merge pull request #717 from Sonicadvance1/32bit_allocator
Finishes 32bit memory allocator
2021-01-27 11:00:46 -08:00
Ryan Houdek a98eaa6a27 Finishes 32bit memory allocator
This was partially implemented before but it had some bugs still.
This fixes those bugs and also implements mremap, shmat, and shmdt.

This now works well enough for 32bit graphical applications
2021-01-27 10:49:28 -08:00
Ryan Houdek 2d8b70bbd5 Merge pull request #710 from Sonicadvance1/brk_improvements
More improvements to BRK handling
2021-01-27 10:49:07 -08:00
Stefanos Kornilios Mitsis Poiitidis cc1000751a Merge pull request #719 from Sonicadvance1/ioctl32
Implements ioctl32 for x86-64
2021-01-27 20:18:11 +02:00
Stefanos Kornilios Mitsis Poiitidis f872e178ac Merge pull request #718 from Sonicadvance1/32bit_sigret_handler
Forces SIGRET handler to the lower 32bits VA on 32bit processes
2021-01-27 20:16:51 +02:00
Ryan Houdek a5d9e62cc6 More improvements to BRK handling
Base size is now only one page in size. We will then increment that BRK
size by 8MB alignments. 256MB for 32bit applications was causing some
applications on the edge to run out of virtual memory
I was hitting some 32bit applications that were being fairly mean with
BRK. They were allocating all of BRK space then running out of virtual
memory space with its mmap handler fallback after freeing BRK space.

This means we now munmap BRK pages on release for the guest, similar to
behaviour that Linux does.

Additionally I had an application that was getting very upset that BRK
wasn't actually at the end of program space. So allocate at the end of
program space like expected.

brk test now passes from gvisor
2021-01-27 10:12:25 -08:00
Stefanos Kornilios Mitsis Poiitidis ec31dc0872 Merge pull request #716 from Sonicadvance1/describe_tag
Puts the FEX describe tag in CPUID
2021-01-27 20:10:46 +02:00
Stefanos Kornilios Mitsis Poiitidis 4c4b797f51 Merge pull request #715 from Sonicadvance1/fix_vixl_asserts
Fixes some asserts from vixl and a bug in ExitSpillSRA
2021-01-27 20:09:33 +02:00
Ryan Houdek 15df787710 Implements ioctl32 for x86-64
If you have the COMPAT feature enabled then this just works there.
AArch64 not implemented currently since changes aren't upstream
2021-01-27 08:02:49 -08:00
Ryan Houdek 2437b89c7b Forces SIGRET handler to the lower 32bits VA on 32bit processes
This is visible to the guest process so for 32bit it needs to be in the
lower 32bits.
2021-01-27 07:54:30 -08:00
Ryan Houdek 5673a0e213 Puts the FEX describe tag in CPUID
This is what shows up in /proc/cpuinfo

example on tagged release:   'model name      : FEX-2101'
example on untagged release: 'model name      : FEX-2101-64-g536be23f'

Fixes #661
2021-01-27 07:20:14 -08:00
Ryan Houdek 3df344a78d Fixes some asserts from vixl and a bug in ExitSpillSRA
Assembler by default doesn't allow you to use the assembler. Needs a
CodeBufferCheckScope.

Labels with binds and not links will throw an assert.
ExitSpillSRA was using b instead of bind incorrectly.
Thus branching to a invalid location
2021-01-27 00:53:27 -08:00
Ryan Houdek f58d018934 Enables vixl asserts when in debug mode globally
Fixes an issue where some header asserts were missed
2021-01-27 00:52:37 -08:00
Ryan Houdek 536be23f68 Merge pull request #704 from FEX-Emu/skmp/add-gcc-target-tests
Add gcc target tests
2021-01-26 15:45:12 -08:00
Stefanos Kornilios Mitsis Poiitidis 8ff0b870b1 Merge pull request #712 from Sonicadvance1/fix_timers
Fixes timer syscall API
2021-01-26 15:33:00 +02:00
Ryan Houdek 3679156dd7 Fixes timer syscall API
The glibc and kernel timer interfaces are sufficiently different enough
that unwrapping this through glibc is impossible.
Pass it directly through to the kernel.
2021-01-26 05:02:44 -08:00
Stefanos Kornilios Mitsis Poiitidis 10e398a540 Merge pull request #711 from Sonicadvance1/implement_vbsl_aarch64
Implements VBSL on AArch64
2021-01-26 14:56:22 +02:00
Stefanos Kornilios Mitsis Poiitidis b453ede9ce Tests: Disable gcc target tests that fail only on arm 2021-01-26 14:55:35 +02:00
Ryan Houdek 6bca290611 Implements VBSL on AArch64
This isn't quite optimal without RA constraints
2021-01-26 03:01:48 -08:00
Stefanos Kornilios Mitsis Poiitidis dbfbfe5f0e Merge pull request #709 from Sonicadvance1/remove_boost
Removes all references to Boost
2021-01-26 12:40:43 +02:00
Stefanos Kornilios Mitsis Poiitidis 25bfab25da CI: Disable gcc target tests for x86-32 2021-01-26 12:37:07 +02:00
Ryan Houdek cab317e2f0 Removes all references to Boost
We don't require this
Fixes #699
2021-01-26 02:34:26 -08:00
Ryan Houdek 4d98c890e7 Merge pull request #708 from Sonicadvance1/decode_tsx
Add TSX names to decode tables so we know what they are
2021-01-26 02:24:47 -08:00
Ryan Houdek 513cb27825 Add TSX names to decode tables so we know what they are
Rather than just stating UND

We can't support this until we have an ARM CPU that supports TME and
more robust multiblock.
2021-01-26 02:03:42 -08:00
Stefanos Kornilios Mitsis Poiitidis 29a8fd0a4a Tests: Add gcc-target-tests for 32 and 64 bit 2021-01-26 12:00:45 +02:00
Stefanos Kornilios Mitsis Poiitidis e466dfb480 Merge pull request #707 from Sonicadvance1/update_gvisor_tests
Updates gvisor test results that have changed due to OS update
2021-01-26 11:59:38 +02:00
Ryan Houdek 8cbfef2e74 Updates gvisor test results that have changed due to OS update 2021-01-26 01:11:40 -08:00
Stefanos Kornilios Mitsis Poiitidis 53582fdf47 Merge pull request #701 from FEX-Emu/skmp/opdisp-32-bit-fixes
OpDisp: Address most 32-bit rip size issues
2021-01-25 15:31:55 +02:00
Stefanos Kornilios Mitsis Poiitidis b0365871b6 OpDisp: Address most 32-bit rip size issues 2021-01-25 15:04:12 +02:00
Ryan Houdek 2721418fb7 Merge pull request #698 from phire/better_defaults
Better Default Configuration
2021-01-22 16:56:29 -08:00
Scott Mansell 3cae0a0932 Better Default Configuration
Make FEX have good performance out of the box

resolves #685
2021-01-23 13:47:53 +13:00
Ryan Houdek d4ffea5c31 Merge pull request #670 from FEX-Emu/skmp/faster-ra
IR: Optimize runtime of optimization passes
2021-01-22 04:54:08 -08:00
Stefanos Kornilios Mitsis Poiitidis 0546da7a75 Address review feedback 2021-01-22 11:01:36 +02:00
Ryan Houdek ce1b9fcdc9 Merge pull request #674 from FEX-Emu/skmp/split-irparser
IR: Split parser from IRLoader, add IR::Parse(), Fixes, Tests
2021-01-21 23:09:53 -08:00
Stefanos Kornilios Mitsis Poiitidis 8b025d4c26 IRLoader: Actually set EntryRIP if the IR was parsed 2021-01-21 22:19:16 +02:00
Stefanos Kornilios Mitsis Poiitidis e7f2a650b5 Fix compilation 2021-01-21 22:18:48 +02:00
Stefanos Kornilios Mitsis Poiitidis 3bd8731b26 IR: Add CONFIG_VALIDATE_IR_PARSER, enable it on TestHarnessRunner 2021-01-21 21:38:41 +02:00
Stefanos Kornilios Mitsis Poiitidis 13613c6de1 IR: Change ValidateCode to use two uint64_t instead of __uint128_t 2021-01-21 21:37:58 +02:00
Stefanos Kornilios Mitsis Poiitidis 1177c375ff IR: Further parser cleanups, add fence types 2021-01-21 21:17:27 +02:00
Stefanos Kornilios Mitsis Poiitidis fb57e5183a IR: Fix TypeDefinition to actually fit used vector sizes 2021-01-21 21:10:11 +02:00
Stefanos Kornilios Mitsis Poiitidis 74160f38c4 IR: Fix IREmitter::Copy 2021-01-21 21:01:27 +02:00
Stefanos Kornilios Mitsis Poiitidis 072067b5fd IR: Split parser from IRLoader, add IR::Parse() 2021-01-21 19:17:51 +02:00
Stefanos Kornilios Mitsis Poiitidis 773e5866b4 Merge pull request #666 from Sonicadvance1/handle_unaligned_caspal32
Handle unaligned cmpxchg and cmpxchg8b on ARMv8.1+
2021-01-21 13:30:18 +02:00
Stefanos Kornilios Mitsis Poiitidis 00ca576644 IRCompaction: Only memset in debug 2021-01-21 13:06:43 +02:00
Ryan Houdek 5226f779c3 Merge pull request #673 from FEX-Emu/skmp/fix-ir-printing
IR: Fix Printing and Parsing for float conds
2021-01-21 00:28:50 -08:00
Stefanos Kornilios Mitsis Poiitidis 036771d825 IR: Fix Printing and Parsing for float conds 2021-01-21 10:16:08 +02:00
Ryan Houdek 633ea045e5 Handles unaligned cmpxchg/cmpxchg8b in SIGBUS handler
cmpxchg/cmpxchg8b doesn't have alignment requirements on x86, which means
applications rely on unaligned behaviour support with it.

Steam relies on this to work for a linked list array of jobs in some
internal job queueing system. It will end up always aligning by offsets
of 4 since it stores a couple of pointers.

Doesn't currently support the case of unaligned cmpxchg8b crossing a
cacheline, which ends up being semi-broken depending on which x86
behaviour the application is expected.
Intel CPUs do the "Big ring lock" or "split locks". Which means accesses
across cachelines are atomic.
AMD CPUs will tear the value across the cacheline, which is expected x86
behaviour by spec.

If they are expecting Intel behaviour, then that application is just
broken on non-Intel platforms unless they are fine with a tear.
2021-01-20 16:07:50 -08:00
Stefanos Kornilios Mitsis Poiitidis 0e7db64ac7 RA: Run compation after spilling, not before 2021-01-21 01:01:02 +02:00
Stefanos Kornilios Mitsis Poiitidis ac1036ac92 IR: Sync Invalid class with RA 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 9c462eb400 ConsProp: Revert to ordered set for identical codegen 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 8aef9cb6a7 RA: Exit after first spill per iteration is found 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis d3fc85e18d RA: Make spans at least 1 offset long 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 9f3ce47917 RA: Expire ending intervals before starting new ones 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis a9332e604a JIT/x64: Fix VAddV 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis e7e6b6654f RA: Cleanups 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis c3682f9a3b RA: Fix several bugs, get rid of virtual registers, remove unused complexity 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 47ee9af96c ConstProp: Keep pools in heap 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 10b6fb02c4 RA: Rename CalculatePrecessors to CalculatePredecessors 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 9c760ff83b RA: Get OP_JUMP/CONDJUMP without loop in CalculatePrecessors 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis f3e19ef015 DSE: No need to loop to find branching op 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 6bdf785223 DSE: Optimize map lookups 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 997650bafc DSE: Merge logic more for perf 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 56ca73191d DSE: Merge Flag/GPR/FPR passes for perf 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 900fc6d2a7 ConstProp: Switch maps to unordred_maps 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis af94e24ba3 RA: Simplify and optimize data structures 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 6380ab6bfd RA: Optimize register conflict handling 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 91ddc1dd38 RCLSE: Optimize FindMemberInfo to be O(1) 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis 77d072e35c RA: Switch active set to also use BucketList 2021-01-20 04:26:06 +02:00
Stefanos Kornilios Mitsis Poiitidis e847f94b69 RA: rework ConstrainedRAPass::CalculateNodeInterference 2021-01-20 04:26:06 +02:00
Ryan Houdek eb6386b921 Merge pull request #671 from FEX-Emu/skmp/ir-prettyprint
IR: Pretty print %Invalid, don't print invalid regs
2021-01-19 18:17:43 -08:00
Stefanos Kornilios Mitsis Poiitidis c42c57af35 IR: Pretty print %Invalid, don't print invalid regs 2021-01-20 04:06:37 +02:00
Stefanos Kornilios Mitsis Poiitidis 814ca4746b Merge pull request #669 from Sonicadvance1/workaround_llvm_native
Works around Clang failing to identify new Kryo CPUs
2021-01-20 03:48:55 +02:00
Ryan Houdek 921867de7e Works around Clang failing to identify new Kryo CPUs
Some of the newer CPU cores in LLVM's source claim to be a Cortex-A73,
which means they become limited to an ARMv8.0 feature set.

This is what you get if you compile FEX with -mcpu=native

To work around this issue, manually parse /proc/cpuinfo ourselves and
pull out the CPU type to pass to clang directly.
This also fixes the issue that we were using -march on AArch64, which no
longer works on newer clang versions. We instead need to use mcpu or
mtune.

Should improve all atomic op performance outside of the JITs, where they
were turning in to loadstore exclusive pairs.
2021-01-19 03:21:10 -08:00
Ryan Houdek fa542a5b9d Merge pull request #665 from FEX-Emu/skmp/continue-on-invalid-insn
Frontend, OpDisp: Don't assert on invalid instructions
2021-01-18 07:32:31 -08:00
Stefanos Kornilios Mitsis Poiitidis 567e315854 Frontend, OpDisp: Don't assert on invalid instructions 2021-01-18 17:19:11 +02:00
Stefanos Kornilios Mitsis Poiitidis 8ad32a135f Merge pull request #660 from FEX-Emu/skmp/remove-debug-printfs
OpDisp: Remove missed debug printf
2021-01-18 11:32:30 +02:00
Stefanos Kornilios Mitsis Poiitidis aa0cd76e06 OpDisp: Remove missed debug printf 2021-01-18 10:03:31 +02:00
198 changed files with 15107 additions and 7877 deletions

No files matched your search

+57
View File
@@ -59,20 +59,77 @@ jobs:
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Posix Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the posixtest
run: cmake --build . --config $BUILD_TYPE --target posix_tests
- name: Posix Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_Posix.log || true
- name: gvisor tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gvisor_tests
- name: GVisor Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GVisor.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the gvisor tests
run: cmake --build . --config $BUILD_TYPE --target gcc_target_tests_64
- name: GCC64 Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_GCC64.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+4
View File
@@ -26,3 +26,7 @@
shallow = true
path = External/fex-gvisor-tests-bins
url = https://github.com/FEX-Emu/fex-gvisor-tests-bins.git
[submodule "External/fex-gcc-target-tests-bins"]
shallow = true
path = External/fex-gcc-target-tests-bins
url = https://github.com/FEX-Emu/fex-gcc-target-tests-bins.git
+40 -16
View File
@@ -2,6 +2,7 @@ cmake_minimum_required(VERSION 3.14)
project(FEX)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
@@ -92,6 +93,9 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
add_definitions(-D_M_ARM_64=1)
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
if(CMAKE_BUILD_TYPE MATCHES DEBUG)
add_definitions(-DVIXL_DEBUG=1)
endif()
endif()
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
@@ -99,6 +103,8 @@ if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
message(FATAL_ERROR "FEX doesn't support getting compiled with GCC!")
endif()
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter Development)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
@@ -143,6 +149,22 @@ if(COMPILER_SUPPORTS_MARCH_NATIVE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -march=native")
endif()
if(_M_ARM_64)
# Due to an oversight in llvm, it declares any reasonably new Kryo CPU to only be ARMv8.0
# Manually detect newer CPU revisions until clang and llvm fixes their bug
# This script will either provide a supported CPU or 'native'
# Additionally -march doesn't work under AArch64+Clang, so you have to use -mcpu or -mtune
execute_process(COMMAND python3 "${PROJECT_SOURCE_DIR}/Scripts/aarch64_fit_native.py" "/proc/cpuinfo"
OUTPUT_VARIABLE AARCH64_CPU)
string(STRIP ${AARCH64_CPU} AARCH64_CPU)
check_cxx_compiler_flag("-mcpu=${AARCH64_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcpu=${AARCH64_CPU}")
endif()
endif()
add_compile_options(-Wall)
add_subdirectory(External/FEXCore)
@@ -156,21 +178,23 @@ if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
include(ExternalProject)
if (BUILD_THUNKS)
include(ExternalProject)
ExternalProject_Add(host-libs
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
ExternalProject_Add(host-libs
PREFIX host-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/HostLibs"
BINARY_DIR "Host"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
ExternalProject_Add(guest-libs
PREFIX guest-libs
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS "-DX86_C_COMPILER:STRING=${X86_C_COMPILER}" "-DX86_CXX_COMPILER:STRING=${X86_CXX_COMPILER}"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
)
endif()
+2 -2
View File
@@ -3,7 +3,7 @@ FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
libboost-dev clang-10 llvm-10 nasm ninja-build libnuma-dev \
clang-10 llvm-10 nasm ninja-build libnuma-dev \
libcap-dev libglfw3-dev libepoxy-dev
COPY . /opt/FEX
@@ -21,7 +21,7 @@ RUN ninja
FROM ubuntu:20.04
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y libboost-dev \
RUN DEBIAN_FRONTEND="noninteractive" apt install -y \
libnuma-dev libcap-dev libglfw3-dev libepoxy-dev
COPY --from=builder /opt/FEX/build/Bin/* /usr/bin/
+9 -2
View File
@@ -19,8 +19,7 @@ include(CheckIncludeFileCXX)
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -fno-operator-names -mcx16")
set(CMAKE_REQUIRED_DEFINITIONS "-fno-operator-names")
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
message(STATUS "Enabling x86-64 JIT")
set(ENABLE_JIT 1)
endif()
@@ -43,6 +42,7 @@ set(CMAKE_EXPORT_COMPILE_COMMANDS ON)
find_package(Git)
set(GIT_SHORT_HASH "Unknown")
set(GIT_DESCRIBE_STRING "FEX-Unknown")
if (GIT_FOUND)
execute_process(
@@ -52,6 +52,13 @@ if (GIT_FOUND)
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
endif()
configure_file(
+1 -1
View File
@@ -236,7 +236,7 @@ def print_ir_arg_printer(ops, defines):
if (SSAArgs != 0):
for i in range(0, SSAArgs):
LastArg = (SSAArgs - i - 1) == 0 and not HasArgs
output_file.write("\tPrintArg(out, IR, Op->Header.Args[%d], RAPass);\n" % i)
output_file.write("\tPrintArg(out, IR, Op->Header.Args[%d], RAData);\n" % i)
if not (LastArg):
output_file.write("\t*out << \", \";\n")
+8 -9
View File
@@ -118,7 +118,7 @@ set (SRCS
Common/SoftFloat-3e/s_f32UIToCommonNaN.c
Interface/Config/Config.cpp
Interface/Context/Context.cpp
Interface/Core/BlockCache.cpp
Interface/Core/LookupCache.cpp
Interface/Core/BlockSamplingData.cpp
Interface/Core/CompileService.cpp
Interface/Core/Core.cpp
@@ -131,6 +131,7 @@ set (SRCS
Interface/Core/X86DebugInfo.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/Interpreter/InterpreterCore.cpp
Interface/Core/Interpreter/InterpreterOps.cpp
Interface/Core/X86Tables/BaseTables.cpp
Interface/Core/X86Tables/DDDTables.cpp
Interface/Core/X86Tables/EVEXTables.cpp
@@ -144,7 +145,8 @@ set (SRCS
Interface/Core/X86Tables/X87Tables.cpp
Interface/Core/X86Tables/XOPTables.cpp
Interface/HLE/Thunks/Thunks.cpp
Interface/IR/IR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
@@ -155,10 +157,8 @@ set (SRCS
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadFlagStoreElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/StaticRegisterAllocationPass.cpp
Interface/IR/Passes/DeadGPRStoreElimination.cpp
Interface/IR/Passes/DeadFPRStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Utils/ELFLoader.cpp
@@ -170,7 +170,9 @@ if (_M_X86_64)
list(APPEND SRCS Interface/Core/Interpreter/x86_64Dispatcher.cpp)
endif()
if(_M_ARM_64)
list(APPEND SRCS Interface/Core/Interpreter/Arm64Dispatcher.cpp)
list(APPEND SRCS
Interface/Core/ArchHelpers/Arm64.cpp
Interface/Core/Interpreter/Arm64Dispatcher.cpp)
endif()
set (JIT_LIBS )
@@ -259,9 +261,6 @@ add_custom_target(IR_INC
check_cxx_compiler_flag(-fdiagnostics-color=always GCC_COLOR)
check_cxx_compiler_flag(-fcolor-diagnostics CLANG_COLOR)
set(LINUX_LIBS
numa)
function(AddLibrary Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
@@ -103,7 +103,9 @@ float32_t
if ( ! sig ) exp = 0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_ROUND_ODD
packReturn:
#endif
uiZ = packToF32UI( sign, exp, sig );
uiZ:
uZ.ui = uiZ;
@@ -107,7 +107,9 @@ float64_t
if ( ! sig ) exp = 0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_ROUND_ODD
packReturn:
#endif
uiZ = packToF64UI( sign, exp, sig );
uiZ:
uZ.ui = uiZ;
+12 -5
View File
@@ -77,7 +77,7 @@ struct X80SoftFloat {
}
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
return extF80_roundToInt(lhs, softfloat_round_near_even, false);
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
@@ -173,19 +173,26 @@ struct X80SoftFloat {
}
operator int16_t() const {
return extF80_to_i32(*this, softfloat_round_near_even, false);
auto rv = extF80_to_i32(*this, softfloat_roundingMode, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
return INT16_MIN;
} else {
return rv;
}
}
operator int32_t() const {
return extF80_to_i32(*this, softfloat_round_near_even, false);
return extF80_to_i32(*this, softfloat_roundingMode, false);
}
operator int64_t() const {
return extF80_to_i64(*this, softfloat_round_near_even, false);
return extF80_to_i64(*this, softfloat_roundingMode, false);
}
operator uint64_t() const {
return extF80_to_ui64(*this, softfloat_round_near_even, false);
return extF80_to_ui64(*this, softfloat_roundingMode, false);
}
void operator=(const float rhs) {
+20 -1
View File
@@ -4,6 +4,7 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Config/Config.h>
#include <map>
namespace FEXCore::Config {
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
@@ -33,7 +34,7 @@ namespace FEXCore::Config {
CTX->Config.TSOEnabled = Config != 0;
break;
case FEXCore::Config::CONFIG_SMC_CHECKS:
CTX->Config.SMCChecks = Config != 0;
CTX->Config.SMCChecks = static_cast<FEXCore::Config::ConfigSMCChecks>(Config);
break;
case FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS:
CTX->Config.ABILocalFlags = Config != 0;
@@ -41,6 +42,15 @@ namespace FEXCore::Config {
case FEXCore::Config::CONFIG_ABI_NO_PF:
CTX->Config.ABINoPF = Config != 0;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
CTX->Config.ValidateIRarser = Config != 0;
break;
case FEXCore::Config::CONFIG_AOTIR_GENERATE:
CTX->Config.AOTIRCapture = Config != 0;
break;
case FEXCore::Config::CONFIG_AOTIR_LOAD:
CTX->Config.AOTIRLoad = Config != 0;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
}
@@ -94,6 +104,15 @@ namespace FEXCore::Config {
case FEXCore::Config::CONFIG_ABI_NO_PF:
return CTX->Config.ABINoPF;
break;
case FEXCore::Config::CONFIG_VALIDATE_IR_PARSER:
return CTX->Config.ValidateIRarser;
break;
case FEXCore::Config::CONFIG_AOTIR_GENERATE:
return CTX->Config.AOTIRCapture;
break;
case FEXCore::Config::CONFIG_AOTIR_LOAD:
return CTX->Config.AOTIRLoad;
break;
default: LogMan::Msg::A("Unknown configuration option");
}
+15
View File
@@ -151,6 +151,21 @@ namespace FEXCore::Context {
return CTX->CPUID.RunFunction(Function);
}
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::istream>(const std::string&)> CacheReader) {
CTX->AOTIRLoader = CacheReader;
}
bool WriteAOTIR(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter) {
return CTX->WriteAOTIRCache(CacheWriter);
}
void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name) {
return CTX->AddNamedRegion(Base, Length, Offset, Name);
}
void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length) {
return CTX->RemoveNamedRegion(Base, Length);
}
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->CompileRIP(CTX->ParentThread, RIP);
+52 -6
View File
@@ -6,13 +6,20 @@
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/Utils/Event.h>
#include <stdint.h>
#include <memory>
#include <map>
#include <unordered_map>
#include <set>
#include <mutex>
#include <istream>
#include <ostream>
#include <functional>
namespace FEXCore {
class ThunkHandler;
@@ -30,6 +37,8 @@ class SyscallHandler;
namespace FEXCore::IR {
class RegisterAllocationPass;
class RegisterAllocationData;
class IRListView;
namespace Validation {
class IRValidation;
}
@@ -59,14 +68,23 @@ namespace FEXCore::Context {
bool Is64BitMode {true};
bool TSOEnabled {true};
bool SMCChecks {false};
FEXCore::Config::ConfigSMCChecks SMCChecks {FEXCore::Config::CONFIG_SMC_MMAN};
bool ABILocalFlags {false};
bool ABINoPF {false};
bool AOTIRCapture {false};
bool AOTIRLoad {false};
std::string DumpIR;
// this is for internal use
bool ValidateIRarser { false };
} Config;
using IntCallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
IntCallbackReturn InterpreterCallbackReturn;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
@@ -90,6 +108,28 @@ namespace FEXCore::Context {
CustomCPUFactoryType FallbackCPUFactory;
std::function<void(uint64_t ThreadId, FEXCore::Context::ExitReason)> CustomExitHandler;
struct AOTIRCacheEntry {
uint64_t start;
uint64_t len;
uint64_t crc;
IR::IRListView *IR;
IR::RegisterAllocationData *RAData;
};
std::function<std::unique_ptr<std::istream>(const std::string&)> AOTIRLoader;
std::unordered_map<std::string, std::map<uint64_t, AOTIRCacheEntry>> AOTIRCache;
struct AddrToFileEntry {
uint64_t Start;
uint64_t Len;
uint64_t Offset;
std::string fileid;
void *CachedFileEntry;
};
std::map<uint64_t, AddrToFileEntry> AddrToFile;
#ifdef BLOCKSTATS
std::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
@@ -136,10 +176,13 @@ namespace FEXCore::Context {
FEXCore::Core::ThreadState *GetThreadState();
void LoadEntryList();
std::tuple<void *, FEXCore::Core::DebugData *> CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileFallbackBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<FEXCore::IR::IRListView *, FEXCore::IR::RegisterAllocationData *, uint64_t, uint64_t, uint64_t, uint64_t> GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
std::tuple<void *, FEXCore::IR::IRListView *, FEXCore::Core::DebugData *, FEXCore::IR::RegisterAllocationData *, bool, uint64_t, uint64_t> CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
uintptr_t CompileBlock(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
bool LoadAOTIRCache(std::istream &stream);
bool WriteAOTIRCache(std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter);
// Used for thread creation from syscalls
void InitializeCompiler(FEXCore::Core::InternalThreadState* State, bool CompileThread);
FEXCore::Core::InternalThreadState* CreateThread(FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID);
@@ -150,12 +193,15 @@ namespace FEXCore::Context {
std::vector<FEXCore::Core::InternalThreadState*> *const GetThreads() { return &Threads; }
void AddNamedRegion(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const std::string &filename);
void RemoveNamedRegion(uintptr_t Base, uintptr_t Size);
#if ENABLE_JITSYMBOLS
FEXCore::JITSymbols Symbols;
#endif
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread, bool AlsoClearIRCache);
private:
void WaitForIdleWithTimeout();
@@ -163,7 +209,7 @@ namespace FEXCore::Context {
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread);
void NotifyPause();
uintptr_t AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr, uint64_t Start, uint64_t Length);
FEXCore::CodeLoader *LocalLoader{};
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,30 @@
#pragma once
#include <stdint.h>
namespace FEXCore::ArchHelpers::Arm64 {
constexpr uint32_t CASPAL_MASK = 0xBF'E0'FC'00;
constexpr uint32_t CASPAL_INST = 0x08'60'FC'00;
constexpr uint32_t CASAL_MASK = 0x3F'E0'FC'00;
constexpr uint32_t CASAL_INST = 0x08'E0'FC'00;
constexpr uint32_t ATOMIC_MEM_MASK = 0x3B200C00;
constexpr uint32_t ATOMIC_MEM_INST = 0x38200000;
// Load ops are 4 bits
// Acquire and release bits are independent on the instruction
constexpr uint32_t ATOMIC_ADD_OP = 0b0000;
constexpr uint32_t ATOMIC_CLR_OP = 0b0001;
constexpr uint32_t ATOMIC_EOR_OP = 0b0010;
constexpr uint32_t ATOMIC_SET_OP = 0b0011;
constexpr uint32_t ATOMIC_SMAX_OP = 0b0100;
constexpr uint32_t ATOMIC_SMIN_OP = 0b0101;
constexpr uint32_t ATOMIC_UMAX_OP = 0b0110;
constexpr uint32_t ATOMIC_UMIN_OP = 0b0111;
constexpr uint32_t ATOMIC_SWAP_OP = 0b1000;
bool HandleCASPAL(void *_mcontext, void *_info, uint32_t Instr);
bool HandleCASAL(void *_mcontext, void *_info, uint32_t Instr);
bool HandleAtomicMemOp(void *_mcontext, void *_info, uint32_t Instr);
}
+27 -3
View File
@@ -109,6 +109,31 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h() {
return Res;
}
// 2: Cache and TLB information
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_02h() {
FEXCore::CPUID::FunctionResults Res{};
// returns default values from i7 model 1Ah
Res.eax = 0x1 | // Number of iterations needed for all descriptors
(0x5A << 8) |
(0x03 << 16) |
(0x55 << 24);
Res.ebx = 0xE4 |
(0xB2 << 8) |
(0xF0 << 16) |
(0 << 24);
Res.ecx = 0; // null descriptors
Res.edx = 0x2C |
(0x21 << 8) |
(0xCA << 16) |
(0x09 << 24);
return Res;
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_06h() {
FEXCore::CPUID::FunctionResults Res{};
Res.eax = (1 << 2); // Always running APIC
@@ -327,8 +352,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h() {
}
constexpr char ProcessorBrand[48] = {
"FEX-"
GIT_SHORT_HASH
GIT_DESCRIBE_STRING
"\0"
};
@@ -367,7 +391,7 @@ void CPUIDEmu::Init(FEXCore::Context::Context *ctx) {
CTX = ctx;
RegisterFunction(0, std::bind(&CPUIDEmu::Function_0h, this));
RegisterFunction(1, std::bind(&CPUIDEmu::Function_01h, this));
// 2: Cache and TLB information
RegisterFunction(2, std::bind(&CPUIDEmu::Function_02h, this));
// 3: Serial Number(previously), now reserved
// 4: Deterministic cache parameters for each level
// 5: Monitor/mwait
+7 -1
View File
@@ -3,6 +3,7 @@
#include <unordered_map>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Utils/LogManager.h>
namespace FEXCore {
namespace Context {
@@ -25,8 +26,12 @@ public:
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function) {
auto Handler = FunctionHandlers.find(Function);
if (Handler == FunctionHandlers.end())
if (Handler == FunctionHandlers.end()) {
#ifndef NDEBUG
LogMan::Msg::E("Unhandled CPU ID function, 0x%x", Function);
#endif
return Function_Reserved();
}
return Handler->second();
}
@@ -43,6 +48,7 @@ private:
// Functions
FEXCore::CPUID::FunctionResults Function_0h();
FEXCore::CPUID::FunctionResults Function_01h();
FEXCore::CPUID::FunctionResults Function_02h();
FEXCore::CPUID::FunctionResults Function_06h();
FEXCore::CPUID::FunctionResults Function_07h();
FEXCore::CPUID::FunctionResults Function_8000_0000h();
+19 -37
View File
@@ -1,5 +1,5 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/BlockCache.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/CompileService.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -33,7 +33,7 @@ namespace FEXCore {
WorkerThread.join();
}
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
void CompileService::ClearCache(FEXCore::Core::InternalThreadState *Thread) {
// On cache clear we need to spin down the execution thread to ensure it isn't trying to give us more work items
if (CompileMutex.try_lock()) {
// We can only clear these things if we pulled the compile mutex
@@ -57,30 +57,17 @@ namespace FEXCore {
}
}
if (GuestRIP == 0) {
CompileThreadData->IRLists.clear();
}
else {
auto IR = CompileThreadData->IRLists.find(GuestRIP)->second.release();
CompileThreadData->IRLists.clear();
CompileThreadData->IRLists.try_emplace(GuestRIP, IR);
}
LogMan::Throw::A(CompileThreadData->LocalIRCache.size() == 0, "Compile service must never have LocalIRCache");
CompileMutex.unlock();
}
// Clear the inverse cache of what is calling us from the Context ClearCache routine
auto SelectedThread = Thread->IsCompileService ? ParentThread : Thread;
SelectedThread->BlockCache->ClearCache();
SelectedThread->LookupCache->ClearCache();
SelectedThread->CPUBackend->ClearCache();
SelectedThread->IntBackend->ClearCache();
}
void CompileService::RemoveCodeEntry(uint64_t GuestRIP) {
CompileThreadData->IRLists.erase(GuestRIP);
CompileThreadData->DebugData.erase(GuestRIP);
CompileThreadData->BlockCache->Erase(GuestRIP);
}
CompileService::WorkItem *CompileService::CompileCode(uint64_t RIP) {
// Tell the worker thread to compile code for us
@@ -132,33 +119,28 @@ namespace FEXCore {
// If we had a work item then work on it
if (Item) {
// Does the block cache already contain this RIP?
void *CompiledCode = reinterpret_cast<void*>(CompileThreadData->BlockCache->FindBlock(Item->RIP));
FEXCore::Core::DebugData *DebugData = nullptr;
// Make sure it's not in lookup cache by accident
LogMan::Throw::A(CompileThreadData->LookupCache->FindBlock(Item->RIP) == 0, "Compile Service must never have entries in the LookupCache");
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->State.State.rip = Item->RIP;
if (!CompiledCode) {
// Code isn't in cache, compile now
// Set our thread state's RIP
CompileThreadData->State.State.rip = Item->RIP;
auto [Code, Data] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
CompiledCode = Code;
DebugData = Data;
}
auto [CodePtr, IRList, DebugData, RAData, Generated, StartAddr, Length] = CTX->CompileCode(CompileThreadData.get(), Item->RIP);
if (!CompiledCode) {
LogMan::Throw::A(Generated == true, "Compile Service doesn't have IR Cache");
if (!CodePtr) {
// XXX: We currently have the expectation that compile service code will be significantly smaller than regular thread's code
ERROR_AND_DIE("Couldn't compile code for thread at RIP: 0x%lx", Item->RIP);
}
auto BlockMapPtr = CompileThreadData->BlockCache->AddBlockMapping(Item->RIP, CompiledCode);
if (BlockMapPtr == 0) {
// XXX: We currently have the expectation that compiler service block cache will be significantly underutilized compared to regular thread
ERROR_AND_DIE("Couldn't add code to block cache for thread at RIP: 0x%lx", Item->RIP);
}
Item->CodePtr = CompiledCode;
Item->IRList = CompileThreadData->IRLists.find(Item->RIP)->second.get();
Item->CodePtr = CodePtr;
Item->IRList = IRList;
Item->DebugData = DebugData;
Item->RAData = RAData;
Item->StartAddr = StartAddr;
Item->Length = Length;
GCArray.emplace_back(Item);
Item->ServiceWorkDone.NotifyAll();
+8 -3
View File
@@ -16,6 +16,9 @@ namespace Core {
struct InternalThreadState;
}
namespace IR {
class RegisterAllocationData;
};
class CompileService final {
public:
CompileService(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
@@ -28,8 +31,11 @@ class CompileService final {
// Outgoing
void *CodePtr{};
FEXCore::IR::IRListView<true> *IRList{};
FEXCore::IR::IRListView *IRList{};
FEXCore::IR::RegisterAllocationData *RAData{};
FEXCore::Core::DebugData *DebugData{};
uint64_t StartAddr;
uint64_t Length;
// Communication
Event ServiceWorkDone{};
@@ -37,8 +43,7 @@ class CompileService final {
};
WorkItem *CompileCode(uint64_t RIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
void RemoveCodeEntry(uint64_t GuestRIP);
void ClearCache(FEXCore::Core::InternalThreadState *Thread);
private:
FEXCore::Context::Context *CTX;
File diff suppressed because it is too large. Load diff
+82 -60
View File
@@ -329,9 +329,21 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op)
DecodeInst->TableInfo = Info;
// XXX: Once we support 32bit x86 then this will be necessary to support
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_LEGACY_PREFIX, "Legacy Prefix");
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_UNKNOWN, "Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_INVALID, "Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
return false;
}
LogMan::Throw::A(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
@@ -565,9 +577,22 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
DecodeInst->TableInfo = Info;
// XXX: Once we support 32bit x86 then this will be necessary to support
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_LEGACY_PREFIX, "Legacy Prefix");
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_UNKNOWN, "Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_INVALID, "Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
if (Info->Type == FEXCore::X86Tables::TYPE_LEGACY_PREFIX) {
LogMan::Msg::D("Legacy Prefix");
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
LogMan::Msg::D("Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
return false;
}
if (Info->Type == FEXCore::X86Tables::TYPE_INVALID) {
LogMan::Msg::D("Invalid or Unknown instruction: %s 0x%04x 0x%lx", Info->Name, Op, DecodeInst->PC);
return false;
}
LogMan::Throw::A(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX, "REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
@@ -686,36 +711,12 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return NormalOp(LocalInfo, Op);
}
else if (Info->Type == FEXCore::X86Tables::TYPE_GROUP_EVEX) {
uint8_t P1 = ReadByte();
uint8_t P2 = ReadByte();
uint8_t P3 = ReadByte();
/* uint8_t P1 = */ ReadByte();
/* uint8_t P2 = */ ReadByte();
/* uint8_t P3 = */ ReadByte();
uint8_t EVEXOp = ReadByte();
return NormalOp(&EVEXTableOps[EVEXOp], EVEXOp);
}
else if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
LogMan::Throw::A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
if (Op & 0b1000) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_WIDENING;
DecodeFlags::PushOpAddr(&DecodeInst->Flags, DecodeFlags::FLAG_WIDENING_SIZE_LAST);
}
// XGPR_B bit set
if (Op & 0b0001)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
// XGPR_X bit set
if (Op & 0b0010)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
// XGPR_R bit set
if (Op & 0b0100)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
return false;
}
return NormalOp(Info, Op);
}
@@ -723,13 +724,12 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
bool Decoder::DecodeInstruction(uint64_t PC) {
InstructionSize = 0;
Instruction.fill(0);
bool InstructionDecoded = false;
DecodeInst = &DecodedBuffer[DecodedSize];
memset(DecodeInst, 0, sizeof(DecodedInst));
DecodeInst->PC = PC;
while (!InstructionDecoded) {
for(;;) {
uint8_t Op = ReadByte();
switch (Op) {
case 0x0F: {// Escape Op
@@ -753,9 +753,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
// Take a peek at the op just past the displacement
uint8_t LocalOp = ReadByte();
if (NormalOpHeader(&FEXCore::X86Tables::DDDNowOps[LocalOp], LocalOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::DDDNowOps[LocalOp], LocalOp);
break;
}
case 0x38: { // F38 Table!
@@ -770,9 +768,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
Prefix = PF_38_66;
uint16_t LocalOp = (Prefix << 8) | ReadByte();
if (NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::H0F38TableOps[LocalOp], LocalOp);
break;
}
case 0x3A: { // F3A Table!
@@ -788,9 +784,7 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
Prefix |= PF_3A_REX;
uint16_t LocalOp = (Prefix << 8) | ReadByte();
if (NormalOpHeader(&FEXCore::X86Tables::H0F3ATableOps[LocalOp], LocalOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::H0F3ATableOps[LocalOp], LocalOp);
break;
}
default: // Two byte table!
@@ -805,37 +799,29 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
bool NoOverlay66 = (FEXCore::X86Tables::SecondBaseOps[EscapeOp].Flags & InstFlags::FLAGS_NO_OVERLAY66) != 0;
if (NoOverlay) { // This section of the table ignores prefix extention
if (NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0xF3) { // REP
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REP_PREFIX;
if (NormalOpHeader(&FEXCore::X86Tables::RepModOps[EscapeOp], EscapeOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::RepModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0xF2) { // REPNE
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REPNE_PREFIX;
if (NormalOpHeader(&FEXCore::X86Tables::RepNEModOps[EscapeOp], EscapeOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::RepNEModOps[EscapeOp], EscapeOp);
}
else if (DecodeInst->LastEscapePrefix == 0x66 && !NoOverlay66) { // Operand Size
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
DecodeFlags::PopOpAddrIf(&DecodeInst->Flags, DecodeFlags::FLAG_OPERAND_SIZE_LAST);
if (NormalOpHeader(&FEXCore::X86Tables::OpSizeModOps[EscapeOp], EscapeOp)) {
InstructionDecoded = true;
}
return NormalOpHeader(&FEXCore::X86Tables::OpSizeModOps[EscapeOp], EscapeOp);
}
else if (NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp)) {
InstructionDecoded = true;
else {
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
}
break;
}
@@ -891,9 +877,33 @@ bool Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->Flags |= DecodeFlags::FLAG_GS_PREFIX;
break;
default: { // Default base table
if (NormalOpHeader(&FEXCore::X86Tables::BaseOps[Op], Op)) {
InstructionDecoded = true;
auto Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
LogMan::Throw::A(CTX->Config.Is64BitMode, "Got REX prefix in 32bit mode");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
// Widening displacement
if (Op & 0b1000) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_WIDENING;
DecodeFlags::PushOpAddr(&DecodeInst->Flags, DecodeFlags::FLAG_WIDENING_SIZE_LAST);
}
// XGPR_B bit set
if (Op & 0b0001)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
// XGPR_X bit set
if (Op & 0b0010)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
// XGPR_R bit set
if (Op & 0b0100)
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_R;
} else {
return NormalOpHeader(Info, Op);
}
break;
}
}
@@ -982,6 +992,9 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
SymbolMinAddress = EntryPoint;
}
DecodedMinAddress = EntryPoint;
DecodedMaxAddress = EntryPoint;
// Entry is a jump target
BlocksToDecode.emplace(PC);
@@ -1002,11 +1015,20 @@ bool Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC)
while (1) {
ErrorDuringDecoding = !DecodeInstruction(RIPToDecode + PCOffset);
if (ErrorDuringDecoding) {
LogMan::Msg::D("Couldn't Decode something at 0x%lx, Started at 0x%lx", PC + PCOffset, PC);
LogMan::Throw::A(Blocks.size() != 1, "Decode Error in entry block");
CurrentBlockDecoding.HasInvalidInstruction = true;
if (ErrorDuringDecoding && Blocks.size() != 1) {
ErrorDuringDecoding = false;
}
break;
}
DecodedMinAddress = std::min(DecodedMinAddress, RIPToDecode + PCOffset);
DecodedMaxAddress = std::max(DecodedMaxAddress, RIPToDecode + PCOffset + DecodeInst->InstSize);
++TotalInstructions;
++BlockNumberOfInstructions;
++DecodedSize;
+4
View File
@@ -20,6 +20,7 @@ public:
uint64_t Entry{};
uint64_t NumInstructions{};
FEXCore::X86Tables::DecodedInst *DecodedInstructions;
bool HasInvalidInstruction{};
};
Decoder(FEXCore::Context::Context *ctx);
@@ -29,6 +30,9 @@ public:
return &Blocks;
}
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
private:
FEXCore::Context::Context *CTX;
@@ -199,13 +199,13 @@ DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Co
auto RipReg = x2;
// Mask the address by the virtual address size so we can check for aliases
LoadConstant(x3, Thread->BlockCache->GetVirtualMemorySize() - 1);
LoadConstant(x3, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, x3);
{
// This is the block cache lookup routine
// It matches what is going on it BlockCache.h::FindBlock
LoadConstant(x0, Thread->BlockCache->GetPagePointer());
// It matches what is going on it LookupCache.h::FindBlock
LoadConstant(x0, Thread->LookupCache->GetPagePointer());
// Offset the address and add to our page pointer
lsr(x1, x3, 12);
@@ -220,16 +220,16 @@ DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Co
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::BlockCache::BlockCacheEntry))));
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::BlockCache::BlockCacheEntry, GuestCode)));
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x1, MemOperand(x0, offsetof(FEXCore::BlockCache::BlockCacheEntry, HostCode)));
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x1, &NoBlock);
// If we've made it here then we have a real compiled block
@@ -540,6 +540,10 @@ void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore:
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
// TODO: Implement this. It is missing from the dispatcher
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = nullptr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
@@ -1,6 +1,6 @@
#pragma once
#include "Interface/Core/BlockCache.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/InternalThreadState.h"
#include <FEXCore/Core/CPUBackend.h>
@@ -22,19 +22,16 @@ public:
explicit InterpreterCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread);
~InterpreterCore() override;
std::string GetName() override { return "Interpreter"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) override;
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
bool NeedsOpDispatch() override { return true; }
void ExecuteCode(FEXCore::Core::InternalThreadState *Thread);
void CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread);
void DeleteAsmDispatch();
using CallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
CallbackReturn ReturnPtr;
bool HandleSIGBUS(int Signal, void *info, void *ucontext);
private:
FEXCore::Context::Context *CTX;
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,42 @@
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::IR {
class IRListView;
}
namespace FEXCore::Core{
struct DebugData;
}
namespace FEXCore::CPU {
enum FallbackABI {
FABI_UNKNOWN,
FABI_VOID_U16,
FABI_F80_F32,
FABI_F80_F64,
FABI_F80_I16,
FABI_F80_I32,
FABI_F32_F80,
FABI_F64_F80,
FABI_I16_F80,
FABI_I32_F80,
FABI_I64_F80,
FABI_I64_F80_F80,
FABI_F80_F80,
FABI_F80_F80_F80,
};
struct FallbackInfo {
FallbackABI ABI;
void *fn;
};
class InterpreterOps {
public:
static void InterpretIR(FEXCore::Core::InternalThreadState *Thread, FEXCore::IR::IRListView *CurrentIR, FEXCore::Core::DebugData *DebugData);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
};
};
@@ -1,5 +1,6 @@
#include "Common/MathUtils.h"
#include "Interface/Core/Interpreter/InterpreterClass.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/X86Enums.h>
#include <cmath>
@@ -17,7 +18,7 @@ class DispatchGenerator : public Xbyak::CodeGenerator {
CPUBackend::AsmDispatch DispatchPtr;
CPUBackend::JITCallback CallbackPtr;
InterpreterCore::CallbackReturn ReturnPtr;
FEXCore::Context::Context::IntCallbackReturn ReturnPtr;
uint64_t ThreadStopHandlerAddress;
uint64_t AbsoluteLoopTopAddress;
@@ -94,13 +95,13 @@ DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Co
AbsoluteLoopTopAddress = getCurr<uint64_t>();
{
mov(r13, Thread->BlockCache->GetPagePointer());
mov(r13, Thread->LookupCache->GetPagePointer());
// Load our RIP
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
mov(rax, rdx);
mov(rbx, Thread->BlockCache->GetVirtualMemorySize() - 1);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
@@ -113,7 +114,7 @@ DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Co
mov (rax, rdx);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::BlockCache::BlockCacheEntry)));
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
@@ -238,7 +239,7 @@ DispatchGenerator::DispatchGenerator(FEXCore::Context::Context *ctx, FEXCore::Co
{
ReturnPtr = getCurr<InterpreterCore::CallbackReturn>();
ReturnPtr = getCurr<FEXCore::Context::Context::IntCallbackReturn>();
// using CallbackReturn = __attribute__((naked)) void(*)(FEXCore::Core::InternalThreadState *Thread, volatile void *Host_RSP);
// rdi = thread
@@ -447,7 +448,9 @@ void InterpreterCore::CreateAsmDispatch(FEXCore::Context::Context *ctx, FEXCore:
Generator = new DispatchGenerator(ctx, Thread);
DispatchPtr = Generator->DispatchPtr;
CallbackPtr = Generator->CallbackPtr;
ReturnPtr = Generator->ReturnPtr;
// TODO: It feels wrong to initialize this way
ctx->InterpreterCallbackReturn = Generator->ReturnPtr;
}
bool InterpreterCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
+14 -5
View File
@@ -51,10 +51,22 @@ DEF_OP(Constant) {
LoadConstant(Dst, Op->Constant);
}
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = IR->GetHeader()->Entry + Op->Offset;
auto Dst = GetReg<RA_64>(Node);
LoadConstant(Dst, Constant);
}
DEF_OP(InlineConstant) {
//nop
}
DEF_OP(InlineEntrypointOffset) {
//nop
}
DEF_OP(CycleCounter) {
#ifdef DEBUG_CYCLES
movz(GetReg<RA_64>(Node), 0);
@@ -1043,17 +1055,15 @@ DEF_OP(FCmp) {
}
}
DEF_OP(F80Cmp) {
LogMan::Msg::D("Unimplemented");
}
#undef DEF_OP
void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
@@ -1095,7 +1105,6 @@ void JITCore::RegisterALUHandlers() {
REGISTER_OP(FLOAT_TOGPR_U, Float_ToGPR_U);
REGISTER_OP(FLOAT_TOGPR_S, Float_ToGPR_S);
REGISTER_OP(FCMP, FCmp);
REGISTER_OP(F80CMP, F80Cmp);
#undef REGISTER_OP
}
@@ -3,6 +3,7 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
using namespace vixl;
@@ -22,9 +23,7 @@ DEF_OP(GuestReturn) {
DEF_OP(SignalReturn) {
// First we must reset the stack
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
// Now branch to our signal return helper
// This can't be a direct branch since the code needs to live at a constant location
@@ -38,9 +37,7 @@ DEF_OP(CallbackReturn) {
SpillStaticRegs();
// First we must reset the stack
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
// We can now lower the ref counter again
LoadConstant(x0, reinterpret_cast<uint64_t>(ThreadSharedData.SignalHandlerRefCounterPtr));
@@ -64,14 +61,12 @@ DEF_OP(ExitFunction) {
Label FullLookup;
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
aarch64::Register RipReg;
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP)) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Literal l_BranchHost{ExitFunctionLinkerAddress};
Literal l_BranchGuest{NewRIP};
@@ -84,9 +79,9 @@ DEF_OP(ExitFunction) {
RipReg = GetReg<RA_64>(Op->Header.Args[0].ID());
// L1 Cache
LoadConstant(x0, State->BlockCache->GetL1Pointer());
LoadConstant(x0, State->LookupCache->GetL1Pointer());
and_(x3, RipReg, BlockCache::L1_ENTRIES_MASK);
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
ldp(x1, x0, MemOperand(x0));
@@ -247,7 +242,8 @@ DEF_OP(Thunk) {
#if _M_X86_64
ERROR_AND_DIE("JIT: OP_THUNK not supported with arm simulator")
#else
LoadConstant(x2, Op->ThunkFnPtr);
auto thunkFn = State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
blr(x2);
#endif
@@ -259,15 +255,23 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
uint8_t *NewCode = (uint8_t *)Op->CodePtr;
uint8_t *OldCode = (uint8_t *)&Op->CodeOriginal;
uint8_t *OldCode = (uint8_t *)&Op->CodeOriginalLow;
int len = Op->CodeLength;
int idx = 0;
LoadConstant(GetReg<RA_64>(Node), 0);
LoadConstant(x0, Op->CodePtr);
LoadConstant(x0, IR->GetHeader()->Entry + Op->Offset);
LoadConstant(x1, 1);
while (len >= 8)
{
ldr(x2, MemOperand(x0, idx));
LoadConstant(x3, *(uint32_t *)(OldCode + idx));
cmp(x2, x3);
csel(GetReg<RA_64>(Node), GetReg<RA_64>(Node), x1, Condition::eq);
len -= 8;
idx += 8;
}
while (len >= 4)
{
ldr(w2, MemOperand(x0, idx));
@@ -306,7 +310,7 @@ DEF_OP(RemoveCodeEntry) {
PushDynamicRegsAndLR();
mov(x0, STATE);
LoadConstant(x1, Op->RIP);
LoadConstant(x1, IR->GetHeader()->Entry);
LoadConstant(x2, reinterpret_cast<uintptr_t>(&Context::Context::RemoveCodeEntry));
SpillStaticRegs();
+420 -166
View File
@@ -1,5 +1,6 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/Arm64.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
@@ -7,6 +8,7 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Core/UContext.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <sys/mman.h>
#include <stdio.h>
@@ -38,8 +40,260 @@ using namespace vixl;
using namespace vixl::aarch64;
void JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Op: %s", std::string(Name).c_str());
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Op: %s", std::string(Name).c_str());
} else {
switch(Info.ABI) {
case FABI_VOID_U16:{
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
LoadConstant(x1, (uintptr_t)Info.fn);
blr(x1);
PopDynamicRegsAndLR();
FillStaticRegs();
}
break;
case FABI_F80_F32:{
SpillStaticRegs();
PushDynamicRegsAndLR();
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
LoadConstant(x0, (uintptr_t)Info.fn);
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
ins(GetDst(Node).V2D(), 0, x0);
ins(GetDst(Node).V8H(), 4, w1);
}
break;
case FABI_F80_F64:{
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
LoadConstant(x0, (uintptr_t)Info.fn);
blr(x0);
PopDynamicRegsAndLR();
FillStaticRegs();
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
ins(GetDst(Node).V2D(), 0, x0);
ins(GetDst(Node).V8H(), 4, w1);
}
break;
case FABI_F80_I16:
case FABI_F80_I32: {
SpillStaticRegs();
PushDynamicRegsAndLR();
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
LoadConstant(x1, (uintptr_t)Info.fn);
blr(x1);
PopDynamicRegsAndLR();
FillStaticRegs();
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
ins(GetDst(Node).V2D(), 0, x0);
ins(GetDst(Node).V8H(), 4, w1);
}
break;
case FABI_F32_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
fmov(GetDst(Node).S(), v0.S());
}
break;
case FABI_F64_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetDst(Node).D(), v0.D());
}
break;
case FABI_I16_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
uxth(GetReg<RA_64>(Node), x0);
}
break;
case FABI_I32_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetReg<RA_32>(Node), w0);
}
break;
case FABI_I64_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetReg<RA_64>(Node), x0);
}
break;
case FABI_I64_F80_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(x3, GetSrc(IROp->Args[1].ID()).V2D(), 1);
LoadConstant(x4, (uintptr_t)Info.fn);
blr(x4);
PopDynamicRegsAndLR();
FillStaticRegs();
mov(GetReg<RA_64>(Node), x0);
}
break;
case FABI_F80_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
LoadConstant(x2, (uintptr_t)Info.fn);
blr(x2);
PopDynamicRegsAndLR();
FillStaticRegs();
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
ins(GetDst(Node).V2D(), 0, x0);
ins(GetDst(Node).V8H(), 4, w1);
}
break;
case FABI_F80_F80_F80:{
SpillStaticRegs();
PushDynamicRegsAndLR();
umov(x0, GetSrc(IROp->Args[0].ID()).V2D(), 0);
umov(x1, GetSrc(IROp->Args[0].ID()).V2D(), 1);
umov(x2, GetSrc(IROp->Args[1].ID()).V2D(), 0);
umov(x3, GetSrc(IROp->Args[1].ID()).V2D(), 1);
LoadConstant(x4, (uintptr_t)Info.fn);
blr(x4);
PopDynamicRegsAndLR();
FillStaticRegs();
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
ins(GetDst(Node).V2D(), 0, x0);
ins(GetDst(Node).V8H(), 4, w1);
}
break;
case FABI_UNKNOWN:
default:
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Fallback abi: %s %d", std::string(Name).c_str(), Info.ABI);
}
}
}
void JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
@@ -294,7 +548,7 @@ bool JITCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSig
memcpy(guest_uctx->__fpregs_mem._xmm, State->State.State.xmm, sizeof(State->State.State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = 0x37F;
guest_uctx->__fpregs_mem.fcw = State->State.State.FCW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
@@ -395,6 +649,8 @@ bool JITCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
PC[-1] = DMB;
PC[0] = LDR;
PC[1] = DMB;
// Back up one instruction and have another go
_mcontext->pc -= 4;
}
else if ( (Instr & 0x3F'FF'FC'00) == 0x08'9F'FC'00) { // STLR*
uint32_t STR = 0b0011'1000'0011'1111'0110'1000'0000'0000;
@@ -404,15 +660,48 @@ bool JITCore::HandleSIGBUS(int Signal, void *info, void *ucontext) {
PC[-1] = DMB;
PC[0] = STR;
PC[1] = DMB;
// Back up one instruction and have another go
_mcontext->pc -= 4;
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASPAL_MASK) == FEXCore::ArchHelpers::Arm64::CASPAL_INST) { // CASPAL
if (FEXCore::ArchHelpers::Arm64::HandleCASPAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASPAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::CASAL_MASK) == FEXCore::ArchHelpers::Arm64::CASAL_INST) { // CASAL
if (FEXCore::ArchHelpers::Arm64::HandleCASAL(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS CASAL: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
}
else if ((Instr & FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_MASK) == FEXCore::ArchHelpers::Arm64::ATOMIC_MEM_INST) { // Atomic memory op
if (FEXCore::ArchHelpers::Arm64::HandleAtomicMemOp(_mcontext, info, Instr)) {
// Skip this instruction now
_mcontext->pc += 4;
return true;
}
else {
uint8_t Op = (PC[0] >> 12) & 0xF;
LogMan::Msg::E("Unhandled JIT SIGBUS Atomic mem op 0x%02x: PC: %p Instruction: 0x%08x\n", Op, PC, PC[0]);
return false;
}
}
else {
LogMan::Msg::E("Unhandled JIT SIGBUS: PC: %p Instruction: 0x%08x\n", PC, PC[0]);
return false;
}
// Back up one instruction and have another go
_mcontext->pc -= 4;
vixl::aarch64::CPU::EnsureIAndDCacheCoherency(&PC[-1], 16);
return true;
}
@@ -658,27 +947,21 @@ void JITCore::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant) {
}
}
struct PhysReg { uint32_t Class; uint32_t VId; };
static IR::PhysicalRegister GetPhys(IR::RegisterAllocationData *RAData, uint32_t Node) {
auto PhyReg = RAData->GetNodeRegister(Node);
static PhysReg GetPhys(IR::RegisterAllocationPass *RAPass, uint32_t Node) {
uint64_t Reg = RAPass->GetNodeRegister(Node);
auto rv = PhysReg {uint32_t(Reg>>32), (uint32_t)Reg};
LogMan::Throw::A(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa%d. Class: %d", Node, PhyReg.Class);
if (rv.VId != ~0U)
return rv;
else
LogMan::Msg::A("Couldn't Allocate register for node: ssa%d. Class: %d", Node, Reg >> 32);
return PhysReg { ~0U, ~0U};
return PhyReg;
}
template<>
aarch64::Register JITCore::GetReg<JITCore::RA_32>(uint32_t Node) {
auto Reg = GetPhys(RAPass, Node);
auto Reg = GetPhys(RAData, Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
return SRA64[Reg.VId].W();
return SRA64[Reg.Reg].W();
} else if (Reg.Class == IR::GPRClass.Val) {
return RA64[Reg.VId].W();
return RA64[Reg.Reg].W();
} else {
LogMan::Throw::A(false, "Unexpected Class: %d", Reg.Class);
}
@@ -686,11 +969,11 @@ auto Reg = GetPhys(RAPass, Node);
template<>
aarch64::Register JITCore::GetReg<JITCore::RA_64>(uint32_t Node) {
auto Reg = GetPhys(RAPass, Node);
auto Reg = GetPhys(RAData, Node);
if (Reg.Class == IR::GPRFixedClass.Val) {
return SRA64[Reg.VId];
return SRA64[Reg.Reg];
} else if (Reg.Class == IR::GPRClass.Val) {
return RA64[Reg.VId];
return RA64[Reg.Reg];
} else {
LogMan::Throw::A(false, "Unexpected Class: %d", Reg.Class);
}
@@ -698,33 +981,33 @@ aarch64::Register JITCore::GetReg<JITCore::RA_64>(uint32_t Node) {
template<>
std::pair<aarch64::Register, aarch64::Register> JITCore::GetSrcPair<JITCore::RA_32>(uint32_t Node) {
uint32_t Reg = GetPhys(RAPass, Node).VId;
uint32_t Reg = GetPhys(RAData, Node).Reg;
return RA32Pair[Reg];
}
template<>
std::pair<aarch64::Register, aarch64::Register> JITCore::GetSrcPair<JITCore::RA_64>(uint32_t Node) {
uint32_t Reg = GetPhys(RAPass, Node).VId;
uint32_t Reg = GetPhys(RAData, Node).Reg;
return RA64Pair[Reg];
}
aarch64::VRegister JITCore::GetSrc(uint32_t Node) {
auto Reg = GetPhys(RAPass, Node);
auto Reg = GetPhys(RAData, Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
return SRAFPR[Reg.VId];
return SRAFPR[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
return RAFPR[Reg.VId];
return RAFPR[Reg.Reg];
} else {
LogMan::Throw::A(false, "Unexpected Class: %d", Reg.Class);
}
}
aarch64::VRegister JITCore::GetDst(uint32_t Node) {
auto Reg = GetPhys(RAPass, Node);
auto Reg = GetPhys(RAData, Node);
if (Reg.Class == IR::FPRFixedClass.Val) {
return SRAFPR[Reg.VId];
return SRAFPR[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
return RAFPR[Reg.VId];
return RAFPR[Reg.Reg];
} else {
LogMan::Throw::A(false, "Unexpected Class: %d", Reg.Class);
}
@@ -744,9 +1027,22 @@ bool JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Va
}
}
bool JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = IR->GetHeader()->Entry + Op->Offset;
}
return true;
} else {
return false;
}
}
FEXCore::IR::RegisterClassType JITCore::GetRegClass(uint32_t Node) {
auto Class = static_cast<uint32_t>(RAPass->GetNodeRegister(Node) >> 32);
return FEXCore::IR::RegisterClassType {Class};
return FEXCore::IR::RegisterClassType {GetPhys(RAData, Node).Class};
}
@@ -762,11 +1058,13 @@ bool JITCore::IsGPR(uint32_t Node) {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData) {
void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
using namespace aarch64;
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->RAData = RAData;
auto HeaderOp = IR->GetHeader();
#ifndef NDEBUG
@@ -778,7 +1076,7 @@ void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
if ((GetCursorOffset() + BufferRange) > CurrentCodeBuffer->Size) {
State->CTX->ClearCodeCache(State, HeaderOp->Entry);
State->CTX->ClearCodeCache(State, false);
}
// AAPCS64
@@ -828,71 +1126,67 @@ void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const
bind(&RunBlock);
}
if (HeaderOp->ShouldInterpret) {
// Make sure RIP is syncronized to the context
LoadConstant(x0, HeaderOp->Entry);
str(x0, MemOperand(STATE, offsetof(FEXCore::Core::ThreadState, State.rip)));
//LogMan::Throw::A(RAData->HasFullRA(), "Arm64 JIT only works with RA");
LoadConstant(x0, ThreadSharedData.InterpreterFallbackHelperAddress);
br(x0);
} else {
LogMan::Throw::A(RAPass->HasFullRA(), "Arm64 JIT only works with RA");
SpillSlots = RAData->SpillSlots();
SpillSlots = RAPass->SpillSlots();
if (SpillSlots) {
if (SpillSlots) {
if (IsImmAddSub(SpillSlots * 16)) {
sub(sp, sp, SpillSlots * 16);
} else {
LoadConstant(x0, SpillSlots * 16);
sub(sp, sp, x0);
}
PendingTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
{
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
uint32_t ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()) - DebugData->Subblocks.back().HostCodeStart;
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel)
{
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
}
PendingTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
{
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
bind(&IsTarget->second);
}
if (DebugData) {
DebugData->Subblocks.push_back({Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()), 0, IR->GetID(BlockNode)});
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
uint32_t ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
if (DebugData) {
DebugData->Subblocks.back().HostCodeSize = Buffer->GetOffsetAddress<uintptr_t>(GetCursorOffset()) - DebugData->Subblocks.back().HostCodeStart;
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel)
{
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
FinalizeCode();
auto CodeEnd = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
@@ -991,7 +1285,7 @@ void JITCore::PopCalleeSavedRegisters() {
uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadState *Thread, uint64_t *record) {
auto GuestRip = record[1];
auto HostCode = Thread->BlockCache->FindBlock(GuestRip);
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
//printf("ExitFunctionLink: Aborting, %lX not in cache\n", GuestRip);
@@ -1007,13 +1301,15 @@ uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadS
// optimal case - can branch directly
// patch the code
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
emit.b(offset);
emit.FinalizeCode();
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->BlockCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
emit.ldr(x0, &l_BranchHost);
emit.blr(x0);
@@ -1026,7 +1322,7 @@ uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadS
record[0] = HostCode;
// Add de-linking handler
Thread->BlockCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
}
@@ -1066,15 +1362,13 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
// }
uint64_t VirtualMemorySize = Thread->BlockCache->GetVirtualMemorySize();
uint64_t VirtualMemorySize = Thread->LookupCache->GetVirtualMemorySize();
Literal l_VirtualMemory {VirtualMemorySize};
Literal l_PagePtr {Thread->BlockCache->GetPagePointer()};
Literal l_PagePtr {Thread->LookupCache->GetPagePointer()};
Literal l_CTX {reinterpret_cast<uintptr_t>(CTX)};
Literal l_Interpreter {reinterpret_cast<uint64_t>(State->IntBackend->CompileCode(nullptr, nullptr))};
Literal l_Sleep {reinterpret_cast<uint64_t>(SleepThread)};
uintptr_t CompileBlockPtr{};
uintptr_t CompileFallbackPtr{};
{
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
@@ -1086,20 +1380,8 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
Ptr.ClassPtr = &FEXCore::Context::Context::CompileBlock;
CompileBlockPtr = Ptr.Data;
}
{
using ClassPtrType = uintptr_t (FEXCore::Context::Context::*)(FEXCore::Core::InternalThreadState *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast Ptr;
Ptr.ClassPtr = &FEXCore::Context::Context::CompileFallbackBlock;
CompileFallbackPtr = Ptr.Data;
}
Literal l_CompileBlock {CompileBlockPtr};
Literal l_CompileFallback {CompileFallbackPtr};
Literal l_ExitFunctionLink {(uintptr_t)&ExitFunctionLink};
// Push all the register we need to save
@@ -1123,9 +1405,7 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
};
// used from signals
aarch64::Label LoopTopFillSRA{};
bind(&LoopTopFillSRA);
AbsoluteLoopTopAddressFillSRA = GetLabelAddress<uint64_t>(&LoopTopFillSRA);
AbsoluteLoopTopAddressFillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
FillStaticRegs();
@@ -1134,8 +1414,6 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
aarch64::Label FullLookup{};
aarch64::Label LoopTop{};
aarch64::Label ExitSpillSRA{};
aarch64::Label ThreadPauseHandlerSpillSRA{};
aarch64::Label ExitFunctionLinker{};
bind(&LoopTop);
AbsoluteLoopTopAddress = GetLabelAddress<uint64_t>(&LoopTop);
@@ -1146,9 +1424,9 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
auto RipReg = x2;
// L1 Cache
LoadConstant(x0, Thread->BlockCache->GetL1Pointer());
LoadConstant(x0, Thread->LookupCache->GetL1Pointer());
and_(x3, RipReg, BlockCache::L1_ENTRIES_MASK);
and_(x3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x3, Shift::LSL, 4));
ldp(x1, x0, MemOperand(x0));
cmp(x0, RipReg);
@@ -1159,12 +1437,12 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it BlockCache.h::FindBlock
// It matches what is going on it LookupCache.h::FindBlock
ldr(x0, &l_PagePtr);
// Mask the address by the virtual address size so we can check for aliases
if (__builtin_popcountl(VirtualMemorySize) == 1) {
and_(x3, RipReg, Thread->BlockCache->GetVirtualMemorySize() - 1);
and_(x3, RipReg, Thread->LookupCache->GetVirtualMemorySize() - 1);
}
else {
ldr(x3, &l_VirtualMemory);
@@ -1186,24 +1464,24 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
and_(x1, x3, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::BlockCache::BlockCacheEntry))));
add(x0, x0, Operand(x1, Shift::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry))));
// Load the guest address first to ensure it maps to the address we are currently at
// This fixes aliasing problems
ldr(x1, MemOperand(x0, offsetof(FEXCore::BlockCache::BlockCacheEntry, GuestCode)));
ldr(x1, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, GuestCode)));
cmp(x1, RipReg);
b(&NoBlock, Condition::ne);
// Now load the actual host block to execute if we can
ldr(x3, MemOperand(x0, offsetof(FEXCore::BlockCache::BlockCacheEntry, HostCode)));
ldr(x3, MemOperand(x0, offsetof(FEXCore::LookupCache::LookupCacheEntry, HostCode)));
cbz(x3, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
LoadConstant(x0, Thread->BlockCache->GetL1Pointer());
LoadConstant(x0, Thread->LookupCache->GetL1Pointer());
and_(x1, RipReg, BlockCache::L1_ENTRIES_MASK);
and_(x1, RipReg, LookupCache::L1_ENTRIES_MASK);
add(x0, x0, Operand(x1, Shift::LSL, 4));
stp(x3, x2, MemOperand(x0));
br(x3);
@@ -1211,7 +1489,7 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
}
{
b(&ExitSpillSRA);
bind(&ExitSpillSRA);
ThreadStopHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
SpillStaticRegs();
ThreadStopHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
@@ -1224,8 +1502,7 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
}
{
bind(&ExitFunctionLinker);
ExitFunctionLinkerAddress = GetLabelAddress<uint64_t>(&ExitFunctionLinker);
ExitFunctionLinkerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
SpillStaticRegs();
@@ -1240,7 +1517,6 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
br(x0);
}
aarch64::Label FallbackCore;
// Need to create the block
{
bind(&NoBlock);
@@ -1254,31 +1530,12 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
blr(x3); // { CTX, ThreadState, RIP}
FillStaticRegs();
// X0 now contains either nullptr or block pointer
cbz(x0, &FallbackCore);
// X0 now contains the block pointer
blr(x0);
b(&LoopTop);
}
// We need to fallback to our fallback core
{
bind(&FallbackCore);
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x3, &l_CompileFallback);
// X2 contains our guest RIP
SpillStaticRegs();
blr(x3); // {ThreadState, RIP}
FillStaticRegs();
// X0 now contains either nullptr or block pointer
cbz(x0, &ExitSpillSRA);
br(x0);
}
{
Label RestoreContextStateHelperLabel{};
bind(&RestoreContextStateHelperLabel);
@@ -1290,20 +1547,6 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
}
{
Label InterpreterFallback{};
bind(&InterpreterFallback);
ThreadSharedData.InterpreterFallbackHelperAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
SpillStaticRegs();
mov(x0, STATE);
ldr(x1, &l_Interpreter);
blr(x1);
FillStaticRegs();
b(&LoopTop);
}
{
bind(&ThreadPauseHandlerSpillSRA);
ThreadPauseHandlerAddressSpillSRA = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
SpillStaticRegs();
ThreadPauseHandlerAddress = Buffer->GetOffsetAddress<uint64_t>(GetCursorOffset());
@@ -1377,10 +1620,8 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
place(&l_VirtualMemory);
place(&l_PagePtr);
place(&l_CTX);
place(&l_Interpreter);
place(&l_Sleep);
place(&l_CompileBlock);
place(&l_CompileFallback);
place(&l_ExitFunctionLink);
FinalizeCode();
@@ -1461,6 +1702,19 @@ void JITCore::PopDynamicRegsAndLR() {
add(sp, sp, SPOffset);
}
void JITCore::ResetStack() {
if (SpillSlots == 0)
return;
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
add(sp, sp, x0);
}
}
FEXCore::CPU::CPUBackend *CreateJITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, bool CompileThread) {
return new JITCore(ctx, Thread, JITCore::AllocateNewCodeBuffer(JITCore::INITIAL_CODE_SIZE), CompileThread);
}
@@ -1,6 +1,6 @@
#pragma once
#include "Interface/Core/BlockCache.h"
#include "Interface/Core/LookupCache.h"
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
@@ -78,7 +78,7 @@ public:
~JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) override;
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -100,7 +100,7 @@ private:
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *State;
FEXCore::IR::IRListView<true> const *IR;
FEXCore::IR::IRListView const *IR;
std::map<IR::OrderedNodeWrapper::NodeOffsetType, aarch64::Label> JumpTargets;
@@ -152,6 +152,7 @@ private:
MemOperand GenerateMemOperand(uint8_t AccessSize, aarch64::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr);
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value);
struct LiveRange {
uint32_t Begin;
@@ -223,8 +224,6 @@ private:
/** @} */
struct CompilerSharedData {
uint64_t InterpreterFallbackHelperAddress{};
uint64_t SignalReturnInstruction{};
uint32_t *SignalHandlerRefCounterPtr{};
@@ -232,6 +231,7 @@ private:
CompilerSharedData ThreadSharedData;
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
uint32_t SpillSlots{};
@@ -241,6 +241,8 @@ private:
void PushDynamicRegsAndLR();
void PopDynamicRegsAndLR();
void ResetStack();
using OpHandler = void (JITCore::*)(FEXCore::IR::IROp_Header *IROp, uint32_t Node);
std::array<OpHandler, FEXCore::IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
@@ -264,7 +266,9 @@ private:
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
DEF_OP(CycleCounter);
DEF_OP(Add);
DEF_OP(Sub);
@@ -308,7 +312,6 @@ private:
DEF_OP(Float_ToGPR_U);
DEF_OP(Float_ToGPR_S);
DEF_OP(FCmp);
DEF_OP(F80Cmp);
///< Atomic ops
DEF_OP(CASPair);
@@ -41,9 +41,7 @@ DEF_OP(Break) {
break;
}
case 6: { // INT3
if (SpillSlots) {
add(sp, sp, SpillSlots * 16);
}
ResetStack();
LoadConstant(TMP1, ThreadPauseHandlerAddressSpillSRA);
br(TMP1);
@@ -29,18 +29,21 @@ DEF_OP(CreateElementPair) {
std::pair<aarch64::Register, aarch64::Register> Dst;
aarch64::Register RegFirst;
aarch64::Register RegSecond;
aarch64::Register RegTmp;
switch (Op->Header.Size) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetReg<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_32>(Op->Header.Args[1].ID());
RegTmp = w0;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetReg<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetReg<RA_64>(Op->Header.Args[1].ID());
RegTmp = x0;
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
@@ -53,7 +56,9 @@ DEF_OP(CreateElementPair) {
mov(Dst.second, RegSecond);
mov(Dst.first, RegFirst);
} else {
LogMan::Msg::A("Unhandled CreateElementPair");
mov(RegTmp, RegFirst);
mov(Dst.second, RegSecond);
mov(Dst.first, RegTmp);
}
}
@@ -909,7 +909,17 @@ DEF_OP(VZip2) {
}
DEF_OP(VBSL) {
LogMan::Msg::A("Unimplemented");
auto Op = IROp->C<IR::IROp_VBSL>();
if (IROp->Size == 16) {
mov(VTMP1.V16B(), GetSrc(Op->Header.Args[0].ID()).V16B());
bsl(VTMP1.V16B(), GetSrc(Op->Header.Args[1].ID()).V16B(), GetSrc(Op->Header.Args[2].ID()).V16B());
mov(GetDst(Node).V16B(), VTMP1.V16B());
}
else {
mov(VTMP1.V8B(), GetSrc(Op->Header.Args[0].ID()).V8B());
bsl(VTMP1.V8B(), GetSrc(Op->Header.Args[1].ID()).V8B(), GetSrc(Op->Header.Args[2].ID()).V8B());
mov(GetDst(Node).V8B(), VTMP1.V8B());
}
}
DEF_OP(VCMPEQ) {
+36 -23
View File
@@ -23,17 +23,28 @@ DEF_OP(Constant) {
mov(GetDst<RA_64>(Node), Op->Constant);
}
DEF_OP(EntrypointOffset) {
auto Op = IROp->C<IR::IROp_EntrypointOffset>();
auto Constant = IR->GetHeader()->Entry + Op->Offset;
mov(GetDst<RA_64>(Node), Constant);
}
DEF_OP(InlineConstant) {
//nop
}
DEF_OP(InlineEntrypointOffset) {
//nop
}
DEF_OP(CycleCounter) {
#ifdef DEBUG_CYCLES
mov (GetDst<RA_64>(Node), 0);
#else
rdtsc();
shl(rdx, 32);
or(rax, rdx);
or_(rax, rdx);
mov (GetDst<RA_64>(Node), rax);
#endif
}
@@ -373,9 +384,9 @@ DEF_OP(Or) {
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
or (rax, Const);
or_(rax, Const);
} else {
or (rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -386,9 +397,9 @@ DEF_OP(And) {
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
and (rax, Const);
and_(rax, Const);
} else {
and(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -399,9 +410,9 @@ DEF_OP(Xor) {
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
uint64_t Const;
if (IsInlineConstant(Op->Header.Args[1], &Const)) {
xor(rax, Const);
xor_(rax, Const);
} else {
xor(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(rax, GetSrc<RA_64>(Op->Header.Args[1].ID()));
}
mov(Dst, rax);
}
@@ -428,7 +439,7 @@ DEF_OP(Lshl) {
};
} else {
mov(rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 4:
@@ -476,7 +487,7 @@ DEF_OP(Lshr) {
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 1:
@@ -534,7 +545,7 @@ DEF_OP(Ashr) {
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 1:
movsx(rax, GetSrc<RA_8>(Op->Header.Args[0].ID()));
@@ -583,7 +594,7 @@ DEF_OP(Ror) {
}
} else {
mov (rcx, GetSrc<RA_64>(Op->Header.Args[1].ID()));
and(rcx, Mask);
and_(rcx, Mask);
switch (OpSize) {
case 4: {
mov(eax, GetSrc<RA_32>(Op->Header.Args[0].ID()));
@@ -790,10 +801,10 @@ DEF_OP(FindLSB) {
bsf(rcx, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, 0x40);
cmovz(rcx, rax);
xor(rax, rax);
xor_(rax, rax);
cmp(GetSrc<RA_64>(Op->Header.Args[0].ID()), 1);
sbb(rax, rax);
or(rax, rcx);
or_(rax, rcx);
mov (GetDst<RA_64>(Node), rax);
}
@@ -870,7 +881,7 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(ax, GetSrc<RA_16>(Op->Header.Args[0].ID()));
xor(ax, 0xF);
xor_(ax, 0xF);
movzx(eax, ax);
L(Skip);
mov(GetDst<RA_32>(Node), eax);
@@ -882,7 +893,7 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(eax, GetSrc<RA_32>(Op->Header.Args[0].ID()));
xor(eax, 0x1F);
xor_(eax, 0x1F);
L(Skip);
mov(GetDst<RA_32>(Node), eax);
break;
@@ -893,7 +904,7 @@ DEF_OP(CountLeadingZeroes) {
Label Skip;
je(Skip);
bsr(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
xor(rax, 0x3F);
xor_(rax, 0x3F);
L(Skip);
mov(GetDst<RA_64>(Node), rax);
break;
@@ -939,15 +950,15 @@ DEF_OP(Bfi) {
mov(Dst, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(TMP2, DestMask);
and(Dst, TMP2);
and_(Dst, TMP2);
mov(TMP2, SourceMask);
and(TMP1, TMP2);
and_(TMP1, TMP2);
shl(TMP1, Op->lsb);
or_(Dst, TMP1);
if (OpSize != 8) {
mov(rcx, uint64_t((1ULL << (OpSize * 8)) - 1));
and(Dst, rcx);
and_(Dst, rcx);
}
}
@@ -987,7 +998,7 @@ DEF_OP(Bfe) {
if (Op->Width != 64) {
mov(rcx, uint64_t((1ULL << Op->Width) - 1));
and(Dst, rcx);
and_(Dst, rcx);
}
}
@@ -1136,21 +1147,21 @@ DEF_OP(FCmp) {
mov(rcx, 0);
setb(cl);
shl(rcx, IR::FCMP_FLAG_LT);
or(rdx, rcx);
or_(rdx, rcx);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_UNORDERED)) {
sahf();
mov(rcx, 0);
setp(cl);
shl(rcx, IR::FCMP_FLAG_UNORDERED);
or(rdx, rcx);
or_(rdx, rcx);
}
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ)) {
sahf();
mov(rcx, 0);
setz(cl);
shl(rcx, IR::FCMP_FLAG_EQ);
or(rdx, rcx);
or_(rdx, rcx);
}
mov (GetDst<RA_64>(Node), rdx);
}
@@ -1161,7 +1172,9 @@ void JITCore::RegisterALUHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &JITCore::Op_##x
REGISTER_OP(TRUNCELEMENTPAIR, TruncElementPair);
REGISTER_OP(CONSTANT, Constant);
REGISTER_OP(ENTRYPOINTOFFSET, EntrypointOffset);
REGISTER_OP(INLINECONSTANT, InlineConstant);
REGISTER_OP(INLINEENTRYPOINTOFFSET, InlineEntrypointOffset);
REGISTER_OP(CYCLECOUNTER, CycleCounter);
REGISTER_OP(ADD, Add);
REGISTER_OP(SUB, Sub);
@@ -155,16 +155,16 @@ DEF_OP(AtomicAnd) {
lock();
switch (Op->Size) {
case 1:
and(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
and(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
and(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
and(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
}
@@ -177,16 +177,16 @@ DEF_OP(AtomicOr) {
lock();
switch (Op->Size) {
case 1:
or(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
or(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
or(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
or(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
}
@@ -199,16 +199,16 @@ DEF_OP(AtomicXor) {
lock();
switch (Op->Size) {
case 1:
xor(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(byte [MemReg], GetSrc<RA_8>(Op->Header.Args[1].ID()));
break;
case 2:
xor(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(word [MemReg], GetSrc<RA_16>(Op->Header.Args[1].ID()));
break;
case 4:
xor(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(dword [MemReg], GetSrc<RA_32>(Op->Header.Args[1].ID()));
break;
case 8:
xor(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(qword [MemReg], GetSrc<RA_64>(Op->Header.Args[1].ID()));
break;
default: LogMan::Msg::A("Unhandled AtomicAdd size: %d", Op->Size);
}
@@ -329,7 +329,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
and(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
and_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -345,7 +345,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
and(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
and_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -362,7 +362,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
and(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
and_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -379,7 +379,7 @@ DEF_OP(AtomicFetchAnd) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
and(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
and_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -406,7 +406,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
or(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
or_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -422,7 +422,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
or(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
or_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -439,7 +439,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
or(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
or_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -456,7 +456,7 @@ DEF_OP(AtomicFetchOr) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
or(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
or_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -483,7 +483,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt8(), TMP1.cvt8());
mov(TMP3.cvt8(), TMP1.cvt8());
xor(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt8(), GetSrc<RA_8>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(byte [MemReg], TMP2.cvt8());
@@ -499,7 +499,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt16(), TMP1.cvt16());
mov(TMP3.cvt16(), TMP1.cvt16());
xor(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt16(), GetSrc<RA_16>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(word [MemReg], TMP2.cvt16());
@@ -516,7 +516,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt32(), TMP1.cvt32());
mov(TMP3.cvt32(), TMP1.cvt32());
xor(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt32(), GetSrc<RA_32>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(dword [MemReg], TMP2.cvt32());
@@ -533,7 +533,7 @@ DEF_OP(AtomicFetchXor) {
L(Loop);
mov(TMP2.cvt64(), TMP1.cvt64());
mov(TMP3.cvt64(), TMP1.cvt64());
xor(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
xor_(TMP2.cvt64(), GetSrc<RA_64>(Op->Header.Args[1].ID()));
// Updates RAX with the value from memory
lock(); cmpxchg(qword [MemReg], TMP2.cvt64());
@@ -3,6 +3,7 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void JITCore::Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
@@ -66,7 +67,7 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP)) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
Label l_BranchHost;
Label l_BranchGuest;
@@ -81,10 +82,10 @@ DEF_OP(ExitFunction) {
Xbyak::Reg RipReg = GetSrc<RA_64>(Op->NewRIP.ID());
// L1 Cache
mov(rcx, ThreadState->BlockCache->GetL1Pointer());
mov(rcx, ThreadState->LookupCache->GetL1Pointer());
mov(rax, RipReg);
and_(rax, BlockCache::L1_ENTRIES_MASK);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
Xbyak::RegExp LookupBase = rcx + rax;
@@ -226,8 +227,10 @@ DEF_OP(Thunk) {
sub(rsp, 8); // Align
mov(rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
mov(rax, reinterpret_cast<uintptr_t>(Op->ThunkFnPtr));
mov(rax, reinterpret_cast<uintptr_t>(thunkFn));
call(rax);
if (NumPush & 1)
@@ -239,12 +242,12 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
uint8_t* OldCode = (uint8_t*)&Op->CodeOriginal;
uint8_t* OldCode = (uint8_t*)&Op->CodeOriginalLow;
int len = Op->CodeLength;
int idx = 0;
xor_(GetDst<RA_64>(Node), GetDst<RA_64>(Node));
mov(rax, Op->CodePtr);
mov(rax, IR->GetHeader()->Entry + Op->Offset);
mov(rbx, 1);
while (len >= 4) {
cmp(dword[rax + idx], *(uint32_t*)(OldCode + idx));
@@ -279,7 +282,7 @@ DEF_OP(RemoveCodeEntry) {
sub(rsp, 8); // Align
mov(rdi, STATE);
mov(rax, Op->RIP); // imm64 move
mov(rax, IR->GetHeader()->Entry); // imm64 move
mov(rsi, rax);
@@ -9,7 +9,7 @@ DEF_OP(GetHostFlag) {
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
shr(rax, Op->Flag);
and(rax, 1);
and_(rax, 1);
mov(GetDst<RA_64>(Node), rax);
}
+408 -185
View File
@@ -11,6 +11,8 @@
#include <cmath>
#include <signal.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
// #define DEBUG_RA 1
// #define DEBUG_CYCLES
@@ -194,7 +196,7 @@ bool JITCore::HandleGuestSignal(int Signal, void *info, void *ucontext, GuestSig
memcpy(guest_uctx->__fpregs_mem._xmm, ThreadState->State.State.xmm, sizeof(ThreadState->State.State.xmm));
// FCW store default
guest_uctx->__fpregs_mem.fcw = 0x37F;
guest_uctx->__fpregs_mem.fcw = ThreadState->State.State.FCW;
// Reconstruct FSW
guest_uctx->__fpregs_mem.fsw =
@@ -337,9 +339,239 @@ bool JITCore::HandleSignalPause(int Signal, void *info, void *ucontext) {
return false;
}
void JITCore::PushRegs() {
for (auto &Xmm : RAXMM_x) {
sub(rsp, 16);
movaps(ptr[rsp], Xmm);
}
for (auto &Reg : RA64)
push(Reg);
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
}
void JITCore::PopRegs() {
auto NumPush = RA64.size();
if (NumPush & 1)
add(rsp, 8); // Align
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
for (uint32_t i = RAXMM_x.size(); i > 0; --i) {
movaps(RAXMM_x[i - 1], ptr[rsp]);
add(rsp, 16);
}
}
void JITCore::Op_Unhandled(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Op: %s", std::string(Name).c_str());
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Op: %s", std::string(Name).c_str());
} else {
switch(Info.ABI) {
case FABI_VOID_U16: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
break;
}
case FABI_F80_F32:{
PushRegs();
movss(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
pxor(GetDst(Node), GetDst(Node));
movq(GetDst(Node), rax);
pinsrw(GetDst(Node), edx, 4);
}
break;
case FABI_F80_F64:{
PushRegs();
movsd(xmm0, GetSrc(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
pxor(GetDst(Node), GetDst(Node));
movq(GetDst(Node), rax);
pinsrw(GetDst(Node), edx, 4);
}
break;
case FABI_F80_I16:
case FABI_F80_I32: {
PushRegs();
mov(edi, GetSrc<RA_32>(IROp->Args[0].ID()));
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
pxor(GetDst(Node), GetDst(Node));
movq(GetDst(Node), rax);
pinsrw(GetDst(Node), edx, 4);
}
break;
case FABI_F32_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
movss(GetDst(Node), xmm0);
}
break;
case FABI_F64_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
movsd(GetDst(Node), xmm0);
}
break;
case FABI_I16_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
movzx(GetDst<RA_64>(Node), ax);
}
break;
case FABI_I32_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
mov(GetDst<RA_32>(Node), eax);
}
break;
case FABI_I64_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
mov(GetDst<RA_64>(Node), rax);
}
break;
case FABI_I64_F80_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
mov(GetDst<RA_64>(Node), rax);
}
break;
case FABI_F80_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
pxor(GetDst(Node), GetDst(Node));
movq(GetDst(Node), rax);
pinsrw(GetDst(Node), edx, 4);
}
break;
case FABI_F80_F80_F80:{
PushRegs();
movq(rdi, GetSrc(IROp->Args[0].ID()));
pextrq(rsi, GetSrc(IROp->Args[0].ID()), 1);
movq(rdx, GetSrc(IROp->Args[1].ID()));
pextrq(rcx, GetSrc(IROp->Args[1].ID()), 1);
mov(rax, (uintptr_t)Info.fn);
call(rax);
PopRegs();
pxor(GetDst(Node), GetDst(Node));
movq(GetDst(Node), rax);
pinsrw(GetDst(Node), edx, 4);
}
break;
case FABI_UNKNOWN:
default:
auto Name = FEXCore::IR::GetName(IROp->Op);
LogMan::Msg::A("Unhandled IR Fallback abi: %s %d", std::string(Name).c_str(), Info.ABI);
}
}
}
void JITCore::Op_NoOp(FEXCore::IR::IROp_Header *IROp, uint32_t Node) {
@@ -460,27 +692,20 @@ void JITCore::ClearCache() {
}
}
uint32_t JITCore::GetPhys(uint32_t Node) {
uint64_t Reg = RAPass->GetNodeRegister(Node);
IR::PhysicalRegister JITCore::GetPhys(uint32_t Node) {
auto PhyReg = RAData->GetNodeRegister(Node);
if ((uint32_t)Reg != ~0U)
return Reg;
else
LogMan::Msg::A("Couldn't Allocate register for node: ssa%d. Class: %d", Node, Reg >> 32);
LogMan::Throw::A(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa%d. Class: %d", Node, PhyReg.Class);
return ~0U;
return PhyReg;
}
bool JITCore::IsFPR(uint32_t Node) {
auto Class = RAPass->GetNodeRegister(Node) >> 32;
return Class == IR::FPRClass.Val;
return RAData->GetNodeRegister(Node).Class == IR::FPRClass.Val;
}
bool JITCore::IsGPR(uint32_t Node) {
auto Class = RAPass->GetNodeRegister(Node) >> 32;
return Class == IR::GPRClass.Val;
return RAData->GetNodeRegister(Node).Class == IR::GPRClass.Val;
}
template<uint8_t RAType>
@@ -489,17 +714,17 @@ Xbyak::Reg JITCore::GetSrc(uint32_t Node) {
// r10
// Callee Saved
// rbx, rbp, r12, r13, r14, r15
uint32_t Reg = GetPhys(Node);
auto PhyReg = GetPhys(Node);
if (RAType == RA_64)
return RA64[Reg].cvt64();
return RA64[PhyReg.Reg].cvt64();
else if (RAType == RA_XMM)
return RAXMM[Reg];
return RAXMM[PhyReg.Reg];
else if (RAType == RA_32)
return RA64[Reg].cvt32();
return RA64[PhyReg.Reg].cvt32();
else if (RAType == RA_16)
return RA64[Reg].cvt16();
return RA64[PhyReg.Reg].cvt16();
else if (RAType == RA_8)
return RA64[Reg].cvt8();
return RA64[PhyReg.Reg].cvt8();
}
template
@@ -515,23 +740,23 @@ template
Xbyak::Reg JITCore::GetSrc<JITCore::RA_8>(uint32_t Node);
Xbyak::Xmm JITCore::GetSrc(uint32_t Node) {
uint32_t Reg = GetPhys(Node);
return RAXMM_x[Reg];
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
template<uint8_t RAType>
Xbyak::Reg JITCore::GetDst(uint32_t Node) {
uint32_t Reg = GetPhys(Node);
auto PhyReg = GetPhys(Node);
if (RAType == RA_64)
return RA64[Reg].cvt64();
return RA64[PhyReg.Reg].cvt64();
else if (RAType == RA_XMM)
return RAXMM[Reg];
return RAXMM[PhyReg.Reg];
else if (RAType == RA_32)
return RA64[Reg].cvt32();
return RA64[PhyReg.Reg].cvt32();
else if (RAType == RA_16)
return RA64[Reg].cvt16();
return RA64[PhyReg.Reg].cvt16();
else if (RAType == RA_8)
return RA64[Reg].cvt8();
return RA64[PhyReg.Reg].cvt8();
}
template
@@ -548,11 +773,11 @@ Xbyak::Reg JITCore::GetDst<JITCore::RA_8>(uint32_t Node);
template<uint8_t RAType>
std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair(uint32_t Node) {
uint32_t Reg = GetPhys(Node);
auto PhyReg = GetPhys(Node);
if (RAType == RA_64)
return RA64Pair[Reg];
return RA64Pair[PhyReg.Reg];
else if (RAType == RA_32)
return {RA64Pair[Reg].first.cvt32(), RA64Pair[Reg].second.cvt32()};
return {RA64Pair[PhyReg.Reg].first.cvt32(), RA64Pair[PhyReg.Reg].second.cvt32()};
}
template
@@ -562,8 +787,8 @@ template
std::pair<Xbyak::Reg, Xbyak::Reg> JITCore::GetSrcPair<JITCore::RA_32>(uint32_t Node);
Xbyak::Xmm JITCore::GetDst(uint32_t Node) {
uint32_t Reg = GetPhys(Node);
return RAXMM_x[Reg];
auto PhyReg = GetPhys(Node);
return RAXMM_x[PhyReg.Reg];
}
bool JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
@@ -580,6 +805,20 @@ bool JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Va
}
}
bool JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) {
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
if (OpHeader->Op == IR::IROps::OP_INLINEENTRYPOINTOFFSET) {
auto Op = OpHeader->C<IR::IROp_InlineEntrypointOffset>();
if (Value) {
*Value = IR->GetHeader()->Entry + Op->Offset;
}
return true;
} else {
return false;
}
}
std::tuple<JITCore::SetCC, JITCore::CMovCC, JITCore::JCC> JITCore::GetCC(IR::CondClassType cond) {
switch (cond.Val) {
case FEXCore::IR::COND_EQ: return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
@@ -587,7 +826,7 @@ std::tuple<JITCore::SetCC, JITCore::CMovCC, JITCore::JCC> JITCore::GetCC(IR::Con
case FEXCore::IR::COND_SGE: return { &CodeGenerator::setge, &CodeGenerator::cmovge, &CodeGenerator::jge };
case FEXCore::IR::COND_SLT: return { &CodeGenerator::setl , &CodeGenerator::cmovl , &CodeGenerator::jl };
case FEXCore::IR::COND_SGT: return { &CodeGenerator::setg , &CodeGenerator::cmovg , &CodeGenerator::jg };
case FEXCore::IR::COND_SLE: return { &CodeGenerator::setle, &CodeGenerator::cmovle, &CodeGenerator::jle };
case FEXCore::IR::COND_SLE: return { &CodeGenerator::setle, &CodeGenerator::cmovle, &CodeGenerator::jle };
case FEXCore::IR::COND_UGE: return { &CodeGenerator::setae, &CodeGenerator::cmovae, &CodeGenerator::jae };
case FEXCore::IR::COND_ULT: return { &CodeGenerator::setb , &CodeGenerator::cmovb , &CodeGenerator::jb };
case FEXCore::IR::COND_UGT: return { &CodeGenerator::seta , &CodeGenerator::cmova , &CodeGenerator::ja };
@@ -608,18 +847,22 @@ std::tuple<JITCore::SetCC, JITCore::CMovCC, JITCore::JCC> JITCore::GetCC(IR::Con
LogMan::Msg::A("Unsupported compare type");
break;
}
// Hope for the best
return { &CodeGenerator::sete , &CodeGenerator::cmove , &CodeGenerator::je };
}
void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData) {
void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
this->RAData = RAData;
auto HeaderOp = IR->GetHeader();
// Fairly excessive buffer range to make sure we don't overflow
uint32_t BufferRange = SSACount * 16;
if ((getSize() + BufferRange) > CurrentCodeBuffer->Size) {
ThreadState->CTX->ClearCodeCache(ThreadState, HeaderOp->Entry);
ThreadState->CTX->ClearCodeCache(ThreadState, false);
}
void *Entry = getCurr<void*>();
@@ -644,140 +887,132 @@ void *JITCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const
L(RunBlock);
}
if (HeaderOp->ShouldInterpret) {
mov(rax, HeaderOp->Entry);
mov(qword [STATE + offsetof(FEXCore::Core::CPUState, rip)], rax);
mov(rax, (uintptr_t)ThreadSharedData.InterpreterFallbackHelperAddress);
jmp(rax);
} else {
LogMan::Throw::A(RAPass->HasFullRA(), "Needs RA");
LogMan::Throw::A(RAData != nullptr, "Needs RA");
SpillSlots = RAPass->SpillSlots();
SpillSlots = RAData->SpillSlots();
if (SpillSlots) {
sub(rsp, SpillSlots * 16);
}
if (SpillSlots) {
sub(rsp, SpillSlots * 16);
}
#ifdef BLOCKSTATS
BlockSamplingData::BlockData *SamplingData = CTX->BlockData->GetBlockData(HeaderOp->Entry);
#ifdef BLOCKSTATS
BlockSamplingData::BlockData *SamplingData = CTX->BlockData->GetBlockData(HeaderOp->Entry);
if (GetSamplingData) {
mov(rcx, reinterpret_cast<uintptr_t>(SamplingData));
rdtsc();
shl(rdx, 32);
or_(rax, rdx);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Start)], rax);
}
auto ExitBlock = [&]() {
if (GetSamplingData) {
mov(rcx, reinterpret_cast<uintptr_t>(SamplingData));
// Get time
rdtsc();
shl(rdx, 32);
or(rax, rdx);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Start)], rax);
or_(rax, rdx);
// Calculate time spent in block
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Start)]);
sub(rax, rdx);
// Add time to total time
add(qword [rcx + offsetof(BlockSamplingData::BlockData, TotalTime)], rax);
// Increment call count
inc(qword [rcx + offsetof(BlockSamplingData::BlockData, TotalCalls)]);
// Calculate min
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Min)]);
cmp(rdx, rax);
cmova(rdx, rax);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Min)], rdx);
// Calculate max
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Max)]);
cmp(rdx, rax);
cmovb(rdx, rax);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Max)], rdx);
}
};
#endif
auto ExitBlock = [&]() {
if (GetSamplingData) {
mov(rcx, reinterpret_cast<uintptr_t>(SamplingData));
// Get time
rdtsc();
shl(rdx, 32);
or(rax, rdx);
PendingTargetLabel = nullptr;
// Calculate time spent in block
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Start)]);
sub(rax, rdx);
// Add time to total time
add(qword [rcx + offsetof(BlockSamplingData::BlockData, TotalTime)], rax);
// Increment call count
inc(qword [rcx + offsetof(BlockSamplingData::BlockData, TotalCalls)]);
// Calculate min
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Min)]);
cmp(rdx, rax);
cmova(rdx, rax);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Min)], rdx);
// Calculate max
mov(rdx, qword [rcx + offsetof(BlockSamplingData::BlockData, Max)]);
cmp(rdx, rax);
cmovb(rdx, rax);
mov(qword [rcx + offsetof(BlockSamplingData::BlockData, Max)], rdx);
}
};
#endif
PendingTargetLabel = nullptr;
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
{
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
// if there is a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
L(IsTarget->second);
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
#ifdef DEBUG_RA
if (IROp->Op != IR::OP_BEGINBLOCK &&
IROp->Op != IR::OP_CONDJUMP &&
IROp->Op != IR::OP_JUMP) {
std::stringstream Inst;
auto Name = FEXCore::IR::GetName(IROp->Op);
if (IROp->HasDest) {
uint64_t PhysReg = RAPass->GetNodeRegister(Node);
if (PhysReg >= GPRPairBase)
Inst << "\tPair" << GetPhys(Node) << " = " << Name << " ";
else if (PhysReg >= XMMBase)
Inst << "\tXMM" << GetPhys(Node) << " = " << Name << " ";
else
Inst << "\tReg" << GetPhys(Node) << " = " << Name << " ";
}
else {
Inst << "\t" << Name << " ";
}
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t ArgNode = IROp->Args[i].ID();
uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
if (PhysReg >= GPRPairBase)
Inst << "Pair" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else if (PhysReg >= XMMBase)
Inst << "XMM" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else
Inst << "Reg" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
}
LogMan::Msg::D("%s", Inst.str().c_str());
}
#endif
uint32_t ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel)
for (auto [BlockNode, BlockHeader] : IR->GetBlocks()) {
using namespace FEXCore::IR;
{
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
uint32_t Node = IR->GetID(BlockNode);
auto IsTarget = JumpTargets.find(Node);
if (IsTarget == JumpTargets.end()) {
IsTarget = JumpTargets.try_emplace(Node).first;
}
// if there is a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
L(IsTarget->second);
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
#ifdef DEBUG_RA
if (IROp->Op != IR::OP_BEGINBLOCK &&
IROp->Op != IR::OP_CONDJUMP &&
IROp->Op != IR::OP_JUMP) {
std::stringstream Inst;
auto Name = FEXCore::IR::GetName(IROp->Op);
if (IROp->HasDest) {
uint64_t PhysReg = RAPass->GetNodeRegister(Node);
if (PhysReg >= GPRPairBase)
Inst << "\tPair" << GetPhys(Node) << " = " << Name << " ";
else if (PhysReg >= XMMBase)
Inst << "\tXMM" << GetPhys(Node) << " = " << Name << " ";
else
Inst << "\tReg" << GetPhys(Node) << " = " << Name << " ";
}
else {
Inst << "\t" << Name << " ";
}
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t ArgNode = IROp->Args[i].ID();
uint64_t PhysReg = RAPass->GetNodeRegister(ArgNode);
if (PhysReg >= GPRPairBase)
Inst << "Pair" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else if (PhysReg >= XMMBase)
Inst << "XMM" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
else
Inst << "Reg" << GetPhys(ArgNode) << (i + 1 == NumArgs ? "" : ", ");
}
LogMan::Msg::D("%s", Inst.str().c_str());
}
#endif
uint32_t ID = IR->GetID(CodeNode);
// Execute handler
OpHandler Handler = OpHandlers[IROp->Op];
(this->*Handler)(IROp, ID);
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel)
{
jmp(*PendingTargetLabel, T_NEAR);
}
PendingTargetLabel = nullptr;
void *Exit = getCurr<void*>();
this->IR = nullptr;
@@ -804,7 +1039,7 @@ static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::InternalT
uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadState *Thread, uint64_t *record) {
auto GuestRip = record[1];
auto HostCode = Thread->BlockCache->FindBlock(GuestRip);
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
if (!HostCode) {
Thread->State.State.rip = GuestRip;
@@ -812,7 +1047,7 @@ uint64_t JITCore::ExitFunctionLink(JITCore *core, FEXCore::Core::InternalThreadS
}
auto LinkerAddress = core->ExitFunctionLinkerAddress;
Thread->BlockCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
@@ -884,21 +1119,21 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
mov(rdx, qword [STATE + offsetof(FEXCore::Core::CPUState, rip)]);
// L1 Cache
mov(r13, Thread->BlockCache->GetL1Pointer());
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rax, rdx);
and_(rax, BlockCache::L1_ENTRIES_MASK);
and_(rax, LookupCache::L1_ENTRIES_MASK);
shl(rax, 4);
cmp(qword[r13 + rax + 8], rdx);
jne(FullLookup);
jmp(qword[r13 + rax + 0]);
L(FullLookup);
mov(r13, Thread->BlockCache->GetPagePointer());
L(FullLookup);
mov(r13, Thread->LookupCache->GetPagePointer());
// Full lookup
mov(rax, rdx);
mov(rbx, Thread->BlockCache->GetVirtualMemorySize() - 1);
mov(rbx, Thread->LookupCache->GetVirtualMemorySize() - 1);
and_(rax, rbx);
shr(rax, 12);
@@ -909,9 +1144,9 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
je(NoBlock);
mov (rax, rdx);
and(rax, 0x0FFF);
and_(rax, 0x0FFF);
shl(rax, (int)log2(sizeof(FEXCore::BlockCache::BlockCacheEntry)));
shl(rax, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// check for aliasing
mov(rcx, qword [rdi + rax + 8]);
@@ -925,13 +1160,13 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
je(NoBlock);
// Update L1
mov(r13, Thread->BlockCache->GetL1Pointer());
mov(r13, Thread->LookupCache->GetL1Pointer());
mov(rcx, rdx);
and_(rcx, BlockCache::L1_ENTRIES_MASK);
and_(rcx, LookupCache::L1_ENTRIES_MASK);
shl(rcx, 1);
mov(qword[r13 + rcx*8 + 8], rdx);
mov(qword[r13 + rcx*8 + 0], rax);
// Real block if we made it here
jmp(rax);
}
@@ -962,7 +1197,7 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
call(rax);
jmp(rax);
}
Label FallbackCore;
// Block creation
{
@@ -996,18 +1231,6 @@ void JITCore::CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread) {
ud2();
}
{
// Interpreter fallback helper code
ThreadSharedData.InterpreterFallbackHelperAddress = getCurr<void*>();
// This will get called so our stack is now misaligned
mov(rdi, STATE);
mov(rax, reinterpret_cast<uint64_t>(ThreadState->IntBackend->CompileCode(nullptr, nullptr)));
call(rax);
jmp(LoopTop);
}
{
// Signal return handler
ThreadSharedData.SignalHandlerReturnAddress = getCurr<uint64_t>();
@@ -1,6 +1,6 @@
#pragma once
#include "Interface/Core/BlockCache.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/BlockSamplingData.h"
#include "Interface/Core/JIT/x86_64/JIT.h"
@@ -14,6 +14,7 @@ using namespace Xbyak;
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <tuple>
@@ -50,15 +51,15 @@ namespace FEXCore::CPU {
using namespace Xbyak::util;
const std::array<Xbyak::Reg, 9> RA64 = { rsi, r8, r9, r10, r11, rbp, r12, r13, r15 };
const std::array<std::pair<Xbyak::Reg, Xbyak::Reg>, 4> RA64Pair = {{ {rsi, r8}, {r9, r10}, {r11, rbp}, {r12, r13} }};
const std::array<Xbyak::Reg, 11> RAXMM = { xmm0, xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10 };
const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm0, xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10 };
const std::array<Xbyak::Reg, 11> RAXMM = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
const std::array<Xbyak::Xmm, 11> RAXMM_x = { xmm1, xmm2, xmm3, xmm4, xmm5, xmm6, xmm7, xmm8, xmm9, xmm10, xmm11};
class JITCore final : public CPUBackend, public Xbyak::CodeGenerator {
public:
explicit JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread, CodeBuffer Buffer, bool CompileThread);
~JITCore() override;
std::string GetName() override { return "JIT"; }
void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) override;
void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
@@ -78,7 +79,7 @@ private:
Label* PendingTargetLabel{};
FEXCore::Context::Context *CTX;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::IR::IRListView<true> const *IR;
FEXCore::IR::IRListView const *IR;
std::unordered_map<IR::OrderedNodeWrapper::NodeOffsetType, Label> JumpTargets;
Xbyak::util::Cpu Features{};
@@ -106,7 +107,7 @@ private:
constexpr static uint8_t RA_64 = 3;
constexpr static uint8_t RA_XMM = 4;
uint32_t GetPhys(uint32_t Node);
IR::PhysicalRegister GetPhys(uint32_t Node);
bool IsFPR(uint32_t Node);
bool IsGPR(uint32_t Node);
@@ -125,9 +126,11 @@ private:
Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr);
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value);
void CreateCustomDispatch(FEXCore::Core::InternalThreadState *Thread);
IR::RegisterAllocationPass *RAPass;
FEXCore::IR::RegisterAllocationData *RAData;
#ifdef BLOCKSTATS
bool GetSamplingData {true};
@@ -165,8 +168,6 @@ private:
uint32_t SignalHandlerRefCounter{};
struct CompilerSharedData {
void *InterpreterFallbackHelperAddress;
uint64_t SignalHandlerReturnAddress{};
uint32_t *SignalHandlerRefCounterPtr{};
@@ -196,6 +197,9 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
void PushRegs();
void PopRegs();
#define DEF_OP(x) void Op_##x(FEXCore::IR::IROp_Header *IROp, uint32_t Node)
///< Unhandled handler
@@ -206,8 +210,10 @@ private:
///< ALU Ops
DEF_OP(TruncElementPair);
DEF_OP(Constant);
DEF_OP(Constant);
DEF_OP(EntrypointOffset);
DEF_OP(InlineConstant);
DEF_OP(InlineEntrypointOffset);
DEF_OP(CycleCounter);
DEF_OP(Add);
DEF_OP(Sub);
@@ -48,7 +48,7 @@ DEF_OP(Break) {
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
// This jump target needs to be a constant offset here
mov(TMP1, ThreadPauseHandlerAddress);
jmp(TMP1);
@@ -89,10 +89,10 @@ DEF_OP(SetRoundingMode) {
mov(TMP1.cvt32(), dword [rsp]);
// Insert the new rounding mode
and(TMP1.cvt32(), ~(0b111 << 13));
and_(TMP1.cvt32(), ~(0b111 << 13));
mov(TMP2.cvt32(), Src);
shl(TMP2.cvt32(), 13);
or(TMP1.cvt32(), TMP2.cvt32());
or_(TMP1.cvt32(), TMP2.cvt32());
// Store it to mxcsr
// Only loads from memory
@@ -28,18 +28,21 @@ DEF_OP(CreateElementPair) {
std::pair<Xbyak::Reg, Xbyak::Reg> Dst;
Xbyak::Reg RegFirst;
Xbyak::Reg RegSecond;
Xbyak::Reg RegTmp;
switch (Op->Header.Size) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetSrc<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_32>(Op->Header.Args[1].ID());
RegTmp = eax;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetSrc<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_64>(Op->Header.Args[1].ID());
RegTmp = rax;
break;
}
default: LogMan::Msg::A("Unknown Size"); break;
@@ -52,7 +55,9 @@ DEF_OP(CreateElementPair) {
mov(Dst.second, RegSecond);
mov(Dst.first, RegFirst);
} else {
LogMan::Msg::A("Unhandled CreateElementPair");
mov(RegTmp, RegFirst);
mov(Dst.second, RegSecond);
mov(Dst.first, RegTmp);
}
}
@@ -262,17 +262,17 @@ DEF_OP(VAddP) {
vpaddw(GetDst(Node), xmm15, xmm14);
switch (Op->Header.ElementSize) {
case 1:
vpunpcklbw(xmm11, xmm15, xmm14);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm11, xmm12);
vpunpckhbw(xmm14, xmm11, xmm12);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpunpcklbw(xmm11, xmm15, xmm14);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm11, xmm12);
vpunpckhbw(xmm14, xmm11, xmm12);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpaddb(GetDst(Node), xmm15, xmm14);
break;
@@ -291,17 +291,17 @@ DEF_OP(VAddP) {
movdqu(xmm15, GetSrc(Op->Header.Args[0].ID()));
movdqu(xmm14, GetSrc(Op->Header.Args[1].ID()));
vpunpcklbw(xmm11, xmm15, xmm14);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm11, xmm12);
vpunpckhbw(xmm14, xmm11, xmm12);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpunpcklbw(xmm11, xmm15, xmm14);
vpunpcklbw(xmm0, xmm15, xmm14);
vpunpckhbw(xmm12, xmm15, xmm14);
vpunpcklbw(xmm15, xmm11, xmm12);
vpunpckhbw(xmm14, xmm11, xmm12);
vpunpcklbw(xmm15, xmm0, xmm12);
vpunpckhbw(xmm14, xmm0, xmm12);
vpaddb(GetDst(Node), xmm15, xmm14);
break;
@@ -327,7 +327,7 @@ DEF_OP(VAddV) {
switch (Op->Header.ElementSize) {
case 2: {
for (int i = Elements; i > 1; i >>= 1) {
phaddw(Dest, Src);
vphaddw(Dest, Src, Dest);
Src = Dest;
}
pextrw(eax, Dest, 0);
@@ -336,7 +336,7 @@ DEF_OP(VAddV) {
}
case 4: {
for (int i = Elements; i > 1; i >>= 1) {
phaddd(Dest, Src);
vphaddd(Dest, Src, Dest);
Src = Dest;
}
pextrd(eax, Dest, 0);
@@ -912,7 +912,10 @@ DEF_OP(VZip2) {
}
DEF_OP(VBSL) {
LogMan::Msg::A("Unimplemented");
auto Op = IROp->C<IR::IROp_VBSL>();
vpand(xmm0, GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
vpandn(xmm12, GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[2].ID()));
vpor(GetDst(Node), xmm0, xmm12);
}
DEF_OP(VCMPEQ) {
@@ -1,10 +1,10 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/Core.h"
#include "Interface/Core/BlockCache.h"
#include "Interface/Core/LookupCache.h"
#include <sys/mman.h>
namespace FEXCore {
BlockCache::BlockCache(FEXCore::Context::Context *CTX)
LookupCache::LookupCache(FEXCore::Context::Context *CTX)
: ctx {CTX} {
// Block cache ends up looking like this
@@ -36,13 +36,13 @@ BlockCache::BlockCache(FEXCore::Context::Context *CTX)
VirtualMemSize = ctx->Config.VirtualMemSize;
}
BlockCache::~BlockCache() {
LookupCache::~LookupCache() {
munmap(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8);
munmap(reinterpret_cast<void*>(PageMemory), CODE_SIZE);
munmap(reinterpret_cast<void*>(L1Pointer), L1_SIZE);
}
void BlockCache::HintUsedRange(uint64_t Address, uint64_t Size) {
void LookupCache::HintUsedRange(uint64_t Address, uint64_t Size) {
// Tell the kernel we will definitely need [Address, Address+Size) mapped for the page pointer
// Page Pointer is allocated per page, so shift by page size
Address >>= 12;
@@ -50,14 +50,22 @@ void BlockCache::HintUsedRange(uint64_t Address, uint64_t Size) {
madvise(reinterpret_cast<void*>(PagePointer + Address), Size, MADV_WILLNEED);
}
void BlockCache::ClearCache() {
void LookupCache::ClearL2Cache() {
// Clear out the page memory
madvise(reinterpret_cast<void*>(PagePointer), ctx->Config.VirtualMemSize / 4096 * 8, MADV_DONTNEED);
madvise(reinterpret_cast<void*>(PageMemory), CODE_SIZE, MADV_DONTNEED);
madvise(reinterpret_cast<void*>(L1Pointer), L1_SIZE, MADV_DONTNEED);
AllocateOffset = 0;
}
void LookupCache::ClearCache() {
// Clear L1
madvise(reinterpret_cast<void*>(L1Pointer), L1_SIZE, MADV_DONTNEED);
// Clear L2
ClearL2Cache();
// All code is gone, remove links
BlockLinks.clear();
// All code is gone, clear the block list
BlockList.clear();
}
}
@@ -2,23 +2,50 @@
#include "Interface/Context/Context.h"
#include <FEXCore/Utils/LogManager.h>
#include <map>
namespace FEXCore {
class BlockCache {
class LookupCache {
public:
struct BlockCacheEntry {
struct LookupCacheEntry {
uintptr_t HostCode;
uintptr_t GuestCode;
};
BlockCache(FEXCore::Context::Context *CTX);
~BlockCache();
LookupCache(FEXCore::Context::Context *CTX);
~LookupCache();
using BlockCacheIter = uintptr_t;
using LookupCacheIter = uintptr_t;
uintptr_t End() { return 0; }
uintptr_t FindBlock(uint64_t Address) {
return FindCodePointerForAddress(Address);
auto HostCode = FindCodePointerForAddress(Address);
if (HostCode) {
return HostCode;
} else {
auto HostCode = BlockList.find(Address);
if (HostCode != BlockList.end()) {
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
} else {
return 0;
}
}
}
std::map<uint64_t, std::vector<uint64_t>> CodePages;
void AddBlockMapping(uint64_t Address, void *HostCode, uint64_t Start, uint64_t Length) {
auto InsertPoint = BlockList.emplace(Address, (uintptr_t)HostCode);
LogMan::Throw::A(InsertPoint.second == true, "Dupplicate block mapping added");
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length) >> 12; CurrentPage <= EndPage; CurrentPage++) {
CodePages[CurrentPage].push_back(Address);
}
// no need to update L1 or L2, they will get updated on first lookup
}
void Erase(uint64_t Address) {
@@ -30,8 +57,11 @@ public:
it->second();
}
// Remove from BlockList
BlockList.erase(Address);
// Do L1
auto &L1Entry = reinterpret_cast<BlockCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = L1Entry.HostCode = 0;
}
@@ -49,14 +79,32 @@ public:
}
// Page exists, just set the offset to zero
auto BlockPointers = reinterpret_cast<BlockCacheEntry*>(LocalPagePointer);
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
BlockPointers[PageOffset].GuestCode = 0;
BlockPointers[PageOffset].HostCode = 0;
}
uintptr_t AddBlockMapping(uint64_t Address, void *Ptr) {
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
BlockLinks.insert({{GuestDestination, HostLink}, delinker});
}
void ClearCache();
void ClearL2Cache();
void HintUsedRange(uint64_t Address, uint64_t Size);
uintptr_t GetL1Pointer() { return L1Pointer; }
uintptr_t GetPagePointer() { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
// Do L1
auto &L1Entry = reinterpret_cast<BlockCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = L1Entry.HostCode = 0;
}
@@ -74,40 +122,23 @@ public:
// Allocate one now if we can
uintptr_t NewPageBacking = AllocateBackingForPage();
if (!NewPageBacking) {
// Couldn't allocate, return so the frontend can recover from this
return 0;
// Couldn't allocate, clear L2 and retry
ClearL2Cache();
CacheBlockMapping(Address, HostCode);
return;
}
Pointers[Address] = NewPageBacking;
LocalPagePointer = NewPageBacking;
}
// Add the new pointer to the page block
auto BlockPointers = reinterpret_cast<BlockCacheEntry*>(LocalPagePointer);
uintptr_t CastPtr = reinterpret_cast<uintptr_t>(Ptr);
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
// This silently replaces existing mappings
BlockPointers[PageOffset].GuestCode = FullAddress;
BlockPointers[PageOffset].HostCode = CastPtr;
return CastPtr;
BlockPointers[PageOffset].HostCode = HostCode;
}
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
BlockLinks.insert({{GuestDestination, HostLink}, delinker});
}
void ClearCache();
void HintUsedRange(uint64_t Address, uint64_t Size);
uintptr_t GetL1Pointer() { return L1Pointer; }
uintptr_t GetPagePointer() { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
private:
uintptr_t AllocateBackingForPage() {
uintptr_t NewBase = AllocateOffset;
uintptr_t NewEnd = AllocateOffset + SIZE_PER_PAGE;
@@ -125,7 +156,7 @@ private:
uintptr_t FindCodePointerForAddress(uint64_t Address) {
// Do L1
auto &L1Entry = reinterpret_cast<BlockCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
@@ -143,7 +174,7 @@ private:
}
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<BlockCacheEntry*>(LocalPagePointer);
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == FullAddress)
{
@@ -172,11 +203,13 @@ private:
}
};
std::map<BlockLinkTag, std::function<void()>> BlockLinks;
std::map<uint64_t, uint64_t> BlockList;
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(BlockCacheEntry);
constexpr static size_t L1_SIZE = L1_ENTRIES * sizeof(BlockCacheEntry);
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(LookupCacheEntry);
constexpr static size_t L1_SIZE = L1_ENTRIES * sizeof(LookupCacheEntry);
size_t AllocateOffset {};
File diff suppressed because it is too large. Load diff
+7 -1
View File
@@ -88,7 +88,8 @@ public:
uint8_t GPRSize = CTX->Config.Is64BitMode ? 8 : 4;
// If we don't have a jump target to a new block then we have to leave
// Set the RIP to the next instruction and leave
_ExitFunction(_Constant(GPRSize * 8, NextRIP));
auto RelocatedNextRIP = _EntrypointOffset(NextRIP - Current_Header->Entry, GPRSize);
_ExitFunction(RelocatedNextRIP);
}
else if (it != JumpTargets.end()) {
_Jump(it->second.BlockEntry);
@@ -349,6 +350,8 @@ public:
void FST(OpcodeArgs);
void FST(OpcodeArgs);
template<bool Truncate>
void FIST(OpcodeArgs);
enum class OpResult {
@@ -381,6 +384,7 @@ public:
void X87TAN(OpcodeArgs);
void X87ATAN(OpcodeArgs);
void X87LDENV(OpcodeArgs);
void X87FLDCW(OpcodeArgs);
void X87FNSTENV(OpcodeArgs);
void X87FSTCW(OpcodeArgs);
void X87LDSW(OpcodeArgs);
@@ -476,9 +480,11 @@ public:
private:
bool DecodeFailure{false};
FEXCore::IR::IROp_IRHeader *Current_Header{};
OrderedNode *Current_HeaderNode{};
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
OrderedNode *GetDynamicPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false);
void StoreResult_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, FEXCore::X86Tables::DecodedOperand const& Operand, OrderedNode *const Src, uint8_t OpSize, int8_t Align);
+41 -2
View File
@@ -3,11 +3,14 @@
#include <cstring>
#include <stdlib.h>
#include <vector>
#include <sys/mman.h>
namespace FEXCore {
constexpr size_t CODE_SIZE = 0x1000;
X86GeneratedCode::X86GeneratedCode() {
// Allocate a page for our emulated guest
CodePtr = malloc(0x1000);
CodePtr = AllocateGuestCodeSpace(CODE_SIZE);
SignalReturn = reinterpret_cast<uint64_t>(CodePtr);
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr) + 2;
@@ -21,7 +24,43 @@ X86GeneratedCode::X86GeneratedCode() {
}
X86GeneratedCode::~X86GeneratedCode() {
free(CodePtr);
munmap(CodePtr, CODE_SIZE);
}
void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
FEXCore::Config::Value<bool> Is64BitMode{FEXCore::Config::CONFIG_IS64BIT_MODE, 0};
if (Is64BitMode()) {
// 64bit mode can have its sigret handler anywhere
return mmap(nullptr, Size, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
}
// First 64bit page
constexpr uintptr_t LOCATION_MAX = 0x1'0000'0000;
// 32bit mode
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void *Ptr = mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED &&
reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
munmap(Ptr, Size);
continue;
}
if (Ptr != MAP_FAILED) {
return Ptr;
}
}
// Can't do anything about this
// Here's hoping the application doesn't use signals
return MAP_FAILED;
}
}
+3
View File
@@ -1,4 +1,6 @@
#pragma once
#include <FEXCore/Config/Config.h>
#include <stdint.h>
namespace FEXCore {
@@ -12,5 +14,6 @@ public:
private:
void *CodePtr{};
void* AllocateGuestCodeSpace(size_t Size);
};
}
-18
View File
@@ -100,24 +100,6 @@ void InitializeInfoTables(Context::OperatingMode Mode) {
InitializeEVEXTables();
#ifndef NDEBUG
auto CheckTable = [&UnknownOp](auto& FinalTable) {
for (size_t i = 0; i < FinalTable.size(); ++i) {
auto const &Op = FinalTable.at(i);
if (Op == UnknownOp) {
LogMan::Msg::A("Unknown Op: 0x%lx", i);
}
}
};
// CheckTable(BaseOps);
// CheckTable(SecondBaseOps);
// CheckTable(RepModOps);
// CheckTable(RepNEModOps);
// CheckTable(OpSizeModOps);
// CheckTable(X87Ops);
X86InstDebugInfo::InstallDebugInfo();
LogMan::Msg::D("X86Tables had %ld total insts, and %ld labeled as understood", Total, NumInsts);
#endif
@@ -125,9 +125,12 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
// GROUP 11
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 0), 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT, 1, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 1), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC6), 7), 1, X86InstInfo{"XABORT", TYPE_INST, FLAGS_MODRM, 1, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 0), 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 1), 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 7), 1, X86InstInfo{"XBEGIN", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_SETS_RIP | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
};
const U16U8InfoStruct PrimaryGroupOpTable_64[] = {
@@ -21,8 +21,8 @@ void InitializeSecondaryModRMTables() {
{((1 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 5), 1, X86InstInfo{"XEND", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 6), 1, X86InstInfo{"XTEST", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{((1 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// REG /3
+12 -7
View File
@@ -7,6 +7,7 @@
#include <string>
#include <map>
#include <array>
#include <Interface/Context/Context.h>
#include "Interface/Core/InternalThreadState.h"
#include "FEXCore/Core/X86Enums.h"
@@ -23,13 +24,17 @@ static thread_local FEXCore::Core::InternalThreadState *Thread;
namespace FEXCore {
struct ExportEntry { const char* Name; ThunkedFunction* Fn; };
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
class ThunkHandler_impl final: public ThunkHandler {
std::shared_mutex ThunksMutex;
std::map<std::string, ThunkedFunction*> Thunks = {
{ "fex:loadlib", &LoadLib}
std::map<IR::SHA256Sum, ThunkedFunction*> Thunks = {
{
// sha256(fex:loadlib)
{ 0x27, 0x7e, 0xb7, 0x69, 0x5b, 0xe9, 0xab, 0x12, 0x6e, 0xf7, 0x85, 0x9d, 0x4b, 0xc9, 0xa2, 0x44, 0x46, 0xcf, 0xbd, 0xb5, 0x87, 0x43, 0xef, 0x28, 0xa2, 0x65, 0xba, 0xfc, 0x89, 0x0f, 0x77, 0x80},
&LoadLib
}
};
/*
@@ -82,8 +87,8 @@ namespace FEXCore {
std::unique_lock lk(That->ThunksMutex);
int i;
for (i = 0; Exports[i].Name; i++) {
That->Thunks[Exports[i].Name] = Exports[i].Fn;
for (i = 0; Exports[i].sha256; i++) {
That->Thunks[*reinterpret_cast<IR::SHA256Sum*>(Exports[i].sha256)] = Exports[i].Fn;
}
LogMan::Msg::D("Loaded %d syms", i);
@@ -92,11 +97,11 @@ namespace FEXCore {
public:
ThunkedFunction* LookupThunk(const char *Name) {
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) {
std::shared_lock lk(ThunksMutex);
auto it = Thunks.find(Name);
auto it = Thunks.find(sha256);
if (it != Thunks.end()) {
return it->second;
+2 -1
View File
@@ -1,4 +1,5 @@
#pragma once
#include <FEXCore/IR/IR.h>
namespace FEXCore::Core {
struct InternalThreadState;
@@ -9,7 +10,7 @@ namespace FEXCore {
class ThunkHandler {
public:
virtual ThunkedFunction* LookupThunk(const char *name) = 0;
virtual ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) = 0;
virtual void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) = 0;
virtual ~ThunkHandler() { }
+47 -11
View File
@@ -28,7 +28,9 @@
"static constexpr FEXCore::IR::RegisterClassType FPRFixedClass {3}",
"static constexpr FEXCore::IR::RegisterClassType GPRPairClass {4}",
"static constexpr FEXCore::IR::RegisterClassType ComplexClass {5}",
"static constexpr FEXCore::IR::RegisterClassType InvalidClass {~0U}",
"static constexpr FEXCore::IR::RegisterClassType InvalidClass {7}",
"",
"static constexpr uint8_t InvalidReg {31}",
"",
"static const FEXCore::IR::TypeDefinition i8 {TypeDefinition::Create(1, 0)}",
"static const FEXCore::IR::TypeDefinition i16 {TypeDefinition::Create(2, 0)}",
@@ -78,8 +80,7 @@
],
"Args": [
"uint64_t", "Entry",
"uint32_t", "BlockCount",
"bool", "ShouldInterpret"
"uint32_t", "BlockCount"
]
},
"CodeBlock": {
@@ -129,18 +130,16 @@
"DestClass": "GPR",
"DestSize": "8",
"Args": [
"__uint128_t", "CodeOriginal",
"uint64_t", "CodePtr",
"uint64_t", "CodeOriginalLow",
"uint64_t", "CodeOriginalHigh",
"int64_t", "Offset",
"uint8_t", "CodeLength"
]
},
"RemoveCodeEntry": {
"HasSideEffects": true,
"OpClass": "Misc",
"Args": [
"uint64_t", "RIP"
]
"OpClass": "Misc"
},
"GuestCallDirect": {
@@ -281,6 +280,32 @@
]
},
"EntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"OpClass": "ALU",
"HasDest": true,
"DestClass": "GPR",
"DestSize": "RegisterSize",
"Args": [
"int64_t", "Offset"
],
"HelperArgs": [
"uint8_t", "RegisterSize"
]
},
"InlineEntrypointOffset": {
"Desc": ["Returns the <entrypoint> + Offset address"],
"OpClass": "ALU",
"DestSize": "RegisterSize",
"Args": [
"int64_t", "Offset"
],
"HelperArgs": [
"uint8_t", "RegisterSize"
]
},
"Constant": {
"Desc": ["Generates a 64bit constant inside of a GPR",
"Unsupported to create a constant in FPR"
@@ -641,8 +666,7 @@
"ArgPtr"
],
"Args":[
"const char*", "ThunkName",
"uintptr_t", "ThunkFnPtr"
"SHA256Sum", "ThunkNameHash"
]
},
@@ -3307,6 +3331,15 @@
]
},
"F80LoadFCW": {
"OpClass": "Vector",
"HasSideEffects": true,
"SSAArgs": "1",
"SSANames": [
"X80FCW"
]
},
"F80Add": {
"OpClass": "Vector",
"HasDest": true,
@@ -3436,6 +3469,9 @@
"SSANames": [
"X80Src"
],
"Args": [
"bool", "Truncate"
],
"HelperArgs": [
"uint8_t", "Size"
]
@@ -3,6 +3,8 @@
#include <FEXCore/Utils/LogManager.h>
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <iomanip>
namespace FEXCore::IR {
#define IROP_GETNAME_IMPL
#define IROP_GETRAARGS_IMPL
@@ -12,16 +14,22 @@ namespace FEXCore::IR {
#include <FEXCore/IR/IRDefines.inc>
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, uint64_t Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, const SHA256Sum &Arg) {
*out << "sha256:";
for(auto byte: Arg.data)
*out << std::hex << std::setfill('0') << std::setw(2) << (unsigned int)byte;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, uint64_t Arg) {
*out << "#0x" << std::hex << Arg;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, const char* Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, const char* Arg) {
*out << Arg;
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, CondClassType Arg) {
std::array<std::string, 14> CondNames = {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, CondClassType Arg) {
std::array<std::string, 22> CondNames = {
"EQ",
"NEQ",
"UGE",
@@ -36,12 +44,20 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
"SLT",
"SGT",
"SLE",
"Invalid Cond",
"Invalid Cond",
"FLU",
"FGE",
"FLEU",
"FGT",
"FU",
"FNU"
};
*out << CondNames[Arg];
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, MemOffsetType Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, MemOffsetType Arg) {
std::array<std::string, 3> Names = {
"SXTX",
"UXTW",
@@ -51,7 +67,7 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
*out << Names[Arg];
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, RegisterClassType Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, RegisterClassType Arg) {
if (Arg == GPRClass.Val)
*out << "GPR";
else if (Arg == GPRFixedClass.Val)
@@ -66,26 +82,33 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
*out << "Unknown Registerclass " << Arg;
}
static void PrintArg(std::stringstream *out, IRListView<false> const* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationPass *RAPass) {
static void PrintArg(std::stringstream *out, IRListView const* IR, OrderedNodeWrapper Arg, IR::RegisterAllocationData *RAData) {
auto [CodeNode, IROp] = IR->at(Arg)();
*out << "%ssa" << std::to_string(Arg.ID());
if (RAPass) {
uint64_t RegClass = RAPass->GetNodeRegister(Arg.ID());
FEXCore::IR::RegisterClassType Class {uint32_t(RegClass >> 32)};
uint32_t Reg = RegClass;
switch (Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(GPRFixed"; break;
case FEXCore::IR::FPRClass.Val: *out << "(FPR"; break;
case FEXCore::IR::FPRFixedClass.Val: *out << "(FPRFixed"; break;
case FEXCore::IR::GPRPairClass.Val: *out << "(GPRPair"; break;
case FEXCore::IR::ComplexClass.Val: *out << "(Complex"; break;
case FEXCore::IR::InvalidClass.Val: *out << "(Invalid"; break;
default: *out << "(Unknown"; break;
}
if (Arg.ID() == 0) {
*out << "%Invalid";
} else {
*out << "%ssa" << std::to_string(Arg.ID());
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(Arg.ID());
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(GPRFixed"; break;
case FEXCore::IR::FPRClass.Val: *out << "(FPR"; break;
case FEXCore::IR::FPRFixedClass.Val: *out << "(FPRFixed"; break;
case FEXCore::IR::GPRPairClass.Val: *out << "(GPRPair"; break;
case FEXCore::IR::ComplexClass.Val: *out << "(Complex"; break;
case FEXCore::IR::InvalidClass.Val: *out << "(Invalid"; break;
default: *out << "(Unknown"; break;
}
*out << std::dec << Reg << ")";
if (PhyReg.Class != FEXCore::IR::InvalidClass.Val) {
*out << std::dec << (uint32_t)PhyReg.Reg << ")";
} else {
*out << ")";
}
}
}
if (IROp->HasDest) {
@@ -108,15 +131,7 @@ static void PrintArg(std::stringstream *out, IRListView<false> const* IR, Ordere
}
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, IR::TypeDefinition Arg) {
*out << "i" << std::dec << static_cast<uint32_t>(Arg.Bytes() * 8);
if (Arg.Elements()) {
*out << "v" << std::dec << static_cast<uint32_t>(Arg.Elements());
}
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false> const* IR, FEXCore::IR::FenceType Arg) {
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, FEXCore::IR::FenceType Arg) {
if (Arg == IR::Fence_Load) {
*out << "Loads";
}
@@ -131,7 +146,7 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView<false>
}
}
void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAllocationPass *RAPass) {
void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationData *RAData) {
auto HeaderOp = IR->GetHeader();
int8_t CurrentIndent = 0;
@@ -188,11 +203,9 @@ void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAlloc
*out << "%ssa" << std::to_string(ID);
if (RAPass) {
uint64_t RegClass = RAPass->GetNodeRegister(ID);
FEXCore::IR::RegisterClassType Class {uint32_t(RegClass >> 32)};
uint32_t Reg = RegClass;
switch (Class) {
if (RAData) {
auto PhyReg = RAData->GetNodeRegister(ID);
switch (PhyReg.Class) {
case FEXCore::IR::GPRClass.Val: *out << "(GPR"; break;
case FEXCore::IR::GPRFixedClass.Val: *out << "(GPRFixed"; break;
case FEXCore::IR::FPRClass.Val: *out << "(FPR"; break;
@@ -202,8 +215,11 @@ void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAlloc
case FEXCore::IR::InvalidClass.Val: *out << "(Invalid"; break;
default: *out << "(Unknown"; break;
}
*out << std::dec << Reg << ")";
if (PhyReg.Class != FEXCore::IR::InvalidClass.Val) {
*out << std::dec << (uint32_t)PhyReg.Reg << ")";
} else {
*out << ")";
}
}
*out << " i" << std::dec << (ElementSize * 8);
@@ -245,9 +261,9 @@ void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAlloc
auto [NodeNode, IROp] = NodeBegin();
auto PhiOp = IROp->C<IR::IROp_PhiValue>();
*out << "[ ";
PrintArg(out, IR, PhiOp->Value, RAPass);
PrintArg(out, IR, PhiOp->Value, RAData);
*out << ", ";
PrintArg(out, IR, PhiOp->Block, RAPass);
PrintArg(out, IR, PhiOp->Block, RAData);
*out << " ]";
if (PhiOp->Next.ID())
+653
View File
@@ -0,0 +1,653 @@
#include <string>
#include <vector>
#include <istream>
#include <unordered_map>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
namespace FEXCore::IR {
namespace {
enum class DecodeFailure {
DECODE_OKAY,
DECODE_UNKNOWN_TYPE,
DECODE_INVALID,
DECODE_INVALIDCHAR,
DECODE_INVALIDRANGE,
DECODE_INVALIDREGISTERCLASS,
DECODE_UNKNOWN_SSA,
DECODE_INVALID_CONDFLAG,
DECODE_INVALID_MEMOFFSETTYPE,
DECODE_INVALID_FENCETYPE,
};
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string DecodeErrorToString(DecodeFailure Failure) {
switch (Failure) {
case DecodeFailure::DECODE_OKAY: return "Okay";
case DecodeFailure::DECODE_UNKNOWN_TYPE: return "Unknown Type";
case DecodeFailure::DECODE_INVALID: return "Invalid";
case DecodeFailure::DECODE_INVALIDCHAR: return "Invalid starting char";
case DecodeFailure::DECODE_INVALIDRANGE: return "Invalid integer range";
case DecodeFailure::DECODE_INVALIDREGISTERCLASS: return "Invalid register class";
case DecodeFailure::DECODE_UNKNOWN_SSA: return "Unknown SSA value";
case DecodeFailure::DECODE_INVALID_CONDFLAG: return "Invalid Conditional name";
case DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE: return "Invalid Memory Offset Type";
case DecodeFailure::DECODE_INVALID_FENCETYPE: return "Invalid Fence Type";
};
}
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
class IRParser: public FEXCore::IR::IREmitter {
public:
template<typename Type>
std::pair<DecodeFailure, Type> DecodeValue(std::string &Arg) {
return {DecodeFailure::DECODE_UNKNOWN_TYPE, {}};
}
template<>
std::pair<DecodeFailure, uint8_t> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, bool> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE || Result > 1) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result != 0};
}
template<>
std::pair<DecodeFailure, uint16_t> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint16_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, uint32_t> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint32_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, uint64_t> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint64_t Result = strtoull(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, int64_t> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
int64_t Result = (int64_t)strtoull(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, IR::SHA256Sum> DecodeValue(std::string &Arg) {
IR::SHA256Sum Result;
if (Arg.at(0) != 's' || Arg.at(1) != 'h' || Arg.at(2) != 'a' || Arg.at(3) != '2' || Arg.at(4) != '5' || Arg.at(5) != '6' || Arg.at(6) != ':')
return {DecodeFailure::DECODE_INVALIDCHAR, Result};
auto GetDigit = [](const std::string &Arg, int pos, uint8_t *val) {
auto chr = Arg.at(pos);
if (chr >= '0' && chr <= '9') {
*val = chr - '0';
return true;
} else if (chr >= 'a' && chr <= 'f') {
*val = 10 + chr - 'a';
return true;
} else {
return false;
}
};
for (size_t i = 0; i < sizeof(Result.data); i++) {
uint8_t high, low;
if (!GetDigit(Arg, 7 + 2 * i + 0, &high) || !GetDigit(Arg, 7 + 2 * i + 1, &low)) {
return {DecodeFailure::DECODE_INVALIDRANGE, Result};
}
Result.data[i] = high * 16 + low;
}
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType> DecodeValue(std::string &Arg) {
if (Arg == "GPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRClass};
}
else if (Arg == "FPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRClass};
}
else if (Arg == "GPRPair") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRPairClass};
}
else if (Arg == "Complex") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::ComplexClass};
}
return {DecodeFailure::DECODE_INVALIDREGISTERCLASS, FEXCore::IR::InvalidClass};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition> DecodeValue(std::string &Arg) {
uint8_t Size{}, Elements{1};
int NumArgs = sscanf(Arg.c_str(), "i%hhdv%hhd", &Size, &Elements);
if (NumArgs != 1 && NumArgs != 2) {
return {DecodeFailure::DECODE_INVALID, {}};
}
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::TypeDefinition::Create(Size / 8, Elements)};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::CondClassType> DecodeValue(std::string &Arg) {
std::array<std::string, 22> CondNames = {
"EQ",
"NEQ",
"UGE",
"ULT",
"MI",
"PL",
"VS",
"VC",
"UGT",
"ULE",
"SGE",
"SLT",
"SGT",
"SLE",
"Invalid Cond",
"Invalid Cond",
"FLU",
"FGE",
"FLEU",
"FGT",
"FU",
"FNU"
};
for (size_t i = 0; i < CondNames.size(); ++i) {
if (CondNames[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, CondClassType{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_CONDFLAG, {}};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType> DecodeValue(std::string &Arg) {
std::array<std::string, 3> Names = {
"SXTX",
"UXTW",
"SXTW",
};
for (size_t i = 0; i < Names.size(); ++i) {
if (Names[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, MemOffsetType{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE, {}};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::FenceType> DecodeValue(std::string &Arg) {
std::array<std::string, 3> Names = {
"Loads",
"Stores",
"LoadStores",
};
for (size_t i = 0; i < Names.size(); ++i) {
if (Names[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, FenceType{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_FENCETYPE, {}};
}
template<>
std::pair<DecodeFailure, OrderedNode*> DecodeValue(std::string &Arg) {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
// Strip off the type qualifier from the ssa value
size_t ArgEnd = std::string::npos;
std::string SSAName = trim(Arg);
ArgEnd = SSAName.find_first_of(" ");
if (ArgEnd != std::string::npos) {
SSAName = SSAName.substr(0, ArgEnd);
}
// Forward declarations may make this not succed
auto Op = SSANameMapper.find(SSAName);
if (Op == SSANameMapper.end()) {
return {DecodeFailure::DECODE_UNKNOWN_SSA, nullptr};
}
return {DecodeFailure::DECODE_OKAY, Op->second};
}
struct LineDefinition {
size_t LineNumber;
bool HasDefinition{};
std::string Definition{};
FEXCore::IR::TypeDefinition Size{};
std::string IROp{};
FEXCore::IR::IROps OpEnum;
bool HasArgs{};
std::vector<std::string> Args;
OrderedNode *Node{};
};
std::vector<std::string> Lines;
std::unordered_map<std::string, OrderedNode*> SSANameMapper;
std::vector<LineDefinition> Defs;
LineDefinition *CurrentDef{};
IRParser(std::istream *text) {
InitializeStaticTables();
std::string TmpLine;
while (!text->eof()) {
std::getline(*text, TmpLine);
if (text->eof()) {
break;
}
if (text->fail()) {
LogMan::Msg::E("Failed to getline on line: %ld", Lines.size());
return;
}
Lines.emplace_back(TmpLine);
}
ResetWorkingList();
Loaded = Parse();
}
bool Loaded = false;
#define IROP_PARSER_ALLOCATE_HELPERS
#include <FEXCore/IR/IRDefines.inc>
bool Parse() {
auto CheckPrintError = [&](LineDefinition &Def, DecodeFailure Failure) -> bool {
if (Failure != DecodeFailure::DECODE_OKAY) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Value Couldn't be decoded due to %s", DecodeErrorToString(Failure).c_str());
return false;
}
return true;
};
// String parse every line for our definitions
for (size_t i = 0; i < Lines.size(); ++i) {
std::string Line = Lines[i];
LineDefinition Def{};
CurrentDef = &Def;
Def.LineNumber = i;
Line = trim(Line);
// Skip empty lines
if (Line.empty()) {
continue;
}
if (Line[0] == ';') {
// This is a comment line
// Skip it
continue;
}
size_t CurrentPos{};
// Let's see if this node is assigning something first
if (Line[0] == '%') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of("=", CurrentPos)) != std::string::npos) {
Def.Definition = Line.substr(0, DefinitionEnd);
Def.Definition = trim(Def.Definition);
Def.HasDefinition = true;
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA declaration without assignment");
return false;
}
}
// Check if we are pulling in some IR from the IR Printer
// Prints (%ssa%d) at the start of lines without a definition
if (Line[0] == '(') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of(")", CurrentPos)) != std::string::npos) {
size_t SSAEnd = std::string::npos;
if ((SSAEnd = Line.find_last_of(" ", DefinitionEnd)) != std::string::npos) {
std::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
Type = trim(Type);
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
Def.Size = DefinitionSize.second;
}
Def.Definition = trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
CurrentPos = DefinitionEnd + 1;
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA value with numbered SSA provided but no closing parentheses");
return false;
}
}
if (Def.HasDefinition) {
// Let's check if we have a size declared with this variable
size_t NameEnd = std::string::npos;
if ((NameEnd = Def.Definition.find_first_of(" ")) != std::string::npos) {
std::string Type = Def.Definition.substr(NameEnd + 1);
Type = trim(Type);
Def.Definition = trim(Def.Definition.substr(0, NameEnd));
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
Def.Size = DefinitionSize.second;
}
if (Def.Definition == "%Invalid") {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Definition tried to define reserved %Invalid ssa node");
return false;
}
}
// Let's get the IR op
size_t OpNameEnd = std::string::npos;
std::string RemainingLine = trim(Line.substr(CurrentPos));
CurrentPos = 0;
if ((OpNameEnd = RemainingLine.find_first_of(" \t\n\r\0", CurrentPos)) != std::string::npos) {
Def.IROp = RemainingLine.substr(CurrentPos, OpNameEnd);
Def.IROp = trim(Def.IROp);
Def.HasArgs = true;
CurrentPos = OpNameEnd;
}
else {
if (RemainingLine.empty()) {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Line without an IROp?");
return false;
}
Def.IROp = RemainingLine;
Def.HasArgs = false;
}
if (Def.HasArgs) {
RemainingLine = trim(RemainingLine.substr(CurrentPos));
CurrentPos = 0;
if (RemainingLine.empty()) {
// How did we get here?
Def.HasArgs = false;
}
else {
while (!RemainingLine.empty()) {
size_t ArgEnd = std::string::npos;
ArgEnd = RemainingLine.find_first_of(",");
std::string Arg = RemainingLine.substr(0, ArgEnd);
Arg = trim(Arg);
Def.Args.emplace_back(Arg);
RemainingLine.erase(0, ArgEnd+1); // +1 to ensure we go past the ','
if (ArgEnd == std::string::npos)
break;
}
}
}
Defs.emplace_back(Def);
}
// Ensure all of the ops are real ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
auto Op = NameToOpMap.find(Def.IROp);
if (Op == NameToOpMap.end()) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("IROp '%s' doesn't exist", Def.IROp.c_str());
return false;
}
Def.OpEnum = Op->second;
}
// Emit the header op
IRPair<IROp_IRHeader> IRHeader;
{
auto &Def = Defs[0];
CurrentDef = &Def;
if (Def.OpEnum != FEXCore::IR::IROps::OP_IRHEADER) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("First op needs to be IRHeader. Was '%s'", Def.IROp.c_str());
return false;
}
auto Entry = DecodeValue<uint64_t>(Def.Args[0]);
auto CodeBlockCount = DecodeValue<uint64_t>(Def.Args[2]);
if (!CheckPrintError(Def, Entry.first)) return false;
if (!CheckPrintError(Def, CodeBlockCount.first)) return false;
IRHeader = _IRHeader(InvalidNode, Entry.second, CodeBlockCount.second);
}
SetWriteCursor(nullptr); // isolate the header from everything following
// Initialize SSANameMapper with Invalid value
SSANameMapper["%Invalid"] = Invalid();
// Spin through the blocks and generate basic block ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
if (Def.OpEnum == FEXCore::IR::IROps::OP_CODEBLOCK) {
auto CodeBlock = _CodeBlock(InvalidNode, InvalidNode);
SSANameMapper[Def.Definition] = CodeBlock.Node;
Def.Node = CodeBlock.Node;
if (i == 1) {
// First code block is the entry block
// Link the header to the first block
IRHeader.first->Blocks = CodeBlock.Node->Wrapped(ListData.Begin());
}
CodeBlocks.emplace_back(CodeBlock.Node);
}
}
SetWriteCursor(nullptr); // isolate the block headers too
// Spin through all the definitions and add the ops to the basic blocks
OrderedNode *CurrentBlock{};
FEXCore::IR::IROp_CodeBlock *CurrentBlockOp{};
for(size_t i = 1; i < Defs.size(); ++i) {
auto &Def = Defs[i];
CurrentDef = &Def;
switch (Def.OpEnum) {
// Special handled
case FEXCore::IR::IROps::OP_IRHEADER:
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("IRHEADER used in the middle of the block!");
return false; // only one OP_IRHEADER allowed per block
case FEXCore::IR::IROps::OP_CODEBLOCK: {
SetWriteCursor(nullptr); // isolate from previous block
if (CurrentBlock != nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("CodeBlock being used inside of already existing codeblock!");
return false;
}
CurrentBlock = Def.Node;
CurrentBlockOp = CurrentBlock->Op(Data.Begin())->CW<FEXCore::IR::IROp_CodeBlock>();
break;
}
case FEXCore::IR::IROps::OP_BEGINBLOCK: {
if (CurrentBlock == nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("EndBlock being used outside of a block!");
return false;
}
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
if (!CheckPrintError(Def, Adjust.first)) return false;
Def.Node = _BeginBlock(Adjust.second);
CurrentBlockOp->Begin = Def.Node->Wrapped(ListData.Begin());
break;
}
case FEXCore::IR::IROps::OP_ENDBLOCK: {
if (CurrentBlock == nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("EndBlock being used outside of a block!");
return false;
}
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
if (!CheckPrintError(Def, Adjust.first)) return false;
Def.Node = _EndBlock(Adjust.second);
CurrentBlockOp->Last = Def.Node->Wrapped(ListData.Begin());
CurrentBlock = nullptr;
CurrentBlockOp = nullptr;
break;
}
case FEXCore::IR::IROps::OP_DUMMY: {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Dummy op must not be used");
break;
}
#define IROP_PARSER_SWITCH_HELPERS
#include <FEXCore/IR/IRDefines.inc>
default: {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Unhandled Op enum '%s' in parser", Def.IROp.c_str());
return false;
break;
}
}
if (Def.HasDefinition) {
auto IROp = Def.Node->Op(Data.Begin());
if (Def.Size.Elements()) {
IROp->Size = Def.Size.Bytes() * Def.Size.Elements();
IROp->ElementSize = Def.Size.Bytes();
}
else {
IROp->Size = Def.Size.Bytes();
IROp->ElementSize = 0;
}
SSANameMapper[Def.Definition] = Def.Node;
}
}
return true;
}
void InitializeStaticTables() {
if (NameToOpMap.size() == 0) {
for (FEXCore::IR::IROps Op = FEXCore::IR::IROps::OP_DUMMY;
Op <= FEXCore::IR::IROps::OP_LAST;
Op = static_cast<FEXCore::IR::IROps>(static_cast<uint32_t>(Op) + 1)) {
NameToOpMap[FEXCore::IR::GetName(Op)] = Op;
}
}
}
};
} // anon namespace
IREmitter* Parse(std::istream *in) {
auto parser = new IRParser(in);
if (parser->Loaded) {
return parser;
} else {
delete parser;
return nullptr;
}
}
}
+2 -4
View File
@@ -1,6 +1,6 @@
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Config/Config.h>
@@ -11,9 +11,7 @@ void PassManager::AddDefaultPasses(bool InlineConstants, bool StaticRegisterAllo
if (!DisablePasses()) {
InsertPass(CreateContextLoadStoreElimination());
InsertPass(CreateDeadFlagStoreElimination());
InsertPass(CreateDeadGPRStoreElimination());
InsertPass(CreateDeadFPRStoreElimination());
InsertPass(CreateDeadStoreElimination());
InsertPass(CreatePassDeadCodeElimination());
InsertPass(CreateConstProp(InlineConstants));
+2 -3
View File
@@ -3,14 +3,13 @@
namespace FEXCore::IR {
class Pass;
class RegisterAllocationPass;
class RegisterAllocationData;
FEXCore::IR::Pass* CreateConstProp(bool InlineConstants);
FEXCore::IR::Pass* CreateContextLoadStoreElimination();
FEXCore::IR::Pass* CreateSyscallOptimization();
FEXCore::IR::Pass* CreateDeadFlagCalculationEliminination();
FEXCore::IR::Pass* CreateDeadFlagStoreElimination();
FEXCore::IR::Pass* CreateDeadGPRStoreElimination();
FEXCore::IR::Pass* CreateDeadFPRStoreElimination();
FEXCore::IR::Pass* CreateDeadStoreElimination();
FEXCore::IR::Pass* CreatePassDeadCodeElimination();
FEXCore::IR::Pass* CreateIRCompaction();
FEXCore::IR::RegisterAllocationPass* CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass, bool OptimizeSRA);
+23 -12
View File
@@ -12,6 +12,8 @@
namespace FEXCore::IR {
class ConstProp final : public FEXCore::IR::Pass {
std::unordered_map<uint64_t, OrderedNode*> ConstPool;
std::map<OrderedNode*, uint64_t> AddressgenConsts;
public:
bool Run(IREmitter *IREmit) override;
bool InlineConstants;
@@ -164,22 +166,21 @@ bool ConstProp::Run(IREmitter *IREmit) {
auto HeaderOp = CurrentIR.GetHeader();
{
std::map<uint64_t, OrderedNode*> Consts;
// constants are pooled per block
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_CONSTANT) {
auto Op = IROp->C<IR::IROp_Constant>();
if (Consts.count(Op->Constant)) {
IREmit->ReplaceAllUsesWith(CodeNode, Consts[Op->Constant]);
if (ConstPool.count(Op->Constant)) {
IREmit->ReplaceAllUsesWith(CodeNode, ConstPool[Op->Constant]);
Changed = true;
} else {
Consts[Op->Constant] = CodeNode;
ConstPool[Op->Constant] = CodeNode;
}
}
}
Consts.clear();
ConstPool.clear();
}
}
@@ -265,14 +266,13 @@ bool ConstProp::Run(IREmitter *IREmit) {
// LoadMem / StoreMem imm pooling
// If imms are close by, use address gen to generate the values instead of using a new imm
std::map<OrderedNode*, uint64_t> Consts;
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_LOADMEM || IROp->Op == OP_STOREMEM) {
uint64_t Addr;
if (IREmit->IsValueConstant(IROp->Args[0], &Addr) && IROp->Args[1].IsInvalid()) {
for (auto& Const: Consts) {
for (auto& Const: AddressgenConsts) {
if ((Addr - Const.second) < 65536) {
IREmit->ReplaceNodeArgument(CodeNode, 0, Const.first);
IREmit->ReplaceNodeArgument(CodeNode, 1, IREmit->_Constant(Addr - Const.second));
@@ -280,14 +280,14 @@ bool ConstProp::Run(IREmitter *IREmit) {
}
}
Consts[IREmit->UnwrapNode(IROp->Args[0])] = Addr;
AddressgenConsts[IREmit->UnwrapNode(IROp->Args[0])] = Addr;
}
doneOp:
;
}
IREmit->SetWriteCursor(CodeNode);
}
Consts.clear();
AddressgenConsts.clear();
}
for (auto [CodeNode, IROp] : CurrentIR.GetAllCode()) {
@@ -485,7 +485,7 @@ bool ConstProp::Run(IREmitter *IREmit) {
auto Op = IROp->CW<IR::IROp_LoadMem>();
auto AddressHeader = IREmit->GetOpHeader(Op->Header.Args[0]);
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8 && !Header->ShouldInterpret) {
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8) {
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, Op->Size, AddressHeader);
@@ -503,7 +503,7 @@ bool ConstProp::Run(IREmitter *IREmit) {
auto Op = IROp->CW<IR::IROp_StoreMem>();
auto AddressHeader = IREmit->GetOpHeader(Op->Header.Args[0]);
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8 && !Header->ShouldInterpret) {
if (AddressHeader->Op == OP_ADD && AddressHeader->Size == 8) {
auto [OffsetType, OffsetScale, Arg0, Arg1] = MemExtendedAddressing(IREmit, Op->Size, AddressHeader);
Op->OffsetType = OffsetType;
@@ -712,7 +712,9 @@ bool ConstProp::Run(IREmitter *IREmit) {
uint64_t amt = __builtin_ctzl(Constant2);
IREmit->SetWriteCursor(CodeNode);
auto shift = IREmit->_Lshl(CurrentIR.GetNode(Op->Header.Args[0]), IREmit->_Constant(amt));
shift.first->Header.Size = IROp->Size; // force Lshl to be the same size as the original Mul
IREmit->ReplaceAllUsesWith(CodeNode, shift);
Changed = true;
}
}
break;
@@ -747,7 +749,7 @@ bool ConstProp::Run(IREmitter *IREmit) {
}
// constant inlining
if (!HeaderOp->ShouldInterpret && InlineConstants) {
if (InlineConstants) {
for (auto [CodeNode, IROp] : CurrentIR.GetAllCode()) {
switch(IROp->Op) {
case OP_LSHR:
@@ -852,6 +854,15 @@ bool ConstProp::Run(IREmitter *IREmit) {
IREmit->ReplaceNodeArgument(CodeNode, 0, IREmit->_InlineConstant(Constant));
Changed = true;
} else {
auto NewRIP = IREmit->GetOpHeader(Op->NewRIP);
if (NewRIP->Op == OP_ENTRYPOINTOFFSET) {
auto EO = NewRIP->C<IR::IROp_EntrypointOffset>();
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->NewRIP));
IREmit->ReplaceNodeArgument(CodeNode, 0, IREmit->_InlineEntrypointOffset(EO->Offset, EO->Header.Size));
Changed = true;
}
}
break;
}
@@ -28,9 +28,12 @@ namespace {
FEXCore::IR::OrderedNode *StoreNode;
};
using ContextInfo = std::vector<ContextMemberInfo>;
struct ContextInfo {
std::vector<ContextMemberInfo*> Lookup;
std::vector<ContextMemberInfo> ClassificationInfo;
};
constexpr static std::array<LastAccessType, 14> DefaultAccess = {
constexpr static std::array<LastAccessType, 15> DefaultAccess = {
ACCESS_NONE,
ACCESS_NONE,
ACCESS_INVALID, // PAD
@@ -45,9 +48,12 @@ namespace {
ACCESS_INVALID, // PAD
ACCESS_NONE,
ACCESS_NONE,
ACCESS_NONE,
};
static void ClassifyContextStruct(ContextInfo *ContextClassification) {
static void ClassifyContextStruct(ContextInfo *ContextClassificationInfo) {
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, rip),
@@ -88,15 +94,6 @@ namespace {
});
}
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gs),
sizeof(FEXCore::Core::CPUState::gs),
},
DefaultAccess[4],
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, es),
@@ -133,6 +130,15 @@ namespace {
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gs),
sizeof(FEXCore::Core::CPUState::gs),
},
DefaultAccess[4],
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo{
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, fs),
@@ -185,17 +191,38 @@ namespace {
});
}
// FCW
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, FCW),
sizeof(FEXCore::Core::CPUState::FCW),
},
DefaultAccess[14],
FEXCore::IR::InvalidClass,
});
size_t ClassifiedStructSize{};
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
for (auto &it : *ContextClassification) {
LogMan::Throw::A(it.Class.Offset == ContextClassificationInfo->Lookup.size(), "Offset missmatch %d %d", it.Class.Offset == ContextClassificationInfo->Lookup.size());
for (int i = 0; i < it.Class.Size; i++) {
ContextClassificationInfo->Lookup.push_back(&it);
}
ClassifiedStructSize += it.Class.Size;
}
LogMan::Throw::A(ClassifiedStructSize == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! %ld != %ld",
ClassifiedStructSize, sizeof(FEXCore::Core::CPUState));
LogMan::Throw::A(ContextClassificationInfo->Lookup.size() == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! %ld != %ld",
ContextClassificationInfo->Lookup.size(), sizeof(FEXCore::Core::CPUState));
}
static void ResetClassificationAccesses(ContextInfo *ContextClassification) {
static void ResetClassificationAccesses(ContextInfo *ContextClassificationInfo) {
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
auto SetAccess = [&](size_t Offset, auto Access) {
ContextClassification->at(Offset).Accessed = Access;
ContextClassification->at(Offset).AccessRegClass = FEXCore::IR::InvalidClass;
@@ -234,6 +261,8 @@ namespace {
for (size_t i = 0; i < 32; ++i) {
SetAccess(Offset++, DefaultAccess[13]);
}
SetAccess(Offset++, DefaultAccess[14]);
}
struct BlockInfo {
@@ -265,20 +294,8 @@ private:
bool RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit);
};
ContextMemberInfo *RCLSE::FindMemberInfo(ContextInfo *ClassifiedInfo, uint32_t Offset, uint8_t Size) {
ContextMemberInfo *Info{};
// Just linearly scan to find the info
for (size_t i = 0; i < ClassifiedInfo->size(); ++i) {
ContextMemberInfo *LocalInfo = &ClassifiedInfo->at(i);
if (LocalInfo->Class.Offset <= Offset &&
(LocalInfo->Class.Offset + LocalInfo->Class.Size) > Offset) {
Info = LocalInfo;
break;
}
}
LogMan::Throw::A(Info != nullptr, "Couldn't find Context Member to record to");
return Info;
ContextMemberInfo *RCLSE::FindMemberInfo(ContextInfo *ContextClassificationInfo, uint32_t Offset, uint8_t Size) {
return ContextClassificationInfo->Lookup.at(Offset);
}
ContextMemberInfo *RCLSE::RecordAccess(ContextMemberInfo *Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size, LastAccessType AccessType, FEXCore::IR::OrderedNode *Node, FEXCore::IR::OrderedNode *StoreNode) {
@@ -404,7 +421,7 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter *IREmit) {
// XXX: Walk the list and calculate the control flow
ContextInfo LocalInfo = ClassifiedStruct;
ContextInfo &LocalInfo = ClassifiedStruct;
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
@@ -1,170 +0,0 @@
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
// Higher values might result in more stores getting eliminated but will make the optimization take more time
constexpr int PropagationRounds = 5;
namespace FEXCore::IR {
class DeadFPRStoreElimination final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
struct FPRInfo {
uint64_t reads { 0 };
uint64_t writes { 0 };
uint64_t kill { 0 };
};
bool IsFPR(uint32_t Offset) {
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto end = offsetof(FEXCore::Core::ThreadState, State.xmm[17][0]);
if (Offset < begin || Offset >= end)
return false;
return true;
}
bool IsTrackedWriteFPR(uint32_t Offset, uint8_t Size) {
if (Size != 16 && Size != 8 && Size != 4)
return false;
if (Offset & 15)
return false;
return IsFPR(Offset);
}
uint64_t FPRBit(uint32_t Offset, uint32_t Size) {
if (!IsFPR(Offset)) {
return 0;
}
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto regn = (Offset - begin)/16;
auto bitn = regn * 3;
if (!IsTrackedWriteFPR(Offset, Size))
return 7UL << (bitn);
if (Size == 16)
return 7UL << (bitn);
else if (Size == 8)
return 3UL << (bitn);
else if (Size == 4)
return 1UL << (bitn);
else
LogMan::Throw::A(false, "Unexpected FPR size %d", Size);
}
/**
* @brief This is a temporary pass to detect simple multiblock dead FPR stores
*
* First pass computes which FPRs are read and written per block
*
* Second pass computes which FPRs are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead FPRs across multiple blocks.
*
* Third pass removes the dead stores.
*
*/
bool DeadFPRStoreElimination::Run(IREmitter *IREmit) {
std::map<OrderedNode*, FPRInfo> FPRMap;
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
// Pass 1
// Compute FPRs read/writes per block
// This is conservative and doesn't try to be smart about loads after writes
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
if (IsTrackedWriteFPR(Op->Offset, IROp->Size))
FPRMap[BlockNode].writes |= FPRBit(Op->Offset, IROp->Size);
else
FPRMap[BlockNode].reads |= FPRBit(Op->Offset, IROp->Size);
}
else if (IROp->Op == OP_STORECONTEXTINDEXED ||
IROp->Op == OP_LOADCONTEXTINDEXED) {
// We can't track through these
FPRMap[BlockNode].reads = -1;
}
else if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->CW<IR::IROp_LoadContext>();
FPRMap[BlockNode].reads |= FPRBit(Op->Offset, IROp->Size);
}
}
}
}
// Pass 2
// Compute FPRs that are stored, but always ovewritten in the next blocks
// Propagate the information a few times to eliminate more
for (int i = 0; i < PropagationRounds; i++)
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_JUMP) {
auto Op = IROp->CW<IR::IROp_Jump>();
OrderedNode *TargetNode = CurrentIR.GetNode(Op->Header.Args[0]);
// stores to remove are written by the next block but not read
FPRMap[BlockNode].kill = FPRMap[TargetNode].writes & ~(FPRMap[TargetNode].reads) & ~FPRMap[BlockNode].reads;
// FPRs that are written by the next block can be considered as written by this block, if not read
FPRMap[BlockNode].writes |= FPRMap[BlockNode].kill & ~FPRMap[BlockNode].reads;
}
else if (IROp->Op == OP_CONDJUMP) {
auto Op = IROp->CW<IR::IROp_CondJump>();
OrderedNode *TrueTargetNode = CurrentIR.GetNode(Op->TrueBlock);
OrderedNode *FalseTargetNode = CurrentIR.GetNode(Op->FalseBlock);
// stores to remove are written by the next blocks but not read
FPRMap[BlockNode].kill = FPRMap[TrueTargetNode].writes & ~(FPRMap[TrueTargetNode].reads) & ~FPRMap[BlockNode].reads;
FPRMap[BlockNode].kill &= FPRMap[FalseTargetNode].writes & ~(FPRMap[FalseTargetNode].reads) & ~FPRMap[BlockNode].reads;
// FPRs that are written by the next blocks can be considered as written by this block, if not read
FPRMap[BlockNode].writes |= FPRMap[BlockNode].kill & ~FPRMap[BlockNode].reads;
}
}
}
}
// Pass 3
// Remove the dead stores
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
// If this OP_STORECONTEXT is never read, remove it
if ((FPRMap[BlockNode].kill & FPRBit(Op->Offset, IROp->Size)) == FPRBit(Op->Offset, IROp->Size) && (FPRBit(Op->Offset, IROp->Size) != 0)) {
IREmit->Remove(CodeNode);
Changed = true;
}
}
}
}
}
return Changed;
}
FEXCore::IR::Pass* CreateDeadFPRStoreElimination() {
return new DeadFPRStoreElimination{};
}
}
@@ -1,119 +0,0 @@
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
class DeadFlagStoreElimination final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
struct FlagInfo {
uint64_t reads { 0 };
uint64_t writes { 0 };
uint64_t kill { 0 };
};
/**
* @brief This is a temporary pass to detect simple multiblock dead flag stores
*
* First pass computes which flags are read and written per block
*
* Second pass computes which flags are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead flags across multiple blocks.
*
* Third pass removes the dead stores.
*
*/
bool DeadFlagStoreElimination::Run(IREmitter *IREmit) {
std::map<OrderedNode*, FlagInfo> FlagMap;
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
// Pass 1
// Compute flags read/writes per block
// This is conservative and doesn't try to be smart about loads after writes
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->CW<IR::IROp_StoreFlag>();
FlagMap[BlockNode].writes |= 1UL << Op->Flag;
}
else if (IROp->Op == OP_INVALIDATEFLAGS) {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
FlagMap[BlockNode].writes |= Op->Flags;
}
else if (IROp->Op == OP_LOADFLAG) {
auto Op = IROp->CW<IR::IROp_LoadFlag>();
FlagMap[BlockNode].reads |= 1UL << Op->Flag;
}
}
}
}
// Pass 2
// Compute flags that are stored, but always ovewritten in the next blocks
// Propagate the information a few times to eliminate more
for (int i = 0; i < 5; i++)
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_JUMP) {
auto Op = IROp->CW<IR::IROp_Jump>();
OrderedNode *TargetNode = CurrentIR.GetNode(Op->Header.Args[0]);
// stores to remove are written by the next block but not read
FlagMap[BlockNode].kill = FlagMap[TargetNode].writes & ~(FlagMap[TargetNode].reads) & ~FlagMap[BlockNode].reads;
// Flags that are written by the next block can be considered as written by this block, if not read
FlagMap[BlockNode].writes |= FlagMap[BlockNode].kill & ~FlagMap[BlockNode].reads;
}
else if (IROp->Op == OP_CONDJUMP) {
auto Op = IROp->CW<IR::IROp_CondJump>();
OrderedNode *TrueTargetNode = CurrentIR.GetNode(Op->TrueBlock);
OrderedNode *FalseTargetNode = CurrentIR.GetNode(Op->FalseBlock);
// stores to remove are written by the next blocks but not read
FlagMap[BlockNode].kill = FlagMap[TrueTargetNode].writes & ~(FlagMap[TrueTargetNode].reads) & ~FlagMap[BlockNode].reads;
FlagMap[BlockNode].kill &= FlagMap[FalseTargetNode].writes & ~(FlagMap[FalseTargetNode].reads) & ~FlagMap[BlockNode].reads;
// Flags that are written by the next blocks can be considered as written by this block, if not read
FlagMap[BlockNode].writes |= FlagMap[BlockNode].kill & ~FlagMap[BlockNode].reads;
}
}
}
}
// Pass 3
// Remove the dead stores
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->CW<IR::IROp_StoreFlag>();
// If this StoreFlag is never read, remove it
if (FlagMap[BlockNode].kill & (1UL << Op->Flag)) {
IREmit->Remove(CodeNode);
Changed = true;
}
}
}
}
}
return Changed;
}
FEXCore::IR::Pass* CreateDeadFlagStoreElimination() {
return new DeadFlagStoreElimination{};
}
}
@@ -1,154 +0,0 @@
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
// Higher values might result in more stores getting eliminated but will make the optimization take more time
constexpr int PropagationRounds = 5;
namespace FEXCore::IR {
class DeadGPRStoreElimination final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
struct GPRInfo {
uint32_t reads { 0 };
uint32_t writes { 0 };
uint32_t kill { 0 };
};
bool IsFullGPR(uint32_t Offset, uint8_t Size) {
if (Size != 8)
return false;
if (Offset & 7)
return false;
if (Offset < 8 || Offset >= (17 * 8))
return false;
return true;
}
bool IsGPR(uint32_t Offset) {
if (Offset < 8 || Offset >= (17 * 8))
return false;
return true;
}
uint32_t GPRBit(uint32_t Offset) {
if (!IsGPR(Offset)) {
return 0;
}
return 1 << ((Offset - 8)/8);
}
/**
* @brief This is a temporary pass to detect simple multiblock dead GPR stores
*
* First pass computes which GPRs are read and written per block
*
* Second pass computes which GPRs are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead GPRs across multiple blocks.
*
* Third pass removes the dead stores.
*
*/
bool DeadGPRStoreElimination::Run(IREmitter *IREmit) {
std::map<OrderedNode*, GPRInfo> GPRMap;
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
// Pass 1
// Compute GPRs read/writes per block
// This is conservative and doesn't try to be smart about loads after writes
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
if (IsFullGPR(Op->Offset, IROp->Size))
GPRMap[BlockNode].writes |= GPRBit(Op->Offset);
else
GPRMap[BlockNode].reads |= GPRBit(Op->Offset);
}
else if (IROp->Op == OP_STORECONTEXTINDEXED ||
IROp->Op == OP_LOADCONTEXTINDEXED) {
// We can't track through these
GPRMap[BlockNode].reads = -1;
}
else if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->CW<IR::IROp_LoadContext>();
GPRMap[BlockNode].reads |= GPRBit(Op->Offset);
}
}
}
}
// Pass 2
// Compute GPRs that are stored, but always ovewritten in the next blocks
// Propagate the information a few times to eliminate more
for (int i = 0; i < PropagationRounds; i++)
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_JUMP) {
auto Op = IROp->CW<IR::IROp_Jump>();
OrderedNode *TargetNode = CurrentIR.GetNode(Op->Header.Args[0]);
// stores to remove are written by the next block but not read
GPRMap[BlockNode].kill = GPRMap[TargetNode].writes & ~(GPRMap[TargetNode].reads) & ~GPRMap[BlockNode].reads;
// GPRs that are written by the next block can be considered as written by this block, if not read
GPRMap[BlockNode].writes |= GPRMap[BlockNode].kill & ~GPRMap[BlockNode].reads;
}
else if (IROp->Op == OP_CONDJUMP) {
auto Op = IROp->CW<IR::IROp_CondJump>();
OrderedNode *TrueTargetNode = CurrentIR.GetNode(Op->TrueBlock);
OrderedNode *FalseTargetNode = CurrentIR.GetNode(Op->FalseBlock);
// stores to remove are written by the next blocks but not read
GPRMap[BlockNode].kill = GPRMap[TrueTargetNode].writes & ~(GPRMap[TrueTargetNode].reads) & ~GPRMap[BlockNode].reads;
GPRMap[BlockNode].kill &= GPRMap[FalseTargetNode].writes & ~(GPRMap[FalseTargetNode].reads) & ~GPRMap[BlockNode].reads;
// GPRs that are written by the next blocks can be considered as written by this block, if not read
GPRMap[BlockNode].writes |= GPRMap[BlockNode].kill & ~GPRMap[BlockNode].reads;
}
}
}
}
// Pass 3
// Remove the dead stores
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
// If this OP_STORECONTEXT is never read, remove it
if (GPRMap[BlockNode].kill & GPRBit(Op->Offset)) {
IREmit->Remove(CodeNode);
Changed = true;
}
}
}
}
}
return Changed;
}
FEXCore::IR::Pass* CreateDeadGPRStoreElimination() {
return new DeadGPRStoreElimination{};
}
}
@@ -0,0 +1,331 @@
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
namespace FEXCore::IR {
constexpr int PropagationRounds = 5;
class DeadStoreElimination final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
struct FlagInfo {
uint64_t reads { 0 };
uint64_t writes { 0 };
uint64_t kill { 0 };
};
struct GPRInfo {
uint32_t reads { 0 };
uint32_t writes { 0 };
uint32_t kill { 0 };
};
bool IsFullGPR(uint32_t Offset, uint8_t Size) {
if (Size != 8)
return false;
if (Offset & 7)
return false;
if (Offset < 8 || Offset >= (17 * 8))
return false;
return true;
}
bool IsGPR(uint32_t Offset) {
if (Offset < 8 || Offset >= (17 * 8))
return false;
return true;
}
uint32_t GPRBit(uint32_t Offset) {
if (!IsGPR(Offset)) {
return 0;
}
return 1 << ((Offset - 8)/8);
}
struct FPRInfo {
uint64_t reads { 0 };
uint64_t writes { 0 };
uint64_t kill { 0 };
};
bool IsFPR(uint32_t Offset) {
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto end = offsetof(FEXCore::Core::ThreadState, State.xmm[17][0]);
if (Offset < begin || Offset >= end)
return false;
return true;
}
bool IsTrackedWriteFPR(uint32_t Offset, uint8_t Size) {
if (Size != 16 && Size != 8 && Size != 4)
return false;
if (Offset & 15)
return false;
return IsFPR(Offset);
}
uint64_t FPRBit(uint32_t Offset, uint32_t Size) {
if (!IsFPR(Offset)) {
return 0;
}
auto begin = offsetof(FEXCore::Core::ThreadState, State.xmm[0][0]);
auto regn = (Offset - begin)/16;
auto bitn = regn * 3;
if (!IsTrackedWriteFPR(Offset, Size))
return 7UL << (bitn);
if (Size == 16)
return 7UL << (bitn);
else if (Size == 8)
return 3UL << (bitn);
else if (Size == 4)
return 1UL << (bitn);
else
LogMan::Msg::A("Unexpected FPR size %d", Size);
return 7UL << (bitn); // Return maximum on failure case
}
struct Info {
FlagInfo flag;
GPRInfo gpr;
FPRInfo fpr;
};
/**
* @brief This is a temporary pass to detect simple multiblock dead flag/gpr/fpr stores
*
* First pass computes which flags/gprs/fprs are read and written per block
*
* Second pass computes which flags/gprs/fprs are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead flags/gprs/fprs across multiple blocks.
*
* Third pass removes the dead stores.
*
*/
bool DeadStoreElimination::Run(IREmitter *IREmit) {
std::unordered_map<OrderedNode*, Info> InfoMap;
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
// Pass 1
// Compute flags/gprs/fprs read/writes per block
// This is conservative and doesn't try to be smart about loads after writes
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
//// Flags ////
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
auto& BlockInfo = InfoMap[BlockNode];
BlockInfo.flag.writes |= 1UL << Op->Flag;
} else if (IROp->Op == OP_INVALIDATEFLAGS) {
auto Op = IROp->C<IR::IROp_InvalidateFlags>();
auto& BlockInfo = InfoMap[BlockNode];
BlockInfo.flag.writes |= Op->Flags;
} else if (IROp->Op == OP_LOADFLAG) {
auto Op = IROp->C<IR::IROp_LoadFlag>();
auto& BlockInfo = InfoMap[BlockNode];
BlockInfo.flag.reads |= 1UL << Op->Flag;
} else if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->C<IR::IROp_StoreContext>();
auto& BlockInfo = InfoMap[BlockNode];
//// GPR ////
if (IsFullGPR(Op->Offset, IROp->Size))
BlockInfo.gpr.writes |= GPRBit(Op->Offset);
else
BlockInfo.gpr.reads |= GPRBit(Op->Offset);
//// FPR ////
if (IsTrackedWriteFPR(Op->Offset, IROp->Size))
BlockInfo.fpr.writes |= FPRBit(Op->Offset, IROp->Size);
else
BlockInfo.fpr.reads |= FPRBit(Op->Offset, IROp->Size);
} else if (IROp->Op == OP_STORECONTEXTINDEXED ||
IROp->Op == OP_LOADCONTEXTINDEXED) {
auto& BlockInfo = InfoMap[BlockNode];
//// GPR ////
// We can't track through these
BlockInfo.gpr.reads = -1;
//// FPR ////
// We can't track through these
BlockInfo.fpr.reads = -1;
} else if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->C<IR::IROp_LoadContext>();
auto& BlockInfo = InfoMap[BlockNode];
//// GPR ////
BlockInfo.gpr.reads |= GPRBit(Op->Offset);
//// FPR ////
BlockInfo.fpr.reads |= FPRBit(Op->Offset, IROp->Size);
}
}
}
}
// Pass 2
// Compute flags/gprs/fprs that are stored, but always ovewritten in the next blocks
// Propagate the information a few times to eliminate more
for (int i = 0; i < PropagationRounds; i++)
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
auto CodeBlock = BlockIROp->C<IROp_CodeBlock>();
auto IROp = CurrentIR.GetNode(CurrentIR.GetNode(CodeBlock->Last)->Header.Previous)->Op(CurrentIR.GetData());
if (IROp->Op == OP_JUMP) {
auto Op = IROp->C<IR::IROp_Jump>();
OrderedNode *TargetNode = CurrentIR.GetNode(Op->Header.Args[0]);
auto& BlockInfo = InfoMap[BlockNode];
auto& TargetInfo = InfoMap[TargetNode];
//// Flags ////
// stores to remove are written by the next block but not read
BlockInfo.flag.kill = TargetInfo.flag.writes & ~(TargetInfo.flag.reads) & ~BlockInfo.flag.reads;
// Flags that are written by the next block can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
//// GPRs ////
// stores to remove are written by the next block but not read
BlockInfo.gpr.kill = TargetInfo.gpr.writes & ~(TargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
// GPRs that are written by the next block can be considered as written by this block, if not read
BlockInfo.gpr.writes |= BlockInfo.gpr.kill & ~BlockInfo.gpr.reads;
//// FPRs ////
// stores to remove are written by the next block but not read
BlockInfo.fpr.kill = TargetInfo.fpr.writes & ~(TargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
// FPRs that are written by the next block can be considered as written by this block, if not read
BlockInfo.fpr.writes |= BlockInfo.fpr.kill & ~BlockInfo.fpr.reads;
} else if (IROp->Op == OP_CONDJUMP) {
auto Op = IROp->C<IR::IROp_CondJump>();
OrderedNode *TrueTargetNode = CurrentIR.GetNode(Op->TrueBlock);
OrderedNode *FalseTargetNode = CurrentIR.GetNode(Op->FalseBlock);
auto& BlockInfo = InfoMap[BlockNode];
auto& TrueTargetInfo = InfoMap[TrueTargetNode];
auto& FalseTargetInfo = InfoMap[FalseTargetNode];
//// Flags ////
// stores to remove are written by the next blocks but not read
BlockInfo.flag.kill = TrueTargetInfo.flag.writes & ~(TrueTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
BlockInfo.flag.kill &= FalseTargetInfo.flag.writes & ~(FalseTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
// Flags that are written by the next blocks can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
//// GPRs ////
// stores to remove are written by the next blocks but not read
BlockInfo.gpr.kill = TrueTargetInfo.gpr.writes & ~(TrueTargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
BlockInfo.gpr.kill &= FalseTargetInfo.gpr.writes & ~(FalseTargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
// GPRs that are written by the next blocks can be considered as written by this block, if not read
BlockInfo.gpr.writes |= BlockInfo.gpr.kill & ~BlockInfo.gpr.reads;
//// FPRs ////
// stores to remove are written by the next blocks but not read
BlockInfo.fpr.kill = TrueTargetInfo.fpr.writes & ~(TrueTargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
BlockInfo.fpr.kill &= FalseTargetInfo.fpr.writes & ~(FalseTargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
// FPRs that are written by the next blocks can be considered as written by this block, if not read
BlockInfo.fpr.writes |= BlockInfo.fpr.kill & ~BlockInfo.fpr.reads;
}
}
}
// Pass 3
// Remove the dead stores
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
//// Flags ////
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
auto& BlockInfo = InfoMap[BlockNode];
// If this StoreFlag is never read, remove it
if (BlockInfo.flag.kill & (1UL << Op->Flag)) {
IREmit->Remove(CodeNode);
Changed = true;
}
} else if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->C<IR::IROp_StoreContext>();
auto& BlockInfo = InfoMap[BlockNode];
//// GPRs ////
// If this OP_STORECONTEXT is never read, remove it
if (BlockInfo.gpr.kill & GPRBit(Op->Offset)) {
IREmit->Remove(CodeNode);
Changed = true;
}
//// FPRs ////
// If this OP_STORECONTEXT is never read, remove it
if ((BlockInfo.fpr.kill & FPRBit(Op->Offset, IROp->Size)) == FPRBit(Op->Offset, IROp->Size) && (FPRBit(Op->Offset, IROp->Size) != 0)) {
IREmit->Remove(CodeNode);
Changed = true;
}
}
}
}
}
return Changed;
}
FEXCore::IR::Pass* CreateDeadStoreElimination() {
return new DeadStoreElimination{};
}
}
+25 -15
View File
@@ -6,6 +6,13 @@
namespace FEXCore::IR {
// struct to avoid zero-initialization
struct RemapNode {
IR::OrderedNodeWrapper::NodeOffsetType NodeID;
};
static_assert(sizeof(RemapNode) == 4);
class IRCompaction final : public FEXCore::IR::Pass {
public:
IRCompaction();
@@ -14,7 +21,7 @@ public:
private:
static constexpr size_t AlignSize = 0x2000;
OpDispatchBuilder LocalBuilder;
std::vector<IR::OrderedNodeWrapper::NodeOffsetType> OldToNewRemap;
std::vector<RemapNode> OldToNewRemap;
struct CodeBlockData {
OrderedNode *OldNode;
OrderedNode *NewNode;
@@ -35,7 +42,10 @@ bool IRCompaction::Run(IREmitter *IREmit) {
if (OldToNewRemap.size() < NodeCount) {
OldToNewRemap.resize(std::max(OldToNewRemap.size() * 2U, AlignUp(NodeCount, AlignSize)));
}
memset(&OldToNewRemap.at(0), 0xFF, NodeCount * sizeof(IR::OrderedNodeWrapper::NodeOffsetType));
#ifndef NDEBUG
memset(&OldToNewRemap.at(0), 0xFF, NodeCount * sizeof(RemapNode));
#endif
GeneratedCodeBlocks.clear();
// Reset our local working list
@@ -46,7 +56,6 @@ bool IRCompaction::Run(IREmitter *IREmit) {
uintptr_t LocalDataBegin = LocalIR.GetData();
uintptr_t ListBegin = CurrentIR.GetListData();
uintptr_t DataBegin = CurrentIR.GetData();
auto HeaderNode = CurrentIR.GetHeaderNode();
auto HeaderOp = CurrentIR.GetHeader();
@@ -67,9 +76,9 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// Then create all the ops inside the code blocks
// Zero is always zero(invalid)
OldToNewRemap[0] = 0;
auto LocalHeaderOp = LocalBuilder._IRHeader(OrderedNodeWrapper::WrapOffset(0).GetNode(ListBegin), HeaderOp->Entry, HeaderOp->BlockCount, HeaderOp->ShouldInterpret);
OldToNewRemap[CurrentIR.GetID(HeaderNode)] = LocalIR.GetID(LocalHeaderOp.Node);
OldToNewRemap[0].NodeID = 0;
auto LocalHeaderOp = LocalBuilder._IRHeader(OrderedNodeWrapper::WrapOffset(0).GetNode(ListBegin), HeaderOp->Entry, HeaderOp->BlockCount);
OldToNewRemap[CurrentIR.GetID(HeaderNode)].NodeID = LocalIR.GetID(LocalHeaderOp.Node);
{
// Generate our codeblocks and link them together
@@ -77,7 +86,7 @@ bool IRCompaction::Run(IREmitter *IREmit) {
LogMan::Throw::A(BlockHeader->Op == OP_CODEBLOCK, "IR type failed to be a code block");
auto LocalBlockIRNode = LocalBuilder._CodeBlock(LocalHeaderOp, LocalHeaderOp); // Use LocalHeaderOp as a dummy arg for now
OldToNewRemap[CurrentIR.GetID(BlockNode)] = LocalIR.GetID(LocalBlockIRNode.Node);
OldToNewRemap[CurrentIR.GetID(BlockNode)].NodeID = LocalIR.GetID(LocalBlockIRNode.Node);
GeneratedCodeBlocks.emplace_back(CodeBlockData{BlockNode, LocalBlockIRNode});
}
@@ -112,7 +121,7 @@ bool IRCompaction::Run(IREmitter *IREmit) {
// Set our map remapper to map the new location
// Even nodes that don't have a destination need to be in this map
// Need to be able to remap branch targets any other bits
OldToNewRemap[CurrentIR.GetID(CodeNode)] = LocalIR.GetID(LocalPair.Node);
OldToNewRemap[CurrentIR.GetID(CodeNode)].NodeID = LocalIR.GetID(LocalPair.Node);
if (i == 0) {
FirstNode.OldNode = CodeNode;
@@ -137,20 +146,21 @@ bool IRCompaction::Run(IREmitter *IREmit) {
{
// Fixup the arguments of all the IROps
for (auto &Block : GeneratedCodeBlocks) {
auto BlockIROp = CurrentIR.GetOp<FEXCore::IR::IROp_CodeBlock>(Block.OldNode);
auto BlockIROp = LocalIR.GetOp<FEXCore::IR::IROp_CodeBlock>(Block.NewNode);
LogMan::Throw::A(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
for (auto [CodeNode, IROp] : CurrentIR.GetCode(Block.OldNode)) {
auto [LocalNode, LocalIROp] = LocalIR.at(OldToNewRemap[CurrentIR.GetID(CodeNode)])();
for (auto [LocalNode, LocalIROp] : LocalIR.GetCode(Block.NewNode)) {
// Now that we have the op copied over, we need to modify SSA values to point to the new correct locations
// This doesn't use IR::GetArgs(Op) because we need to remap all SSA nodes
// Including ones that we don't RA
uint8_t NumArgs = IROp->NumArgs;
uint8_t NumArgs = LocalIROp->NumArgs;
for (uint8_t i = 0; i < NumArgs; ++i) {
uint32_t OldArg = IROp->Args[i].ID();
LogMan::Throw::A(OldToNewRemap[OldArg] != ~0U, "Tried remapping unfound node %%ssa%d", OldArg);
LocalIROp->Args[i].NodeOffset = OldToNewRemap[OldArg] * sizeof(OrderedNode);
uint32_t OldArg = LocalIROp->Args[i].ID();
#ifndef NDEBUG
LogMan::Throw::A(OldToNewRemap[OldArg].NodeID != ~0U, "Tried remapping unfound node %%ssa%d", OldArg);
#endif
LocalIROp->Args[i].NodeOffset = OldToNewRemap[OldArg].NodeID * sizeof(OrderedNode);
}
}
}
+12 -10
View File
@@ -50,11 +50,13 @@ bool IRValidation::Run(IREmitter *IREmit) {
auto HeaderOp = CurrentIR.GetHeader();
LogMan::Throw::A(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
IR::RegisterAllocationPass * RAPass{};
if (Manager->HasRAPass() && !HeaderOp->ShouldInterpret) {
RAPass = Manager->GetRAPass();
IR::RegisterAllocationData * RAData{};
if (Manager->HasRAPass()) {
RAData = Manager->GetRAPass() ? Manager->GetRAPass()->GetAllocationData() : nullptr;
}
NodeIsLive.Set(1); // IRHEADER
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LogMan::Throw::A(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
@@ -81,12 +83,12 @@ bool IRValidation::Run(IREmitter *IREmit) {
Warnings << "%ssa" << ID << ": Destination created but had no uses" << std::endl;
}
if (RAPass) {
if (RAData) {
// If we have a register allocator then the destination needs to be assigned a register and class
uint64_t Reg = RAPass->GetNodeRegister(ID);
auto PhyReg = RAData->GetNodeRegister(ID);
FEXCore::IR::RegisterClassType ExpectedClass = IR::GetRegClass(IROp->Op);
FEXCore::IR::RegisterClassType AssignedClass = FEXCore::IR::RegisterClassType{uint32_t(Reg >> 32)};
FEXCore::IR::RegisterClassType AssignedClass = FEXCore::IR::RegisterClassType{PhyReg.Class};
// If no register class was assigned
if (AssignedClass == IR::InvalidClass) {
@@ -95,7 +97,7 @@ bool IRValidation::Run(IREmitter *IREmit) {
}
// If no physical register was assigned
if ((uint32_t)Reg == ~0U) {
if (PhyReg.Reg == IR::InvalidReg) {
HadError |= true;
Errors << "%ssa" << ID << ": Had destination but with no register assigned" << std::endl;
}
@@ -260,7 +262,7 @@ bool IRValidation::Run(IREmitter *IREmit) {
HadWarning = false;
if (HadError || HadWarning) {
FEXCore::IR::Dump(&Out, &CurrentIR, RAPass);
FEXCore::IR::Dump(&Out, &CurrentIR, RAData);
if (HadError) {
Out << "Errors:" << std::endl << Errors.str() << std::endl;
@@ -269,8 +271,8 @@ bool IRValidation::Run(IREmitter *IREmit) {
if (HadWarning) {
Out << "Warnings:" << std::endl << Warnings.str() << std::endl;
}
LogMan::Msg::E("%s", Out.str().c_str());
fprintf(stderr, "%s", Out.str().c_str());
}
return false;
File diff suppressed because it is too large. Load diff
@@ -1,15 +1,14 @@
#pragma once
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/RegisterAllocationData.h>
#include <vector>
namespace FEXCore::IR {
template<bool>
class IRListView;
class RegisterAllocationPass : public FEXCore::IR::Pass {
public:
bool HasFullRA() const { return HadFullRA; }
uint32_t SpillSlots() const { return SpillSlotCount; }
virtual void AllocateRegisterSet(uint32_t RegisterCount, uint32_t ClassCount) = 0;
virtual void AddRegisters(FEXCore::IR::RegisterClassType Class, uint32_t RegisterCount) = 0;
@@ -33,10 +32,14 @@ class RegisterAllocationPass : public FEXCore::IR::Pass {
* @{ */
/**
* @brief Returns the register and class encoded together
* Top 32bits is the class, lower 32bits is the register
* @brief Returns the register and class map array
*/
virtual uint64_t GetNodeRegister(uint32_t Node) = 0;
virtual RegisterAllocationData *GetAllocationData() = 0;
/**
* @brief Returns and transfers ownership of the register and class map array
*/
virtual std::unique_ptr<RegisterAllocationData, RegisterAllocationDataDeleter> PullAllocationData() = 0;
/** @} */
protected:
@@ -44,13 +44,10 @@ bool IsStaticAllocFpr(uint32_t Offset, RegisterClassType Class, bool AllowGpr) {
bool StaticRegisterAllocationPass::Run(IREmitter *IREmit) {
auto CurrentIR = IREmit->ViewIR();
if (CurrentIR.GetHeader()->ShouldInterpret)
return false;
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
IREmit->SetWriteCursor(CodeNode);
if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->CW<IR::IROp_LoadContext>();
@@ -79,7 +76,7 @@ bool StaticRegisterAllocationPass::Run(IREmitter *IREmit) {
}
auto StaticClass = GeneralClass == GPRClass ? GPRFixedClass : FPRFixedClass;
OrderedNode *sraReg = IREmit->_StoreRegister(val, false, Op->Offset, GeneralClass, StaticClass, Op->Header.Size);
IREmit->_StoreRegister(val, false, Op->Offset, GeneralClass, StaticClass, Op->Header.Size);
IREmit->Remove(CodeNode);
}
@@ -81,6 +81,8 @@ bool ValueDominanceValidation::Run(IREmitter *IREmit) {
uint8_t NumArgs = IR::GetArgs(IROp->Op);
for (uint32_t i = 0; i < NumArgs; ++i) {
if (IROp->Args[i].IsInvalid()) continue;
if (CurrentIR.GetOp<IROp_Header>(IROp->Args[i])->Op == OP_IRHEADER) continue;
OrderedNodeWrapper Arg = IROp->Args[i];
// We must ensure domininance of all SSA arguments
+14 -15
View File
@@ -1,3 +1,4 @@
#include <FEXCore/Utils/Common/MathUtils.h>
#include <FEXCore/Utils/ELFLoader.h>
#include <FEXCore/Utils/LogManager.h>
#include <cstring>
@@ -201,6 +202,9 @@ bool ELFContainer::LoadELF_32() {
DynamicProgram = Header._32.e_type != ET_EXEC;
// Default BRK size
BRKSize = 4096;
return true;
}
@@ -239,6 +243,9 @@ bool ELFContainer::LoadELF_64() {
DynamicProgram = Header._64.e_type != ET_EXEC;
// Default BRK size
BRKSize = 0x1000'0000;
return true;
}
@@ -321,8 +328,8 @@ void ELFContainer::CalculateMemoryLayouts() {
//
// We need to ignore such empty sections, or we will mistakenly assume the elf starts at zero.
if (hdr->p_memsz > 0) {
MinPhysAddr = std::min(MinPhysAddr, hdr->p_paddr);
MaxPhysAddr = std::max(MaxPhysAddr, hdr->p_paddr + hdr->p_memsz);
MinPhysAddr = std::min(MinPhysAddr, static_cast<uint64_t>(hdr->p_paddr));
MaxPhysAddr = std::max(MaxPhysAddr, static_cast<uint64_t>(hdr->p_paddr + hdr->p_memsz));
}
if (hdr->p_type == PT_TLS) {
TLSHeader._64 = hdr;
@@ -330,6 +337,11 @@ void ELFContainer::CalculateMemoryLayouts() {
}
}
// Calculate BRK
MaxPhysAddr = AlignUp(MaxPhysAddr, 4096);
BRKBase = MaxPhysAddr;
MaxPhysAddr += BRKSize;
PhysMemSize = MaxPhysAddr - MinPhysAddr;
MinPhysicalMemoryLocation = MinPhysAddr;
@@ -674,8 +686,6 @@ void ELFContainer::PrintProgramHeaders() const {
if (Mode == MODE_32BIT) {
LogMan::Throw::A(Header._32.e_shstrndx < SectionHeaders.size(),
"String index section is wrong index!");
Elf32_Shdr const *StrHeader = SectionHeaders.at(Header._32.e_shstrndx)._32;
char const *SHStrings = &RawFile.at(StrHeader->sh_offset);
for (uint32_t i = 0; i < ProgramHeaders.size(); ++i) {
Elf32_Phdr const *hdr = ProgramHeaders.at(i)._32;
LogMan::Msg::I("Type: %d", hdr->p_type);
@@ -691,8 +701,6 @@ void ELFContainer::PrintProgramHeaders() const {
else {
LogMan::Throw::A(Header._64.e_shstrndx < SectionHeaders.size(),
"String index section is wrong index!");
Elf64_Shdr const *StrHeader = SectionHeaders.at(Header._64.e_shstrndx)._64;
char const *SHStrings = &RawFile.at(StrHeader->sh_offset);
for (uint32_t i = 0; i < ProgramHeaders.size(); ++i) {
Elf64_Phdr const *hdr = ProgramHeaders.at(i)._64;
LogMan::Msg::I("Type: %d", hdr->p_type);
@@ -884,9 +892,6 @@ void ELFContainer::FixupRelocations(void *ELFBase, uint64_t GuestELFBase, Symbol
Elf64_Shdr const *GOTHeader {nullptr};
Elf64_Shdr const *DynSymHeader {nullptr};
Elf64_Shdr const *StrHeader = SectionHeaders.at(Header._64.e_shstrndx)._64;
char const *SHStrings = &RawFile.at(StrHeader->sh_offset);
Elf64_Shdr const *StringTableHeader{nullptr};
char const *StrTab{nullptr};
@@ -1206,9 +1211,6 @@ void ELFContainer::GetInitLocations(uint64_t GuestELFBase, std::vector<uint64_t>
for (uint32_t i = 0; i < SectionHeaders.size(); ++i) {
Elf32_Shdr const *hdr = SectionHeaders.at(i)._32;
if (hdr->sh_type == SHT_DYNAMIC) {
Elf32_Shdr const *StrHeader = SectionHeaders.at(hdr->sh_link)._32;
char const *SHStrings = &RawFile.at(StrHeader->sh_offset);
size_t Entries = hdr->sh_size / hdr->sh_entsize;
for (size_t j = 0; i < Entries; ++j) {
Elf32_Dyn const *Dynamic = reinterpret_cast<Elf32_Dyn const*>(&RawFile.at(hdr->sh_offset + j * hdr->sh_entsize));
@@ -1236,9 +1238,6 @@ void ELFContainer::GetInitLocations(uint64_t GuestELFBase, std::vector<uint64_t>
for (uint32_t i = 0; i < SectionHeaders.size(); ++i) {
Elf64_Shdr const *hdr = SectionHeaders.at(i)._64;
if (hdr->sh_type == SHT_DYNAMIC) {
Elf64_Shdr const *StrHeader = SectionHeaders.at(hdr->sh_link)._64;
char const *SHStrings = &RawFile.at(StrHeader->sh_offset);
size_t Entries = hdr->sh_size / hdr->sh_entsize;
for (size_t j = 0; i < Entries; ++j) {
Elf64_Dyn const *Dynamic = reinterpret_cast<Elf64_Dyn const*>(&RawFile.at(hdr->sh_offset + j * hdr->sh_entsize));
+2 -2
View File
@@ -79,10 +79,10 @@ This is an intrusive allocator that is used by the `OpDispatchBuilder` for stori
### OpDispatchBuilder
OpDispatchBuilder provides two routines for handling the IR outside of the class
* `IRListView<false> ViewIR();`
* `IRListView ViewIR();`
* Returns a wrapper container class the allows you to view the IR. This doesn't take ownership of the IR data.
* If the OpDispatcherBuilder changes its IR then changes are also visible to this class
* `IRListView<true> *CreateIRCopy()`
* `IRListView *CreateIRCopy()`
* As the name says, it creates a new copy of the IR that is in the OpDispatchBuilder
* Copying the IR only copies the memory used and doesn't have any free space for optimizations after this copy operation
* Useful for tiered recompilers, AOT, and offline analysis
+11
View File
@@ -3,7 +3,9 @@
#include <list>
#include <memory>
#include <optional>
#include <stdint.h>
#include <unordered_map>
namespace FEXCore::Config {
enum ConfigOption {
@@ -22,6 +24,7 @@ namespace FEXCore::Config {
CONFIG_ABI_LOCAL_FLAGS,
CONFIG_ABI_NO_PF,
CONFIG_DUMPIR,
CONFIG_VALIDATE_IR_PARSER,
CONFIG_SILENTLOGS,
CONFIG_ENVIRONMENT,
CONFIG_OUTPUTLOG,
@@ -31,6 +34,8 @@ namespace FEXCore::Config {
CONFIG_INTERPRETER_INSTALLED,
CONFIG_APP_FILENAME,
CONFIG_DEBUG_DISABLE_OPTIMIZATION_PASSES,
CONFIG_AOTIR_GENERATE,
CONFIG_AOTIR_LOAD
};
enum ConfigCore {
@@ -39,6 +44,12 @@ namespace FEXCore::Config {
CONFIG_CUSTOM,
};
enum ConfigSMCChecks {
CONFIG_SMC_NONE,
CONFIG_SMC_MMAN,
CONFIG_SMC_FULL,
};
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config);
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, std::string const &Config);
uint64_t GetConfig(FEXCore::Context::Context *CTX, ConfigOption Option);
+2 -2
View File
@@ -5,8 +5,8 @@
namespace FEXCore {
namespace IR {
template<bool Copy>
class IRListView;
class RegisterAllocationData;
}
namespace Core {
@@ -44,7 +44,7 @@ class LLVMCore;
* @return An executable function pointer that is theoretically compiled from this point.
* Is actually a function pointer of type `void (FEXCore::Core::ThreadState *Thread)
*/
virtual void *CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) = 0;
virtual void *CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) = 0;
/**
* @brief Function for mapping memory in to the CPUBackend's visible space. Allows setting up virtual mappings if required
+10
View File
@@ -6,6 +6,10 @@
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/CPUID.h>
#include <istream>
#include <ostream>
#include <memory>
namespace FEXCore {
class CodeLoader;
}
@@ -229,4 +233,10 @@ namespace FEXCore::Context {
void SetSignalDelegator(FEXCore::Context::Context *CTX, FEXCore::SignalDelegator *SignalDelegation);
void SetSyscallHandler(FEXCore::Context::Context *CTX, FEXCore::HLE::SyscallHandler *Handler);
FEXCore::CPUID::FunctionResults RunCPUIDFunction(FEXCore::Context::Context *CTX, uint32_t Function, uint32_t Leaf);
void AddNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length, uintptr_t Offset, const std::string& Name);
void RemoveNamedRegion(FEXCore::Context::Context *CTX, uintptr_t Base, uintptr_t Length);
void SetAOTIRLoader(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::istream>(const std::string&)> CacheReader);
bool WriteAOTIR(FEXCore::Context::Context *CTX, std::function<std::unique_ptr<std::ostream>(const std::string&)> CacheWriter);
void FlushCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length);
}
+6
View File
@@ -22,6 +22,7 @@ namespace FEXCore::Core {
struct {
uint32_t base;
} gdt[32];
uint16_t FCW;
};
static_assert(offsetof(CPUState, xmm) % 16 == 0, "xmm needs to be 128bit aligned!");
@@ -47,6 +48,11 @@ namespace FEXCore::Core {
static_assert(std::is_standard_layout<ThreadState>::value, "This needs to be standard layout");
#ifdef PAGE_SIZE
static_assert(PAGE_SIZE == 4096, "FEX only supports 4k pages");
#undef PAGE_SIZE
#endif
constexpr uint64_t PAGE_SIZE = 4096;
std::string_view const& GetFlagName(unsigned Flag);
+6 -4
View File
@@ -2,8 +2,6 @@
#include <cstddef>
#include <cstdint>
#include <bits/types/sigset_t.h>
namespace FEXCore {
namespace x86_64 {
// uc_flags flags
@@ -74,18 +72,22 @@ namespace FEXCore {
};
static_assert(sizeof(FEXCore::x86_64::mcontext_t) == 256, "This needs to be the right size");
struct __attribute__((packed)) sigset_t {
uint64_t val[16];
};
static_assert(sizeof(FEXCore::x86_64::sigset_t) == 128, "This needs to be the right size");
struct __attribute__((packed)) ucontext_t {
uint64_t uc_flags;
FEXCore::x86_64::ucontext_t *uc_link;
FEXCore::x86_64::stack_t uc_stack;
FEXCore::x86_64::mcontext_t uc_mcontext;
sigset_t uc_sigmask;
FEXCore::x86_64::sigset_t uc_sigmask;
FEXCore::x86_64::_libc_fpstate __fpregs_mem;
uint64_t __ssp[4];
};
static_assert(offsetof(FEXCore::x86_64::ucontext_t, uc_mcontext) == 40, "Needs to be correct");
static_assert(sizeof(sigset_t) == 128, "This needs to be the right size");
static_assert(sizeof(FEXCore::x86_64::ucontext_t) == 968, "This needs to be the right size");
}
@@ -3,12 +3,14 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Event.h>
#include <map>
#include <unordered_map>
#include <thread>
namespace FEXCore {
class BlockCache;
class LookupCache;
class CompileService;
}
@@ -59,6 +61,14 @@ namespace FEXCore::Core {
SIGNALEVENT_RETURN,
};
struct LocalIREntry {
uint64_t StartAddr;
uint64_t Length;
std::unique_ptr<FEXCore::IR::IRListView, FEXCore::IR::IRListViewDeleter> IR;
std::unique_ptr<FEXCore::IR::RegisterAllocationData, FEXCore::IR::RegisterAllocationDataDeleter> RAData;
std::unique_ptr<FEXCore::Core::DebugData> DebugData;
};
struct InternalThreadState {
FEXCore::Core::ThreadState State;
@@ -71,14 +81,10 @@ namespace FEXCore::Core {
std::unique_ptr<FEXCore::IR::OpDispatchBuilder> OpDispatcher;
std::shared_ptr<FEXCore::CPU::CPUBackend> CPUBackend;
std::shared_ptr<FEXCore::CPU::CPUBackend> IntBackend;
std::unique_ptr<FEXCore::CPU::CPUBackend> FallbackBackend;
std::unique_ptr<FEXCore::CPU::CPUBackend> CPUBackend;
std::unique_ptr<FEXCore::LookupCache> LookupCache;
std::unique_ptr<FEXCore::BlockCache> BlockCache;
std::unordered_map<uint64_t, std::unique_ptr<FEXCore::IR::IRListView<true>>> IRLists;
std::unordered_map<uint64_t, FEXCore::Core::DebugData> DebugData;
std::unordered_map<uint64_t, LocalIREntry> LocalIRCache;
std::unique_ptr<FEXCore::Frontend::Decoder> FrontendDecoder;
std::unique_ptr<FEXCore::IR::PassManager> PassManager;
+16 -8
View File
@@ -2,11 +2,13 @@
#include <array>
#include <cassert>
#include <cstdint>
#include <string.h>
#include <sstream>
#include <tuple>
namespace FEXCore::IR {
class RegisterAllocationPass;
class RegisterAllocationData;
/**
* @brief The IROp_Header is an dynamically sized array
@@ -287,29 +289,29 @@ struct MemOffsetType final {
};
struct TypeDefinition final {
uint8_t Val;
operator uint8_t() const {
uint16_t Val;
operator uint16_t() const {
return Val;
}
static TypeDefinition Create(uint8_t Bytes) {
TypeDefinition Type{};
Type.Val = Bytes << 2;
Type.Val = Bytes << 8;
return Type;
}
static TypeDefinition Create(uint8_t Bytes, uint8_t Elements) {
TypeDefinition Type{};
Type.Val = (Bytes << 2) | (Elements & 0b11);
Type.Val = (Bytes << 8) | (Elements & 255);
return Type;
}
uint8_t Bytes() const {
return Val >> 2;
return Val >> 8;
}
uint8_t Elements() const {
return Val & 0b11;
return Val & 255;
}
};
@@ -324,6 +326,11 @@ struct FenceType final {
constexpr bool operator!=(FenceType const &rhs) const { return !operator==(rhs); }
};
struct SHA256Sum final {
uint8_t data[32];
bool operator<(SHA256Sum const &rhs) const { return memcmp(data, rhs.data, sizeof(data)) < 0; }
};
class NodeIterator;
/* This iterator can be used to step though nodes.
@@ -455,10 +462,11 @@ public:
}
};
template<bool>
class IRListView;
class IREmitter;
void Dump(std::stringstream *out, IRListView<false> const* IR, IR::RegisterAllocationPass *RAPass);
void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationData *RAData);
IREmitter* Parse(std::istream *in);
template<typename Type>
inline uint32_t NodeWrapperBase<Type>::ID() const { return NodeOffset / sizeof(IR::OrderedNode); }
+8 -4
View File
@@ -21,8 +21,8 @@ friend class FEXCore::IR::PassManager;
ResetWorkingList();
}
IRListView<false> ViewIR() { return IRListView<false>(&Data, &ListData); }
IRListView<true> *CreateIRCopy() { return new IRListView<true>(&Data, &ListData); }
IRListView ViewIR() { return IRListView(&Data, &ListData, false); }
IRListView *CreateIRCopy() { return new IRListView(&Data, &ListData, true); }
void ResetWorkingList();
/**
@@ -40,7 +40,8 @@ friend class FEXCore::IR::PassManager;
IRPair<IROp_Constant> _Constant(uint8_t Size, uint64_t Constant) {
auto Op = AllocateOp<IROp_Constant, IROps::OP_CONSTANT>();
Op.first->Constant = Constant;
uint64_t Mask = ~0ULL >> (Size - 64);
Op.first->Constant = (Constant & Mask);
Op.first->Header.Size = Size / 8;
Op.first->Header.ElementSize = Size / 8;
Op.first->Header.NumArgs = 0;
@@ -504,9 +505,12 @@ friend class FEXCore::IR::PassManager;
LogMan::Throw::A(rhs.ListData.BackingSize() <= ListData.BackingSize(), "Trying to take ownership of data that is too large");
Data.CopyData(rhs.Data);
ListData.CopyData(rhs.ListData);
InvalidNode = rhs.InvalidNode;
InvalidNode = rhs.InvalidNode->Wrapped(rhs.ListData.Begin()).GetNode(ListData.Begin());
CurrentWriteCursor = rhs.CurrentWriteCursor;
CodeBlocks = rhs.CodeBlocks;
for (auto& CodeBlock: CodeBlocks) {
CodeBlock = CodeBlock->Wrapped(rhs.ListData.Begin()).GetNode(ListData.Begin());
}
}
void SetWriteCursor(OrderedNode *Node) {
+43 -10
View File
@@ -8,6 +8,8 @@
#include <cstring>
#include <tuple>
#include <vector>
#include <istream>
#include <ostream>
namespace FEXCore::IR {
/**
@@ -62,17 +64,16 @@ class IntrusiveAllocator final {
uintptr_t Data;
};
template<bool Copy>
class IRListView final {
public:
IRListView() = delete;
IRListView(IRListView<Copy> &&) = delete;
IRListView(IRListView &&) = delete;
IRListView(IntrusiveAllocator *Data, IntrusiveAllocator *List) {
IRListView(IntrusiveAllocator *Data, IntrusiveAllocator *List, bool _IsCopy) : IsCopy(_IsCopy) {
DataSize = Data->Size();
ListSize = List->Size();
if (Copy) {
if (IsCopy) {
IRData = malloc(DataSize + ListSize);
ListData = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(IRData) + DataSize);
memcpy(IRData, reinterpret_cast<void*>(Data->Begin()), DataSize);
@@ -85,24 +86,46 @@ public:
}
}
IRListView<true>(IRListView<true> *Old) {
IRListView(IRListView *Old, bool _IsCopy) : IsCopy(_IsCopy) {
DataSize = Old->DataSize;
ListSize = Old->ListSize;
if (IsCopy) {
IRData = malloc(DataSize + ListSize);
ListData = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(IRData) + DataSize);
memcpy(IRData, Old->IRData, DataSize);
memcpy(ListData, Old->ListData, ListSize);
} else {
IRData = Old->IRData;
ListData = Old->ListData;
}
}
IRListView(std::istream& stream) : IsCopy(true) {
stream.read((char*)&DataSize, sizeof(DataSize));
stream.read((char*)&ListSize, sizeof(ListSize));
IRData = malloc(DataSize + ListSize);
ListData = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(IRData) + DataSize);
memcpy(IRData, Old->IRData, DataSize);
memcpy(ListData, Old->ListData, ListSize);
stream.read((char*)IRData, DataSize);
stream.read((char*)ListData, ListSize);
}
~IRListView() {
if (Copy) {
if (IsCopy) {
free (IRData);
// ListData is just offset from IRData
}
}
IRListView<true> *CreateCopy() {
return new IRListView<true>(this);
void Serialize(std::ostream& stream) {
stream.write((char*)&DataSize, sizeof(DataSize));
stream.write((char*)&ListSize, sizeof(ListSize));
stream.write((char*)IRData, DataSize);
stream.write((char*)ListData, ListSize);
}
IRListView *CreateCopy() {
return new IRListView(this, true);
}
uintptr_t const GetData() const { return reinterpret_cast<uintptr_t>(IRData); }
@@ -149,6 +172,7 @@ public:
return Wrapper.GetNode(GetListData());
}
bool IsShared {false};
private:
struct BlockRange {
using iterator = NodeIterator;
@@ -259,6 +283,15 @@ private:
void *ListData;
size_t DataSize;
size_t ListSize;
bool IsCopy;
};
struct IRListViewDeleter {
void operator()(IRListView* r) {
if (!r->IsShared) {
delete r;
}
}
};
}
@@ -0,0 +1,55 @@
#pragma once
#include "IR.h"
namespace FEXCore::IR {
union PhysicalRegister {
uint8_t Raw;
struct {
uint8_t Reg: 5;
uint8_t Class: 3;
};
bool operator==(const PhysicalRegister &Other) const {
return Raw == Other.Raw;
}
PhysicalRegister(RegisterClassType Class, uint8_t Reg) : Reg(Reg), Class(Class.Val) { }
static const PhysicalRegister Invalid() {
return PhysicalRegister(InvalidClass, InvalidReg);
}
bool IsInvalid() {
return *this == Invalid();
}
};
static_assert(sizeof(PhysicalRegister) == 1);
class RegisterAllocationData {
public:
uint32_t SpillSlotCount {};
uint32_t MapCount {};
bool IsShared {false};
PhysicalRegister Map[0];
PhysicalRegister GetNodeRegister(uint32_t Node) const {
return Map[Node];
}
uint32_t SpillSlots() const { return SpillSlotCount; }
static size_t Size(uint32_t NodeCount) {
return sizeof(RegisterAllocationData) + NodeCount * sizeof(Map[0]);
}
};
struct RegisterAllocationDataDeleter {
void operator()(RegisterAllocationData* r) {
if (!r->IsShared) {
free(r);
}
}
};
}
+29
View File
@@ -8,6 +8,15 @@
#include <unordered_map>
#include <vector>
// Add macros which are missing in some versions of <elf.h>
#ifndef ELF32_ST_VISIBILITY
#define ELF32_ST_VISIBILITY(o) ((o) & 0x3)
#endif
#ifndef ELF64_ST_VISIBILITY
#define ELF64_ST_VISIBILITY(o) ((o) & 0x3)
#endif
namespace ELFLoader {
struct ELFSymbol {
uint64_t FileOffset;
@@ -40,6 +49,15 @@ public:
PhysicalMemorySize);
}
struct BRKInfo {
uint64_t Base;
uint64_t Size;
};
BRKInfo GetBRKInfo() const {
return {BRKBase, BRKSize};
}
// Data, Physical, Size
using MemoryWriter = std::function<void(void *, uint64_t, uint64_t)>;
void WriteLoadableSections(MemoryWriter Writer, uint64_t Offset = 0);
@@ -67,6 +85,14 @@ public:
void GetInitLocations(uint64_t GuestELFBase, std::vector<uint64_t> *Locations);
bool HasTLS() const { return TLSHeader._64 != nullptr; }
uint64_t GetTLSBase() const {
if (GetMode() == ELFMode::MODE_64BIT) {
return TLSHeader._64->p_vaddr;
}
else {
return TLSHeader._32->p_vaddr;
}
}
enum ELFMode {
MODE_32BIT,
@@ -122,6 +148,9 @@ private:
uint64_t MinPhysicalMemoryLocation{0};
uint64_t MaxPhysicalMemoryLocation{0};
uint64_t PhysicalMemorySize{0};
uint64_t BRKBase{};
uint64_t BRKSize{};
ProgramHeader InterpreterHeader{};
bool DynamicProgram{false};
std::string DynamicLinker;
+1
View File
@@ -1,3 +1,4 @@
#pragma once
#define GIT_SHORT_HASH "@GIT_SHORT_HASH@"
#define GIT_DESCRIBE_STRING "@GIT_DESCRIBE_STRING@"
-1
View File
@@ -7,7 +7,6 @@ This is the frontend application and tooling used for development and debugging
* imgui
* json-maker
* tiny-json
* boost interprocess (sadly)
* A C++17 compliant compiler (There are assumptions made about using Clang and LTO)
* clang-tidy if you want the code cleaned up
* cmake
+68
View File
@@ -0,0 +1,68 @@
#!/usr/bin/python3
import re
import sys
import subprocess
# Order this list from oldest to newest
# try not to list something newer than our minimum compiler supported version
BigCoreIDs = {
# ARM
tuple([0x41, 0xd07]): "cortex-a57",
tuple([0x41, 0xd08]): "cortex-a72",
tuple([0x41, 0xd09]): "cortex-a73",
tuple([0x41, 0xd0a]): "cortex-a75",
tuple([0x41, 0xd0b]): "cortex-a76",
tuple([0x41, 0xd0d]): "cortex-a77",
tuple([0x41, 0xd41]): "cortex-a78",
tuple([0x41, 0xd44]): "cortex-x1",
tuple([0x41, 0xd0c]): "neoverse-n1",
tuple([0x41, 0xd49]): "neoverse-n2",
## Nvidia
tuple([0x4e, 0x004]): "carmel", # Carmel
# Qualcomm
tuple([0x51, 0x800]): "cortex-a73", # Kryo 2xx Gold
tuple([0x51, 0x802]): "cortex-a75", # Kryo 3xx Gold
tuple([0x51, 0x804]): "cortex-a76", # Kryo 4xx Gold
}
LittleCoreIDs = {
# ARM
tuple([0x41, 0xd04]): "cortex-a35",
tuple([0x41, 0xd03]): "cortex-a53",
tuple([0x41, 0xd05]): "cortex-a55",
# Qualcomm
tuple([0x51, 0x801]): "cortex-a53", # Kryo 2xx Silver
tuple([0x51, 0x803]): "cortex-a55", # Kryo 3xx Silver
tuple([0x51, 0x805]): "cortex-a55", # Kryo 4xx/5xx Silver
}
# Args: </proc/cpuinfo file>
if (len(sys.argv) < 2):
sys.exit()
cpuinfo = []
with open(sys.argv[1]) as cpuinfo_file:
current_implementer = 0
current_part = 0
for line in cpuinfo_file:
line = line.strip()
if "CPU implementer" in line:
current_implementer = int(re.findall(r'0x[0-9A-F]+', line, re.I)[0], 16)
if "CPU part" in line:
current_part = int(re.findall(r'0x[0-9A-F]+', line, re.I)[0], 16)
cpuinfo += {tuple([current_implementer, current_part])}
largest_big = "native"
largest_little = "native"
for core in cpuinfo:
if BigCoreIDs.get(core):
largest_big = BigCoreIDs.get(core)
if LittleCoreIDs.get(core):
largest_little = LittleCoreIDs.get(core)
# We only want the big core output
print(largest_big)
# print(largest_little)
+6 -2
View File
@@ -58,8 +58,12 @@ else:
Process.wait()
ResultCode = Process.returncode
if (not test_name in expected_output or expected_output[test_name] != ResultCode):
if expected_output.get(test_name):
# expect zero by default
if (not test_name in expected_output):
expected_output[test_name] = 0
if (expected_output[test_name] != ResultCode):
if (test_name in expected_output):
print("test failed, expected is", expected_output[test_name], "but got", ResultCode)
else:
print("Test doesn't have expected output,", test_name)
+34 -7
View File
@@ -23,7 +23,7 @@ namespace FEX::ArgLoader {
#else
.choices({"irint", "irjit"})
#endif
.set_default("irint");
.set_default("irjit");
std::string BreakString = "Break";
std::string MultiBlockString = "Multiblock";
@@ -66,11 +66,11 @@ namespace FEX::ArgLoader {
.help("Number of physical hardware threads to tell the process we have")
.set_default(1);
CPUGroup.add_option("--smc-full-checks")
CPUGroup.add_option("--smc-checks")
.dest("SMCChecks")
.action("store_true")
.help("Checks code for modification before execution. Slow.")
.set_default(false);
.choices({"none", "mman", "full"})
.help("Checks code for modification before execution.\n\tnone: No checks\n\tmman: Invalidate on mmap, mprotect, munmap\n\tfull: Validate code before every run (slow)")
.set_default("mman");
CPUGroup.add_option("--unsafe-no-tso")
.dest("TSOEnabled")
@@ -114,6 +114,18 @@ namespace FEX::ArgLoader {
.help("Disables optimization passes for debugging")
.choices({"0"});
EmulationGroup.add_option("--aotir-capture")
.dest("AOTIRCapture")
.help("Captures IR and generates an AOTIR cache for the loaded executable and libs")
.action("store_true")
.set_default(false);
EmulationGroup.add_option("--aotir-load")
.dest("AOTIRLoad")
.help("Loads an AOTIR cache for the loaded executable")
.action("store_true")
.set_default(false);
Parser.add_option_group(EmulationGroup);
}
{
@@ -217,8 +229,13 @@ namespace FEX::ArgLoader {
}
if (Options.is_set_by_user("SMCChecks")) {
bool SMCChecks = Options.get("SMCChecks");
Set(FEXCore::Config::ConfigOption::CONFIG_SMC_CHECKS, std::to_string(SMCChecks));
auto SMCChecks = Options["SMCChecks"];
if (SMCChecks == "none")
Set(FEXCore::Config::ConfigOption::CONFIG_SMC_CHECKS, "0");
else if (SMCChecks == "mman")
Set(FEXCore::Config::ConfigOption::CONFIG_SMC_CHECKS, "1");
else if (SMCChecks == "full")
Set(FEXCore::Config::ConfigOption::CONFIG_SMC_CHECKS, "2");
}
if (Options.is_set_by_user("AbiLocalFlags")) {
bool AbiLocalFlags = Options.get("AbiLocalFlags");
@@ -249,6 +266,16 @@ namespace FEX::ArgLoader {
if (Options.is_set_by_user("O0")) {
Set(FEXCore::Config::ConfigOption::CONFIG_DEBUG_DISABLE_OPTIMIZATION_PASSES, std::to_string(true));
}
if (Options.is_set_by_user("AOTIRCapture")) {
bool AOTIRCapture = Options.get("AOTIRCapture");
Set(FEXCore::Config::ConfigOption::CONFIG_AOTIR_GENERATE, std::to_string(AOTIRCapture));
}
if (Options.is_set_by_user("AOTIRLoad")) {
bool AOTIRLoad = Options.get("AOTIRLoad");
Set(FEXCore::Config::ConfigOption::CONFIG_AOTIR_LOAD, std::to_string(AOTIRLoad));
}
}
{
+8 -1
View File
@@ -2,6 +2,7 @@
#include <FEXCore/Utils/LogManager.h>
#include <array>
#include <cassert>
#include <cstring>
#include <filesystem>
@@ -187,6 +188,8 @@ namespace FEX::Config {
{FEXCore::Config::ConfigOption::CONFIG_ABI_LOCAL_FLAGS, "ABILocalFlags"},
{FEXCore::Config::ConfigOption::CONFIG_ABI_NO_PF, "ABINoPF"},
{FEXCore::Config::ConfigOption::CONFIG_DEBUG_DISABLE_OPTIMIZATION_PASSES, "O0"},
{FEXCore::Config::ConfigOption::CONFIG_AOTIR_GENERATE, "AOTIRCapture"},
{FEXCore::Config::ConfigOption::CONFIG_AOTIR_LOAD, "AOTIRLoad"},
}};
@@ -234,6 +237,8 @@ namespace FEX::Config {
{"ABILocalFlags", FEXCore::Config::ConfigOption::CONFIG_ABI_LOCAL_FLAGS},
{"AbiNoPF", FEXCore::Config::ConfigOption::CONFIG_ABI_NO_PF},
{"O0", FEXCore::Config::ConfigOption::CONFIG_DEBUG_DISABLE_OPTIMIZATION_PASSES},
{"AOTIRCapture", FEXCore::Config::ConfigOption::CONFIG_AOTIR_GENERATE},
{"AOTIRLoad", FEXCore::Config::ConfigOption::CONFIG_AOTIR_LOAD},
}};
void OptionMapper::MapNameToOption(const char *ConfigName, const char *ConfigString) {
@@ -306,7 +311,7 @@ namespace FEX::Config {
}
};
static const std::array<std::pair<std::string, FEXCore::Config::ConfigOption>, 18> ConfigLookup = {{
static const std::array<std::pair<std::string, FEXCore::Config::ConfigOption>, 20> ConfigLookup = {{
{"FEX_CORE", FEXCore::Config::ConfigOption::CONFIG_DEFAULTCORE},
{"FEX_MAXINST", FEXCore::Config::ConfigOption::CONFIG_MAXBLOCKINST},
{"FEX_SINGLESTEP", FEXCore::Config::ConfigOption::CONFIG_SINGLESTEP},
@@ -325,6 +330,8 @@ namespace FEX::Config {
{"FEX_ABINOPF", FEXCore::Config::ConfigOption::CONFIG_ABI_NO_PF},
{"FEX_BREAK", FEXCore::Config::ConfigOption::CONFIG_BREAK_ON_FRONTEND},
{"FEX_DUMP_GPRS", FEXCore::Config::ConfigOption::CONFIG_DUMP_GPRS},
{"FEX_AOT_GENERATE", FEXCore::Config::ConfigOption::CONFIG_AOTIR_GENERATE},
{"FEX_AOT_LOAD", FEXCore::Config::ConfigOption::CONFIG_AOTIR_LOAD},
}};
std::optional<std::string_view> Value;
+3 -10
View File
@@ -30,7 +30,7 @@ namespace HostFactory {
explicit HostCore(FEXCore::Context::Context* CTX, FEXCore::Core::ThreadState *Thread, bool Fallback);
~HostCore() override;
std::string GetName() override { return "Host Core"; }
void* CompileCode(FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) override;
void* CompileCode(FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) override;
void *MapRegion(void *HostPtr, uint64_t VirtualGuestPtr, uint64_t Size) override {
return HostPtr;
@@ -42,10 +42,6 @@ namespace HostFactory {
bool HandleSIGSEGV(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext);
private:
FEXCore::Context::Context* CTX;
FEXCore::Core::ThreadState *ThreadState;
bool IsFallback{};
uint64_t ReturningStackLocation;
uint64_t ThreadStopHandlerAddress;
};
@@ -54,10 +50,7 @@ namespace HostFactory {
}
HostCore::HostCore(FEXCore::Context::Context* CTX, FEXCore::Core::ThreadState *Thread, bool Fallback)
: CodeGenerator(4096)
, CTX {CTX}
, ThreadState {Thread}
, IsFallback {Fallback} {
: CodeGenerator(4096) {
FEXCore::Context::RegisterHostSignalHandler(CTX, SIGSEGV,
[](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
auto InternalThread = reinterpret_cast<FEXCore::Core::InternalThreadState*>(Thread);
@@ -177,7 +170,7 @@ namespace HostFactory {
ready();
}
void* HostCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView<true> const *IR, FEXCore::Core::DebugData *DebugData) {
void* HostCore::CompileCode([[maybe_unused]] FEXCore::IR::IRListView const *IR, FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData) {
return nullptr;
}
+20 -93
View File
@@ -1,81 +1,21 @@
# Boost with minimum version of 1.50, not exact
find_package(Boost 1.50 REQUIRED)
enable_language(ASM_NASM)
if(NOT CMAKE_ASM_NASM_COMPILER_LOADED)
error("Failed to find NASM compatible assembler!")
endif()
add_subdirectory(LinuxSyscalls)
set(SYSCALL_SRCS
LinuxSyscalls/FileManagement.cpp
LinuxSyscalls/EmulatedFiles/EmulatedFiles.cpp
LinuxSyscalls/SignalDelegator.cpp
LinuxSyscalls/Syscalls.cpp
LinuxSyscalls/x32/Syscalls.cpp
LinuxSyscalls/x32/EPoll.cpp
LinuxSyscalls/x32/FD.cpp
LinuxSyscalls/x32/FS.cpp
LinuxSyscalls/x32/Info.cpp
LinuxSyscalls/x32/Memory.cpp
LinuxSyscalls/x32/NotImplemented.cpp
LinuxSyscalls/x32/Semaphore.cpp
LinuxSyscalls/x32/Sched.cpp
LinuxSyscalls/x32/Signals.cpp
LinuxSyscalls/x32/Socket.cpp
LinuxSyscalls/x32/Thread.cpp
LinuxSyscalls/x32/Time.cpp
LinuxSyscalls/x32/Timer.cpp
LinuxSyscalls/x64/EPoll.cpp
LinuxSyscalls/x64/FD.cpp
LinuxSyscalls/x64/IO.cpp
LinuxSyscalls/x64/Ioctl.cpp
LinuxSyscalls/x64/Info.cpp
LinuxSyscalls/x64/Memory.cpp
LinuxSyscalls/x64/Msg.cpp
LinuxSyscalls/x64/NotImplemented.cpp
LinuxSyscalls/x64/Semaphore.cpp
LinuxSyscalls/x64/Sched.cpp
LinuxSyscalls/x64/Signals.cpp
LinuxSyscalls/x64/Socket.cpp
LinuxSyscalls/x64/Thread.cpp
LinuxSyscalls/x64/Syscalls.cpp
LinuxSyscalls/x64/Time.cpp
LinuxSyscalls/Syscalls/EPoll.cpp
LinuxSyscalls/Syscalls/FD.cpp
LinuxSyscalls/Syscalls/FS.cpp
LinuxSyscalls/Syscalls/Info.cpp
LinuxSyscalls/Syscalls/IO.cpp
LinuxSyscalls/Syscalls/Key.cpp
LinuxSyscalls/Syscalls/Memory.cpp
LinuxSyscalls/Syscalls/Msg.cpp
LinuxSyscalls/Syscalls/Sched.cpp
LinuxSyscalls/Syscalls/Semaphore.cpp
LinuxSyscalls/Syscalls/SHM.cpp
LinuxSyscalls/Syscalls/Signals.cpp
LinuxSyscalls/Syscalls/Socket.cpp
LinuxSyscalls/Syscalls/Thread.cpp
LinuxSyscalls/Syscalls/Time.cpp
LinuxSyscalls/Syscalls/Timer.cpp
LinuxSyscalls/Syscalls/NotImplemented.cpp
LinuxSyscalls/Syscalls/Stubs.cpp
)
set(LIBS FEXCore Common CommonCore pthread)
set(NAME FEXLoader)
set(SRCS ELFLoader.cpp ${SYSCALL_SRCS})
set(LIBS FEXCore Common CommonCore)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
add_executable(FEXLoader ELFLoader.cpp)
target_include_directories(FEXLoader PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} ${LIBS})
target_link_libraries(FEXLoader ${LIBS} LinuxEmulation)
install(TARGETS ${NAME}
install(TARGETS FEXLoader
RUNTIME
DESTINATION bin
COMPONENT runtime)
set(FEX_INTERP FEXInterpreter)
install(CODE "
EXECUTE_PROCESS(COMMAND ln -f ${NAME} ${FEX_INTERP}
EXECUTE_PROCESS(COMMAND ln -f FEXLoader ${FEX_INTERP}
WORKING_DIRECTORY ${CMAKE_INSTALL_PREFIX}/bin/
)
")
@@ -115,39 +55,26 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
endif()
set(NAME TestHarness)
set(SRCS TestHarness.cpp)
add_executable(TestHarness TestHarness.cpp)
target_include_directories(TestHarness PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(TestHarness ${LIBS})
target_link_libraries(${NAME} ${LIBS})
add_executable(TestHarnessRunner TestHarnessRunner.cpp)
target_include_directories(TestHarnessRunner PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
set(NAME TestHarnessRunner)
set(SRCS TestHarnessRunner.cpp
${SYSCALL_SRCS})
target_link_libraries(TestHarnessRunner ${LIBS} LinuxEmulation)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
add_executable(UnitTestGenerator UnitTestGenerator.cpp)
target_include_directories(UnitTestGenerator PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} ${LIBS})
target_link_libraries(UnitTestGenerator ${LIBS})
set(NAME UnitTestGenerator)
set(SRCS UnitTestGenerator.cpp)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} ${LIBS})
set(NAME IRLoader)
set(SRCS
add_executable(IRLoader
IRLoader.cpp
IRLoader/Loader.cpp
${SYSCALL_SRCS})
)
target_include_directories(IRLoader PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
add_executable(${NAME} ${SRCS})
target_include_directories(${NAME} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/Source/)
target_link_libraries(${NAME} ${LIBS})
target_link_libraries(IRLoader ${LIBS} LinuxEmulation)
+45 -5
View File
@@ -16,6 +16,8 @@
#include <string>
#include <unistd.h>
#include <vector>
#include <fstream>
#include <filesystem>
namespace {
static bool SilentLog;
@@ -191,10 +193,10 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS_INTERPRETER, IsInterpreter ? "1" : "0");
FEXCore::Config::Set(FEXCore::Config::CONFIG_INTERPRETER_INSTALLED, IsInterpreterInstalled() ? "1" : "0");
FEXCore::Config::Value<uint8_t> CoreConfig{FEXCore::Config::CONFIG_DEFAULTCORE, 0};
FEXCore::Config::Value<uint64_t> BlockSizeConfig{FEXCore::Config::CONFIG_MAXBLOCKINST, 1};
FEXCore::Config::Value<uint8_t> CoreConfig{FEXCore::Config::CONFIG_DEFAULTCORE, 1};
FEXCore::Config::Value<uint64_t> BlockSizeConfig{FEXCore::Config::CONFIG_MAXBLOCKINST, 5000};
FEXCore::Config::Value<bool> SingleStepConfig{FEXCore::Config::CONFIG_SINGLESTEP, false};
FEXCore::Config::Value<bool> MultiblockConfig{FEXCore::Config::CONFIG_MULTIBLOCK, false};
FEXCore::Config::Value<bool> MultiblockConfig{FEXCore::Config::CONFIG_MULTIBLOCK, true};
FEXCore::Config::Value<bool> GdbServerConfig{FEXCore::Config::CONFIG_GDBSERVER, false};
FEXCore::Config::Value<std::string> LDPath{FEXCore::Config::CONFIG_ROOTFSPATH, ""};
FEXCore::Config::Value<std::string> ThunkLibsPath{FEXCore::Config::CONFIG_THUNKLIBSPATH, ""};
@@ -203,10 +205,11 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Config::Value<std::string> OutputLog{FEXCore::Config::CONFIG_OUTPUTLOG, "stderr"};
FEXCore::Config::Value<std::string> DumpIR{FEXCore::Config::CONFIG_DUMPIR, "no"};
FEXCore::Config::Value<bool> TSOEnabledConfig{FEXCore::Config::CONFIG_TSO_ENABLED, true};
FEXCore::Config::Value<bool> SMCChecksConfig{FEXCore::Config::CONFIG_SMC_CHECKS, false};
FEXCore::Config::Value<uint8_t> SMCChecksConfig{FEXCore::Config::CONFIG_SMC_CHECKS, FEXCore::Config::CONFIG_SMC_MMAN};
FEXCore::Config::Value<bool> ABILocalFlags{FEXCore::Config::CONFIG_ABI_LOCAL_FLAGS, false};
FEXCore::Config::Value<bool> AbiNoPF{FEXCore::Config::CONFIG_ABI_NO_PF, false};
FEXCore::Config::Value<bool> AOTIRCapture{FEXCore::Config::CONFIG_AOTIR_GENERATE, false};
FEXCore::Config::Value<bool> AOTIRLoad{FEXCore::Config::CONFIG_AOTIR_LOAD, false};
::SilentLog = SilentLog();
@@ -261,6 +264,8 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Config::SetConfig(CTX, FEXCore::Config::CONFIG_DUMPIR, DumpIR());
FEXCore::Config::Set(FEXCore::Config::CONFIG_APP_FILENAME, std::filesystem::canonical(Program));
FEXCore::Config::Set(FEXCore::Config::CONFIG_IS64BIT_MODE, Loader.Is64BitMode() ? "1" : "0");
FEXCore::Config::SetConfig(CTX, FEXCore::Config::CONFIG_AOTIR_GENERATE, AOTIRCapture());
FEXCore::Config::SetConfig(CTX, FEXCore::Config::CONFIG_AOTIR_LOAD, AOTIRLoad());
std::unique_ptr<FEX::HLE::SignalDelegator> SignalDelegation = std::make_unique<FEX::HLE::SignalDelegator>();
std::unique_ptr<FEX::HLE::SyscallHandler> SyscallHandler{
@@ -269,6 +274,8 @@ int main(int argc, char **argv, char **const envp) {
CTX,
SignalDelegation.get(),
&Loader)};
auto BRKInfo = Loader.GetBRKInfo();
SyscallHandler->DefaultProgramBreak(BRKInfo.Base, BRKInfo.Size);
FEXCore::Context::SetSignalDelegator(CTX, SignalDelegation.get());
FEXCore::Context::SetSyscallHandler(CTX, SyscallHandler.get());
@@ -286,8 +293,41 @@ int main(int argc, char **argv, char **const envp) {
});
}
if (AOTIRLoad() || AOTIRCapture()) {
LogMan::Msg::I("Warning: AOTIR is experimental, and might lead to crashes. Capture doesn't work with programs that fork.");
}
FEXCore::Context::SetAOTIRLoader(CTX, [](const std::string &fileid) -> std::unique_ptr<std::istream> {
auto filepath = std::filesystem::path(getenv("HOME")) / ".fex-emu" / "aotir" / fileid;
return std::make_unique<std::ifstream>(filepath, std::ios::in | std::ios::binary);
});
FEXCore::Context::RunUntilExit(CTX);
if (AOTIRCapture()) {
std::filesystem::create_directories(std::filesystem::path(getenv("HOME")) / ".fex-emu" / "aotir");
auto WroteCache = FEXCore::Context::WriteAOTIR(CTX, [](const std::string& fileid) -> std::unique_ptr<std::ostream> {
auto filepath = std::filesystem::path(getenv("HOME")) / ".fex-emu" / "aotir" / fileid;
auto AOTWrite = std::make_unique<std::ofstream>(filepath, std::ios::out | std::ios::binary);
if (*AOTWrite) {
std::filesystem::resize_file(filepath, 0);
AOTWrite->seekp(0);
LogMan::Msg::I("AOTIR: Storing %s", fileid.c_str());
} else {
LogMan::Msg::I("AOTIR: Failed to store %s", fileid.c_str());
}
return AOTWrite;
});
if (WroteCache) {
LogMan::Msg::I("AOTIR Cache Stored");
} else {
LogMan::Msg::E("AOTIR Cache Store Failed");
}
}
auto ProgramStatus = FEXCore::Context::GetProgramStatus(CTX);
SyscallHandler.reset();
+21 -6
View File
@@ -3,6 +3,7 @@
#include "Common/MathUtils.h"
#include <FEXCore/Core/CodeLoader.h>
#include <array>
#include <bitset>
#include <cassert>
#include <cstring>
@@ -327,6 +328,11 @@ namespace FEX::HarnessHelper {
ConfigStructBase BaseConfig;
};
#ifdef PAGE_SIZE
static_assert(PAGE_SIZE == 4096, "FEX only supports 4k pages");
#undef PAGE_SIZE
#endif
class HarnessCodeLoader final : public FEXCore::CodeLoader {
static constexpr uint32_t PAGE_SIZE = 4096;
@@ -495,13 +501,13 @@ public:
//AuxVariables.emplace_back(auxv_t{24, ~0ULL}); // AT_PLATFORM
// On x86 only allows userspace to check for monitor and fs/gs base writing in CPL3
//AuxVariables.emplace_back(auxv_t{26, 0}); // AT_HWCAP2
AuxVariables.emplace_back(auxv_t{32, 0ULL}); // sysinfo (vDSO)
AuxVariables.emplace_back(auxv_t{33, 0ULL}); // sysinfo (vDSO)
AuxVariables.emplace_back(auxv_t{32, 0}); // AT_SYSINFO - Entry point to syscall
AuxVariables.emplace_back(auxv_t{33, 0}); // AT_SYSINFO_EHDR - Address of the start of VDSO
}
else {
AuxVariables.emplace_back(auxv_t{4, 0x20}); // AT_PHENT
AuxVariables.emplace_back(auxv_t{32, 0ULL}); // sysinfo (vDSO)
AuxVariables.emplace_back(auxv_t{33, 0ULL}); // sysinfo (vDSO)
AuxVariables.emplace_back(auxv_t{32, 0}); // AT_SYSINFO - Entry point to syscall
AuxVariables.emplace_back(auxv_t{33, 0}); // AT_SYSINFO_EHDR - Address of the start of VDSO
}
AuxVariables.emplace_back(auxv_t{3, DB.GetElfBase()}); // Program header
@@ -554,8 +560,10 @@ public:
size_t ArgSize = Args[i].size();
// Set the pointer to this argument
ArgumentPointers[i] = ArgumentBackingBaseGuest + CurrentOffset;
// Copy the string in to the final location
memcpy(reinterpret_cast<void*>(ArgumentBackingBase + CurrentOffset), &Args[i].at(0), ArgSize);
if (ArgSize > 0) {
// Copy the string in to the final location
memcpy(reinterpret_cast<void*>(ArgumentBackingBase + CurrentOffset), &Args[i].at(0), ArgSize);
}
// Set the null terminator for the string
*reinterpret_cast<uint8_t*>(ArgumentBackingBase + CurrentOffset + ArgSize + 1) = 0;
@@ -733,9 +741,16 @@ public:
bool Is64BitMode() const { return File.GetMode() == ::ELFLoader::ELFContainer::MODE_64BIT; }
::ELFLoader::ELFContainer::BRKInfo GetBRKInfo() const {
auto Info = File.GetBRKInfo();
Info.Base += DB.GetElfBase();
return Info;
}
private:
::ELFLoader::ELFContainer File;
::ELFLoader::ELFSymbolDatabase DB;
std::vector<std::string> Args;
std::vector<std::string> EnvironmentVariables;
std::vector<char const*> LoaderArgs;
+1 -2
View File
@@ -81,7 +81,7 @@ class IRCodeLoader final : public FEXCore::CodeLoader {
uint64_t GetFinalRIP() override { return 0; }
virtual void AddIR(IRHandler Handler) override {
Handler(IR->GetEntryRIP(), IR);
Handler(IR->GetEntryRIP(), IR->GetIREmitter());
}
private:
@@ -124,7 +124,6 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Context::SetSignalDelegator(CTX, SignalDelegation.get());
FEX::IRLoader::InitializeStaticTables();
FEX::IRLoader::Loader Loader(Args[0], Args[1]);
int Return{};
+11 -489
View File
@@ -1,181 +1,12 @@
#include "IRLoader/Loader.h"
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <fstream>
namespace {
std::string ltrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_first_not_of(" \t\n\r")) != std::string::npos) {
String.erase(0, pos);
}
return String;
}
std::string rtrim(std::string String) {
size_t pos = std::string::npos;
if ((pos = String.find_last_not_of(" \t\n\r")) != std::string::npos) {
String.erase(String.begin() + pos + 1, String.end());
}
return String;
}
std::string trim(std::string String) {
return rtrim(ltrim(String));
}
std::string DecodeErrorToString(DecodeFailure Failure) {
switch (Failure) {
case DecodeFailure::DECODE_OKAY: return "Okay";
case DecodeFailure::DECODE_UNKNOWN_TYPE: return "Unknown Type";
case DecodeFailure::DECODE_INVALID: return "Invalid";
case DecodeFailure::DECODE_INVALIDCHAR: return "Invalid starting char";
case DecodeFailure::DECODE_INVALIDRANGE: return "Invalid integer range";
case DecodeFailure::DECODE_INVALIDREGISTERCLASS: return "Invalid register class";
case DecodeFailure::DECODE_UNKNOWN_SSA: return "Unknown SSA value";
case DecodeFailure::DECODE_INVALID_CONDFLAG: return "Invalid Conditional name";
case DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE: return "Invalid Memory Offset Type";
};
}
}
namespace FEX::IRLoader {
std::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
template<>
std::pair<DecodeFailure, uint8_t> Loader::DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, uint16_t> Loader::DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint16_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, uint32_t> Loader::DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint32_t Result = strtoul(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, uint64_t> Loader::DecodeValue(std::string &Arg) {
if (Arg.at(0) != '#') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
uint64_t Result = strtoull(&Arg.at(1), nullptr, 0);
if (errno == ERANGE) return {DecodeFailure::DECODE_INVALIDRANGE, 0};
return {DecodeFailure::DECODE_OKAY, Result};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType> Loader::DecodeValue(std::string &Arg) {
if (Arg == "GPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRClass};
}
else if (Arg == "FPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRClass};
}
else if (Arg == "GPRPair") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRPairClass};
}
else if (Arg == "Complex") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::ComplexClass};
}
return {DecodeFailure::DECODE_INVALIDREGISTERCLASS, FEXCore::IR::InvalidClass};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition> Loader::DecodeValue(std::string &Arg) {
uint8_t Size{}, Elements{1};
int NumArgs = sscanf(Arg.c_str(), "i%hhdv%hhd", &Size, &Elements);
if (NumArgs != 1 && NumArgs != 2) {
return {DecodeFailure::DECODE_INVALID, {}};
}
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::TypeDefinition::Create(Size / 8, Elements)};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::CondClassType> Loader::DecodeValue(std::string &Arg) {
std::array<std::string, 14> CondNames = {
"EQ",
"NEQ",
"UGE",
"ULT",
"MI",
"PL",
"VS",
"VC",
"UGT",
"ULE",
"SGE",
"SLT",
"SGT",
"SLE",
};
for (size_t i = 0; i < CondNames.size(); ++i) {
if (CondNames[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, CondClassType{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_CONDFLAG, {}};
}
template<>
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType> Loader::DecodeValue(std::string &Arg) {
std::array<std::string, 3> Names = {
"SXTX",
"UXTW",
"SXTW",
};
for (size_t i = 0; i < Names.size(); ++i) {
if (Names[i] == Arg) {
return {DecodeFailure::DECODE_OKAY, MemOffsetType{static_cast<uint8_t>(i)}};
}
}
return {DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE, {}};
}
template<>
std::pair<DecodeFailure, OrderedNode*> Loader::DecodeValue(std::string &Arg) {
if (Arg.at(0) != '%') return {DecodeFailure::DECODE_INVALIDCHAR, 0};
// Strip off the type qualifier from the ssa value
size_t ArgEnd = std::string::npos;
std::string SSAName = trim(Arg);
ArgEnd = SSAName.find_first_of(" ");
if (ArgEnd != std::string::npos) {
SSAName = SSAName.substr(0, ArgEnd);
}
// Forward declarations may make this not succed
auto Op = SSANameMapper.find(SSAName);
if (Op == SSANameMapper.end()) {
return {DecodeFailure::DECODE_UNKNOWN_SSA, nullptr};
}
return {DecodeFailure::DECODE_OKAY, Op->second};
}
Loader::Loader(std::string const &Filename, std::string const &ConfigFilename) {
Config.Init(ConfigFilename);
std::fstream fp(Filename, std::fstream::binary | std::fstream::in);
@@ -185,324 +16,15 @@ namespace FEX::IRLoader {
return;
}
std::string TmpLine;
while (!fp.eof()) {
std::getline(fp, TmpLine);
if (fp.eof()) {
break;
}
if (fp.fail()) {
LogMan::Msg::E("Failed to getline on line: %ld", Lines.size());
return;
}
Lines.emplace_back(TmpLine);
}
ParsedCode.reset(FEXCore::IR::Parse(&fp));
ResetWorkingList();
Loaded = Parse();
}
bool Loader::Parse() {
auto CheckPrintError = [&](LineDefinition &Def, DecodeFailure Failure) -> bool {
if (Failure != DecodeFailure::DECODE_OKAY) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Value Couldn't be decoded due to %s", DecodeErrorToString(Failure).c_str());
return false;
}
return true;
};
// String parse every line for our definitions
for (size_t i = 0; i < Lines.size(); ++i) {
std::string Line = Lines[i];
LineDefinition Def{};
CurrentDef = &Def;
Def.LineNumber = i;
Line = trim(Line);
// Skip empty lines
if (Line.empty()) {
continue;
}
if (Line[0] == ';') {
// This is a comment line
// Skip it
continue;
}
size_t CurrentPos{};
// Let's see if this node is assigning something first
if (Line[0] == '%') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of("=", CurrentPos)) != std::string::npos) {
Def.Definition = Line.substr(0, DefinitionEnd);
Def.Definition = trim(Def.Definition);
Def.HasDefinition = true;
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA declaration without assignment");
return false;
}
}
// Check if we are pulling in some IR from the IR Printer
// Prints (%ssa%d) at the start of lines without a definition
if (Line[0] == '(') {
size_t DefinitionEnd = std::string::npos;
if ((DefinitionEnd = Line.find_first_of(")", CurrentPos)) != std::string::npos) {
size_t SSAEnd = std::string::npos;
if ((SSAEnd = Line.find_last_of(" ", DefinitionEnd)) != std::string::npos) {
std::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
Type = trim(Type);
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
Def.Size = DefinitionSize.second;
}
Def.Definition = trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
CurrentPos = DefinitionEnd + 1;
}
else {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("SSA value with numbered SSA provided but no closing parentheses");
return false;
}
}
if (Def.HasDefinition) {
// Let's check if we have a size declared with this variable
size_t NameEnd = std::string::npos;
if ((NameEnd = Def.Definition.find_first_of(" ")) != std::string::npos) {
std::string Type = Def.Definition.substr(NameEnd + 1);
Type = trim(Type);
Def.Definition = trim(Def.Definition.substr(0, NameEnd));
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
if (!CheckPrintError(Def, DefinitionSize.first)) return false;
Def.Size = DefinitionSize.second;
}
if (Def.Definition == "%Invalid") {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Definition tried to define reserved %Invalid ssa node");
return false;
}
}
// Let's get the IR op
size_t OpNameEnd = std::string::npos;
std::string RemainingLine = trim(Line.substr(CurrentPos));
CurrentPos = 0;
if ((OpNameEnd = RemainingLine.find_first_of(" \t\n\r\0", CurrentPos)) != std::string::npos) {
Def.IROp = RemainingLine.substr(CurrentPos, OpNameEnd);
Def.IROp = trim(Def.IROp);
Def.HasArgs = true;
CurrentPos = OpNameEnd;
}
else {
if (RemainingLine.empty()) {
LogMan::Msg::E("Error on Line: %d", i);
LogMan::Msg::E("%s", Lines[i].c_str());
LogMan::Msg::E("Line without an IROp?");
return false;
}
Def.IROp = RemainingLine;
Def.HasArgs = false;
}
if (Def.HasArgs) {
RemainingLine = trim(RemainingLine.substr(CurrentPos));
CurrentPos = 0;
if (RemainingLine.empty()) {
// How did we get here?
Def.HasArgs = false;
}
else {
while (!RemainingLine.empty()) {
size_t ArgEnd = std::string::npos;
ArgEnd = RemainingLine.find_first_of(",");
std::string Arg = RemainingLine.substr(0, ArgEnd);
Arg = trim(Arg);
Def.Args.emplace_back(Arg);
RemainingLine.erase(0, ArgEnd+1); // +1 to ensure we go past the ','
if (ArgEnd == std::string::npos)
break;
}
}
}
Defs.emplace_back(Def);
}
// Ensure all of the ops are real ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
auto Op = NameToOpMap.find(Def.IROp);
if (Op == NameToOpMap.end()) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("IROp '%s' doesn't exist", Def.IROp.c_str());
return false;
}
Def.OpEnum = Op->second;
}
// Emit the header op
IRPair<IROp_IRHeader> IRHeader;
{
auto &Def = Defs[0];
CurrentDef = &Def;
if (Def.OpEnum != FEXCore::IR::IROps::OP_IRHEADER) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("First op needs to be IRHeader. Was '%s'", Def.IROp.c_str());
return false;
}
auto Entry = DecodeValue<uint64_t>(Def.Args[0]);
auto CodeBlockCount = DecodeValue<uint64_t>(Def.Args[2]);
if (!CheckPrintError(Def, Entry.first)) return false;
if (!CheckPrintError(Def, CodeBlockCount.first)) return false;
EntryRIP = Entry.second;
IRHeader = _IRHeader(InvalidNode, Entry.second, CodeBlockCount.second, false);
}
SetWriteCursor(nullptr); // isolate the header from everything following
// Initialize SSANameMapper with Invalid value
SSANameMapper["%Invalid"] = Invalid();
// Spin through the blocks and generate basic block ops
for(size_t i = 0; i < Defs.size(); ++i) {
auto &Def = Defs[i];
if (Def.OpEnum == FEXCore::IR::IROps::OP_CODEBLOCK) {
auto CodeBlock = _CodeBlock(InvalidNode, InvalidNode);
SSANameMapper[Def.Definition] = CodeBlock.Node;
Def.Node = CodeBlock.Node;
if (i == 1) {
// First code block is the entry block
// Link the header to the first block
IRHeader.first->Blocks = CodeBlock.Node->Wrapped(ListData.Begin());
}
}
}
SetWriteCursor(nullptr); // isolate the block headers too
// Spin through all the definitions and add the ops to the basic blocks
OrderedNode *CurrentBlock{};
FEXCore::IR::IROp_CodeBlock *CurrentBlockOp{};
for(size_t i = 1; i < Defs.size(); ++i) {
auto &Def = Defs[i];
CurrentDef = &Def;
if (Def.OpEnum == FEXCore::IR::IROps::OP_CODEBLOCK) {
auto DefTarget = SSANameMapper.find(Def.Definition);
if (CurrentBlock != nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("CodeBlock being used inside of already existing codeblock!");
return false;
}
CurrentBlock = DefTarget->second;
CurrentBlockOp = CurrentBlock->Op(Data.Begin())->CW<FEXCore::IR::IROp_CodeBlock>();
}
if (Def.OpEnum == FEXCore::IR::IROps::OP_ENDBLOCK) {
if (CurrentBlock == nullptr) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("EndBlock being used outside of a block!");
return false;
}
auto Adjust = DecodeValue<uint64_t>(Def.Args[0]);
if (!CheckPrintError(Def, Adjust.first)) return false;
Def.Node = _EndBlock(CurrentBlock);
CurrentBlockOp->Last = Def.Node->Wrapped(ListData.Begin());
CurrentBlock = nullptr;
CurrentBlockOp = nullptr;
}
if (Def.OpEnum == FEXCore::IR::IROps::OP_DUMMY) {
auto &PrevDef = Defs[i - 1];
if (PrevDef.OpEnum != FEXCore::IR::IROps::OP_CODEBLOCK) {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Dummy op must be first op in block");
return false;
}
Def.Node = _Dummy();
CurrentBlockOp->Begin = Def.Node->Wrapped(ListData.Begin());
SSANameMapper[Def.Definition] = Def.Node;
}
switch (Def.OpEnum) {
// Special handled
case FEXCore::IR::IROps::OP_IRHEADER:
case FEXCore::IR::IROps::OP_CODEBLOCK:
case FEXCore::IR::IROps::OP_ENDBLOCK:
case FEXCore::IR::IROps::OP_DUMMY:
break;
#define IROP_PARSER_SWITCH_HELPERS
#include <FEXCore/IR/IRDefines.inc>
default: {
LogMan::Msg::E("Error on Line: %d", Def.LineNumber);
LogMan::Msg::E("%s", Lines[Def.LineNumber].c_str());
LogMan::Msg::E("Unhandled Op enum '%s' in parser", Def.IROp.c_str());
return false;
break;
}
}
if (Def.HasDefinition) {
auto IROp = Def.Node->Op(Data.Begin());
if (Def.Size.Elements()) {
IROp->Size = Def.Size.Bytes() * Def.Size.Elements();
IROp->ElementSize = Def.Size.Bytes();
}
else {
IROp->Size = Def.Size.Bytes();
IROp->ElementSize = 0;
}
SSANameMapper[Def.Definition] = Def.Node;
}
}
std::stringstream out;
auto NewIR = ViewIR();
FEXCore::IR::Dump(&out, &NewIR, nullptr);
printf("IR:\n%s\n@@@@@\n", out.str().c_str());
return true;
}
void InitializeStaticTables() {
for (FEXCore::IR::IROps Op = FEXCore::IR::IROps::OP_DUMMY;
Op <= FEXCore::IR::IROps::OP_LAST;
Op = static_cast<FEXCore::IR::IROps>(static_cast<uint32_t>(Op) + 1)) {
NameToOpMap[FEXCore::IR::GetName(Op)] = Op;
if (ParsedCode) {
auto NewIR = ParsedCode->ViewIR();
EntryRIP = NewIR.GetHeader()->Entry;
std::stringstream out;
FEXCore::IR::Dump(&out, &NewIR, nullptr);
printf("IR:\n%s\n@@@@@\n", out.str().c_str());
}
}
}
+4 -83
View File
@@ -12,28 +12,13 @@
using namespace FEXCore::IR;
namespace {
enum class DecodeFailure {
DECODE_OKAY,
DECODE_UNKNOWN_TYPE,
DECODE_INVALID,
DECODE_INVALIDCHAR,
DECODE_INVALIDRANGE,
DECODE_INVALIDREGISTERCLASS,
DECODE_UNKNOWN_SSA,
DECODE_INVALID_CONDFLAG,
DECODE_INVALID_MEMOFFSETTYPE,
};
}
namespace FEX::IRLoader {
class Loader final : public FEXCore::IR::IREmitter {
class Loader final {
public:
Loader(std::string const &Filename, std::string const &ConfigFilename);
bool IsValid() const { return Loaded; }
bool IsValid() const { return ParsedCode != nullptr; }
IREmitter* GetIREmitter() { return ParsedCode.get(); }
uint64_t GetEntryRIP() const { return EntryRIP; }
bool CompareStates(FEXCore::Core::CPUState const* State) {
@@ -50,74 +35,10 @@ namespace FEX::IRLoader {
}
}
#define IROP_PARSER_ALLOCATE_HELPERS
#include <FEXCore/IR/IRDefines.inc>
private:
bool Parse();
uint64_t EntryRIP{};
std::vector<std::string> Lines;
bool Loaded{};
std::unique_ptr<IREmitter> ParsedCode;
std::unordered_map<std::string, OrderedNode*> SSANameMapper;
template<typename Type>
std::pair<DecodeFailure, Type> DecodeValue(std::string &Arg) {
return {DecodeFailure::DECODE_UNKNOWN_TYPE, {}};
}
template<>
std::pair<DecodeFailure, uint8_t>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, uint16_t>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, uint32_t>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, uint64_t>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, FEXCore::IR::CondClassType>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType>
DecodeValue(std::string &Arg);
template<>
std::pair<DecodeFailure, OrderedNode*>
DecodeValue(std::string &Arg);
struct LineDefinition {
size_t LineNumber;
bool HasDefinition{};
std::string Definition{};
FEXCore::IR::TypeDefinition Size{};
std::string IROp{};
FEXCore::IR::IROps OpEnum;
bool HasArgs{};
std::vector<std::string> Args;
OrderedNode *Node{};
};
std::vector<LineDefinition> Defs;
LineDefinition *CurrentDef{};
FEX::HarnessHelper::ConfigLoader Config;
};
void InitializeStaticTables();
}
+56
View File
@@ -0,0 +1,56 @@
add_library(LinuxEmulation STATIC
FileManagement.cpp
EmulatedFiles/EmulatedFiles.cpp
SignalDelegator.cpp
Syscalls.cpp
x32/Syscalls.cpp
x32/EPoll.cpp
x32/FD.cpp
x32/FS.cpp
x32/Info.cpp
x32/Memory.cpp
x32/NotImplemented.cpp
x32/Semaphore.cpp
x32/Sched.cpp
x32/Signals.cpp
x32/Socket.cpp
x32/Thread.cpp
x32/Time.cpp
x32/Timer.cpp
x64/EPoll.cpp
x64/FD.cpp
x64/IO.cpp
x64/Ioctl.cpp
x64/Info.cpp
x64/Memory.cpp
x64/Msg.cpp
x64/NotImplemented.cpp
x64/Semaphore.cpp
x64/Sched.cpp
x64/Signals.cpp
x64/Socket.cpp
x64/Thread.cpp
x64/Syscalls.cpp
x64/Time.cpp
Syscalls/EPoll.cpp
Syscalls/FD.cpp
Syscalls/FS.cpp
Syscalls/Info.cpp
Syscalls/IO.cpp
Syscalls/Key.cpp
Syscalls/Memory.cpp
Syscalls/Msg.cpp
Syscalls/Sched.cpp
Syscalls/Semaphore.cpp
Syscalls/SHM.cpp
Syscalls/Signals.cpp
Syscalls/Socket.cpp
Syscalls/Thread.cpp
Syscalls/Time.cpp
Syscalls/Timer.cpp
Syscalls/NotImplemented.cpp
Syscalls/Stubs.cpp
)
target_link_libraries(LinuxEmulation FEXCore pthread numa)
Loaded 100 of 198 files, more files were not shown because too many files have changed in this diff. Show more