Compare commits

...
549 Commits
Author SHA1 Message Date
Ryan Houdek 320c5f1847 Docs: Update for release FEX-2510 2025-10-08 17:03:06 -07:00
Ryan Houdek 3b88947cbe Merge pull request #4950 from lioncash/syscall
SyscallHandler: Shrink definition struct from 24 to 16 bytes
2025-10-08 14:03:57 -07:00
Ryan Houdek f05be0120e Merge pull request #4949 from lioncash/thread
ThreadManager: Make StatAlloc instance private
2025-10-08 13:44:43 -07:00
Lioncache 10c3b2fb5d SyscallHandler: Shrink definition struct from 24 to 16 bytes
Just a minor space saving, given how many syscall definitions we'll
have.
2025-10-08 16:40:20 -04:00
Lioncache 74ae89865c SHMStats: Add missing virtual destructor
This should be present for universally consistent behavior, regardless
of how the class is used.
2025-10-08 16:17:42 -04:00
Lioncache a2445904e7 ThreadManager: Make StatAlloc instance private
This doesn't need to be public.
2025-10-08 16:13:23 -04:00
Ryan Houdek bd5188e7bc Merge pull request #4948 from lioncash/fmt2
EnumUtils: Remove fmt include
2025-10-07 13:04:22 -07:00
Lioncache d89c119d9d EnumUtils: Remove fmt include
Forgot to remove this in the previous PR.
2025-10-07 12:54:32 -04:00
Ryan Houdek 802f3bed47 Merge pull request #4947 from lioncash/format
EnumUtils: Further simplify enum passthrough formatting
2025-10-07 06:47:56 -07:00
Ryan Houdek 9c606a278d Merge pull request #4946 from lioncash/sys
SysCalls: Remove unimplemented prototypes
2025-10-07 06:46:59 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Lioncache d94b9fae94 SourceCodeResolver: Pass string_view by value for GenerateMap()
Generally this should be passed as a value type unless there's a good
reason not to.
2025-10-07 01:27:15 -04:00
Lioncache 05fcaaa758 Syscalls: Remove unimplemented prototypes 2025-10-07 01:27:11 -04:00
Ryan Houdek 9626a64340 Merge pull request #4944 from lioncash/type
Addressing: Shave 8 bytes off AddressMode
2025-10-06 14:01:09 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Ryan Houdek bfda91ac16 Merge pull request #4943 from lioncash/table
X86Tables: Remove unused LateInitCopyTable()
2025-10-06 12:17:39 -07:00
Lioncache b7117b86ea X86Tables: Remove unused LateInitCopyTable() 2025-10-06 15:04:36 -04:00
LC a94059ad81 Merge pull request #4942 from Sonicadvance1/fix_map
ThreadManager: Fix mmap usage to not take guest space
2025-10-06 14:59:27 -04:00
Ryan Houdek 2cab87fbd0 ThreadManager: Fix mmap usage to not take guest space
SHM stats and CallRet stack was accidentally using system mmap which
means it would take VA space from the guest. Ensure it uses the
Allocator helpers so that doesn't happen.

Gives 32-bit guests 8-ish megabyte of VA space back.
2025-10-06 11:20:00 -07:00
Ryan Houdek 2017100ff3 Merge pull request #4941 from lioncash/unused
OpcodeDispatcher: Remove unimplemented function prototypes
2025-10-06 10:18:18 -07:00
Lioncache 72cb29d36b OpcodeDispatcher: Remove unimplemented function prototypes
Just cleans out the interface a little.
2025-10-06 13:08:07 -04:00
Ryan Houdek a2b0d594fb Merge pull request #4939 from lioncash/invalid
OpcodeDispatcher: Default alignment parameters for store helpers
2025-10-06 09:11:49 -07:00
Lioncache a16d4ff1f3 OpcodeDispatcher: Deduplicate in LEAOp/SMSWOp
We can shorten a few lines here by just storing the op addr value
to a local variable.
2025-10-06 11:13:00 -04:00
Lioncache 305b1ecf2d OpcodeDispatcher: Default alignment parameters for store helpers
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.

Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
2025-10-06 11:04:12 -04:00
Ryan Houdek 3cc0cae249 Merge pull request #4938 from lioncash/prctl
PrctlUtils: Move to include folder
2025-10-05 23:27:38 -07:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
LC 5f8cb0dc54 Merge pull request #4937 from Sonicadvance1/name_vma
Allocator: Name FEX's VMA regions for allocation
2025-10-05 22:40:58 -04:00
Ryan Houdek 45978474f3 Allocator: Name FEX's VMA regions for allocation
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.

Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.

This tracking will be the first step towards seeing where our memory
usage is going.
2025-10-05 19:30:59 -07:00
Ryan Houdek 75206c7a51 Merge pull request #4936 from lioncash/str
Common/StringConv: std::stoull -> std::strtoull for enum handler
2025-10-05 13:14:02 -07:00
Ryan Houdek 06ff0a45a7 Merge pull request #4935 from lioncash/move
IR: Remove Swap1/Swap2 ops
2025-10-05 11:57:35 -07:00
Ryan Houdek c948d532a7 Merge pull request #4934 from lioncash/select
OpcodeDispatcher: Move off implicit _Select
2025-10-05 11:54:59 -07:00
Ryan Houdek 3825483f79 Merge pull request #4933 from lioncash/reg
RedundantFlagCalculationElimination: Eliminate unnecessary vector copy
2025-10-05 11:53:56 -07:00
Lioncache 0b238d942e Common/StringConv: std::stoull -> std::strtoull for enum handler
Missed this when simplifying the conversion facilities, but
we should be using std::strtoull here, as per the programming
concerns.
2025-10-05 14:49:07 -04:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Lioncache c84801adb3 RedundantFlagCalc: Avoid vector copy in OptimizeParity()
Previously this was making a copy of the vector, when we only
need to read from it.
2025-10-05 12:17:45 -04:00
Lioncache 69f990c1f8 RedundantFlagCalc: Organize headers
Also remove incorrect comment that this pass isn't used.
2025-10-05 12:17:45 -04:00
Lioncache 8a83564678 RedundantFlagCalc: Remove unimplemented prototype
Just tidies the interface a little.
2025-10-05 12:17:45 -04:00
Lioncache 43d93f836f RedundantFlagCalc: Mark members as const where applicable
These don't affect member state.
2025-10-05 03:14:21 -04:00
Ryan Houdek 7c2c0f7fe5 Merge pull request #4932 from lioncash/python
json_ir_generator: Minor cleanup
2025-10-04 11:44:08 -07:00
Lioncache 3ff47cf677 json_ir_generator: Add missing error message
ExitError() expects a string argument. Unlikely error case to be hit,
but we should still make the error useful to the reader.
2025-10-04 11:00:05 -04:00
Lioncache 3165dce89e json_ir_generator: Turn IROpNameMap into a set
This is only used to track for duplicates, so we only really need
the key name, since the values are never used.
2025-10-04 11:00:00 -04:00
Lioncache d4fad0a3b0 json_ir_generator: Simplify some looping
We can use enumerate to collapse a few of these
2025-10-04 10:37:34 -04:00
Lioncache 15fcaa794f json_ir_generator: Simplify is_ssa_type
Same behavior, but a little more straightforward
2025-10-04 08:47:08 -04:00
Lioncache 6a1a505336 json_ir_generator: Add type annotations for top level vars
Makes the types a little more explicit and also makes member
suggestions in IDEs (vscode, etc) work a little better.
2025-10-04 08:45:08 -04:00
LC 85937a7bfb Merge pull request #4931 from Sonicadvance1/sse4a
Implement remaining SSE4a instructions
2025-10-04 08:22:58 -04:00
Ryan Houdek 3baa598b9e HostFeatures: Adds flag for SSE4a 2025-10-04 02:51:28 -07:00
Ryan Houdek 6f83b10af6 InstcountCI: Update 2025-10-04 02:46:27 -07:00
Ryan Houdek f52bcb49ab unittests/ASM: Adds new variable SSE4a tests 2025-10-04 02:46:14 -07:00
Ryan Houdek 9cef8ff7ce OpcodeDispatcher: Implement support for SSE4a variable extrq/insertq 2025-10-04 02:44:56 -07:00
Ryan Houdek a499ad6404 IR: Implement two new IR ops
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
2025-10-04 02:43:39 -07:00
Ryan Houdek 7c8767ab32 OpcodeDispatcher: Minor improvement to inserting constant to vector 2025-10-04 02:43:21 -07:00
Ryan Houdek 38a8559943 Merge pull request #4930 from lioncash/forward
IR: Move RegisterAllocationPass forward decl to JITClass
2025-10-03 21:22:17 -07:00
Lioncache 48c6acbc87 IR: Move RegisterAllocationPass forward decl to JITClass
This isn't actually used anywhere in the IR header, so we can
move it to where it's actually used.

Now the IR interface header doesn't have anything related to the
independent passes in it.
2025-10-04 00:08:58 -04:00
Ryan Houdek b5d93bc4a0 Merge pull request #4928 from lioncash/passthru
EnumUtils: Add define for default passthrough formatting
2025-10-03 19:45:41 -07:00
LC a9fffe771b Merge pull request #4929 from Sonicadvance1/sse4a_imm
OpcodeDispatcher: Implement support for imm variants of SSE4a extrq/insertq
2025-10-03 17:40:19 -04:00
Ryan Houdek 69f3ab14b9 InstcountCI: Update 2025-10-03 13:51:51 -07:00
Ryan Houdek a7b9a34ec9 unittests: Adds imm extrq/insertq tests 2025-10-03 13:34:32 -07:00
Ryan Houdek f45be1f59e OpcodeDispatcher: Implement support for imm extrq/insertq
Fairly straightforward to implement, but not exciting in their
performance.
2025-10-03 13:05:43 -07:00
Ryan Houdek 012c2ba851 IR: Fixes vector 64-bit binops
We were only supporting operating size of 256-bit and 128-bit. These
also support 64-bit which wasn't wired up.

SSE4a will want to use 64-bit.
2025-10-03 13:04:09 -07:00
Lioncache a0c2ce0068 EnumUtils: Add define for default passthrough formatting
Handles a normal case where printing an enum type as an integral value
is still desirable.

Mainly just a way to reduce boilerplate.
2025-10-03 14:21:19 -04:00
Ryan Houdek 86898035da Merge pull request #4926 from lioncash/enum
IR: Convert RegisterClassType to an enum class
2025-10-03 10:35:23 -07:00
Ryan Houdek 05eeffe0f7 Merge pull request #4927 from bylaws/ffds
Windows: Avoid redeclaring strtoll as unimplemented
2025-10-03 10:32:39 -07:00
Billy Laws 70cffce5eb Windows: Avoid redeclaring strtoll as unimplemented
Included as part of the musl code we pull in. Fixes ARM64EC/WOW64 FEX.
2025-10-03 16:30:41 +02:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Ryan Houdek c35b81d2c5 Merge pull request #4925 from lioncash/regclass
IREmitter: Add register class helpers
2025-10-02 22:14:41 -07:00
Lioncache fc63dcedb6 IR: Migrate to new helpers
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
2025-10-03 00:53:02 -04:00
Lioncache c879650c4b IR: Add GPR/FPR helpers
These will be utilized in follow up PRs
2025-10-02 23:53:50 -04:00
Ryan Houdek c4d019939b Merge pull request #4924 from lioncash/irdump 2025-10-02 10:36:12 -07:00
Lioncache 3841cc4aa5 IRDumper: Add remaining missing enum values
Now that the compiler can warn against these, we can fill the
remaining list in.
2025-10-02 12:18:36 -04:00
Lioncache 468f2f2d2d IRDumper: Convert remaining printers to lambda style invocations
This allows easily moving the default case out of the switch, making it
easier for compilers to warn about missing values in switches if any
enum members are added in the future but aren't added to the
formatters.
2025-10-02 12:18:28 -04:00
Lioncache be5db91c13 IRDumper: Add missing CheckTF BranchHint 2025-10-02 03:01:11 -04:00
Lioncache 6e7d0a520a IRDumper: Fix error message for FloatCompareOp
If ever printed this would give a misleading error that it was an
unrecognized OpSize type.
2025-10-02 02:57:29 -04:00
Lioncache e37996873a IRDumper: Add missing FenceType entry 2025-10-02 02:55:34 -04:00
Lioncache 25e851b570 IRDumper: Add missing SyscallFlags entry 2025-10-02 02:55:31 -04:00
Ryan Houdek d41b21a626 Merge pull request #4923 from lioncash/py5
IR: Remove TypeDefinition struct
2025-10-01 22:03:00 -07:00
Lioncache d65c54ba21 IR: Remove TypeDefinition struct
These aren't used at all.
2025-10-02 00:50:51 -04:00
Ryan Houdek a5d3bf0f48 Merge pull request #4922 from lioncash/py4
IR: Convert MemOffsetType to enum class
2025-10-01 21:41:54 -07:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Ryan Houdek 96145e8e3f Merge pull request #4921 from lioncash/py3
IR: Convert rounding modes to enum class
2025-10-01 21:03:15 -07:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Ryan Houdek e4a4a529dd Merge pull request #4920 from lioncash/py2 2025-10-01 20:28:23 -07:00
Lioncache 863d1e0007 IR: Convert FenceType to an enum class
Same thing minus an extra struct lingering around.
2025-10-01 23:13:08 -04:00
Ryan Houdek c64d6f2b36 Merge pull request #4919 from lioncash/python
IR: Convert CondClassType over to enum class
2025-10-01 19:28:50 -07:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Ryan Houdek 9db9be9612 Merge pull request #4918 from lioncash/bitcast
General: Migrate to std::bit_cast
2025-09-30 19:32:51 -07:00
Lioncache 6a19c77184 General: Migrate to std::bit_cast
We already have a few cases where we already use bit_cast in the
emitter, so we may as well move all of our temporary helper instances
over as well.
2025-09-30 22:20:10 -04:00
Ryan Houdek c17b9dea30 Merge pull request #4917 from lioncash/collapse
StringConv: Merge integral-handling facilities together
2025-09-30 18:58:13 -07:00
Lioncache 7564b6c69a StringConv: Merge integral-handling facilities together
Same behavior, but we condense all of the integral handling
into one place.
2025-09-30 21:43:57 -04:00
Ryan Houdek 5858c5041f Merge pull request #4916 from lioncash/type
StringConv: Fix std::enable_if usage
2025-09-30 17:13:43 -07:00
LC 18d0dec416 Merge pull request #4911 from Sonicadvance1/evmd_linux
Linux: Implement support for extended volatile metadata
2025-09-30 20:01:26 -04:00
LC 752046108c Merge pull request #4910 from Sonicadvance1/fix_wow64_evmd
Wow64: Ensure extended volatile metadata is removed
2025-09-30 20:01:12 -04:00
Lioncache 773ddbf5c4 StringConv: Fix std::enable_if usage
This neglected to use the ::type qualifier, so this candidate was
always being considered in overload resolution.
2025-09-30 19:53:45 -04:00
Ryan Houdek 88e90a6f3a Merge pull request #4915 from lioncash/emitter
Arm64Emitter: Cull unnecessary includes
2025-09-30 10:50:04 -07:00
Lioncache 84325d6b5c Arm64Emitter: Cull unnecessary includes
Also fixes an indirect include.
2025-09-30 09:32:25 -04:00
Ryan Houdek 1b8e03e1de Merge pull request #4914 from lioncash/header
fextl/string: Correct <filesystem> include to <functional>
2025-09-29 21:26:39 -07:00
Lioncache 569d7297f0 fextl/string: Correct <filesystem> include to <functional>
This was unintentionally putting all the filesystem utilities into headers implicitly.
2025-09-29 23:54:08 -04:00
Ryan Houdek d9520c9498 Merge pull request #4913 from lioncash/help
JITClass: Mark functions as static where applicable
2025-09-29 20:53:09 -07:00
Lioncache 6b61093037 JITClass: Mark functions as static where applicable
These don't depend on any class state.
2025-09-29 23:42:43 -04:00
LC 74b09d5581 Merge pull request #4912 from Sonicadvance1/avx_32bit_overflow
FEXCore: Implement support for AVX gathers with overflow
2025-09-29 21:10:42 -04:00
Ryan Houdek 08da8f6f5c unittests/ASM: Adds gather overflow tests 2025-09-29 16:49:34 -07:00
Ryan Houdek 7f8fffbb24 FEXCore: Implement support for AVX gathers with overflow
Falls down the emulated path so we can zero extend the address
calculation. No way for SVE gathers to use 32-bit addressing.
2025-09-29 16:06:02 -07:00
Ryan Houdek b08e5f9821 Linux: Implement support for extended volatile metadata
Allows applications running entirely under Linux to use the same
extended volatile metadata as Windows.

For example `FEX_EXTENDEDVOLATILEMETADATA=iw4sp.exe\;0xe9da0-0xe9ec7`
this configuration works for both wow64 and Linux to disable the TSO
emulation on the memcpy routine in that game that consumes around 80% of
CPU time in TSO emulation.

Works with Linux native games as well of course.
2025-09-29 14:36:51 -07:00
Ryan Houdek 1758648f56 Common: Move ApplyFEXExtendedVolatileMetadata to Common
It is going to be used by Linux.

Also adds support for file offset handling, which is required on Linux
but not on win32.
2025-09-29 14:36:50 -07:00
Ryan Houdek 418ce0aed9 Wow64: Ensure extended volatile metadata is removed
Was missing this, matches the arm64ec side.
2025-09-29 14:13:44 -07:00
LC 4e926076c9 Merge pull request #4909 from zeyi2/main
Docs: Fix wrong link in `Readme_CN.md`
2025-09-29 10:30:38 -04:00
mtx 3f0aaeba68 Docs: Fix wrong link in Readme_CN.md 2025-09-29 16:53:31 +08:00
LC 6ecec77a74 Merge pull request #4908 from FEX-Emu/bylaws-patch-1
Change required Python version from 3.10 to 3.9
2025-09-28 13:03:45 -04:00
Billy Laws 79917dc9bb Change required Python version from 3.10 to 3.9
Debian bullseye is a supported build environment which only provides 3.9
2025-09-28 16:34:50 +02:00
Ryan Houdek b0b60fb761 Merge pull request #4907 from neobrain/feature_nix_update
Build: Update toolchain for WoA builds
2025-09-25 13:43:23 -07:00
Tony Wasserka 7463152f50 Build: Update toolchain for WoA builds 2025-09-25 22:17:44 +02:00
Ryan Houdek cbd8217c60 Merge pull request #4904 from bylaws/tesrgr
TestHarnessRunner: Don't attempt to build on MinGW
2025-09-24 13:16:47 -07:00
Ryan Houdek 4de1762c21 Merge pull request #4903 from bylaws/fee
vixl: Update submodule
2025-09-24 12:41:00 -07:00
Billy Laws 73962402df TestHarnessRunner: Don't attempt to build on MinGW 2025-09-24 20:28:33 +01:00
Billy Laws d90ccceec7 vixl: Update submodule 2025-09-24 20:26:28 +01:00
Ryan Houdek 0c5ef1ce1d Merge pull request #4902 from lioncash/fmt
Update fmt to 12.0.0
2025-09-23 12:13:34 -07:00
Ryan Houdek e5680e031d Merge pull request #4835 from pmatos/fix/classify
Fix quiet and signalling nan propagation
2025-09-23 11:17:37 -07:00
Paulo Matos eee0cffa57 instcountci: Fix quiet and signalling nan propagation 2025-09-23 13:48:45 +02:00
Paulo Matos c480b0ba41 asm_tests: Fix quiet and signalling nan propagation 2025-09-23 13:48:45 +02:00
Paulo Matos e7a47a647c Fix quiet and signalling nan propagation
This adds a new mode X87StrictReducedPrecision.
The strict reduced precision is like the reduced precision but adds extra checks,
like the currently implemented nan and snan propagations.

Fix for __builtin_issignaling() test of SPEC2017 classify test.
2025-09-23 13:48:45 +02:00
Paulo Matos 500c82bbf8 Do not format NASM include files 2025-09-23 13:48:45 +02:00
Lioncache 008528af91 Update fmt to 12.0.0
Moves us over to the next major version release.
Changelog here: https://github.com/fmtlib/fmt/releases/tag/12.0.0
2025-09-23 09:50:11 +02:00
Tony Wasserka 13199ad203 Merge pull request #4898 from Sonicadvance1/a_snake_a_snake
CMake: Update our python minspec to 3.10
2025-09-23 09:29:33 +02:00
LC 533dd1a54b Merge pull request #4897 from Sonicadvance1/fix_telem
OpcodeDispatcher: Fixes telemetry on legacy segment read
2025-09-23 08:11:10 +02:00
Ryan Houdek b6f2e10416 CMake: Update our python minspec to 3.10 2025-09-22 12:31:29 -07:00
Ryan Houdek c58db36a7c OpcodeDispatcher: Fixes telemetry on legacy segment read 2025-09-22 11:57:34 -07:00
LC 4593099882 Merge pull request #4896 from pmatos/fix/x87IsZero
Remove unused IsZero; NFC
2025-09-22 15:18:05 +02:00
Paulo Matos 21e29af48b Remove unused IsZero; NFC 2025-09-22 14:22:14 +02:00
LC 892e44a900 Merge pull request #4895 from Sonicadvance1/fix_fxsave
OpcodeDispatcher: Fixes fxsave x87 register storing
2025-09-22 13:27:24 +02:00
Ryan Houdek bca29c2549 unittests/ASM: Adds fxsave/fxrstor tests 2025-09-19 14:06:27 -07:00
Ryan Houdek 330e7f628c InstcountCI: Update 2025-09-19 14:05:29 -07:00
Ryan Houdek efe401c7a4 OpcodeDispatcher: Fixes fxsave x87 register storing
These get stored based on the current rotation of TOP. So we need to be
a bit careful with how we do this storing. A smidge of overhead, but
nothing unexpected.
2025-09-19 14:04:32 -07:00
Ryan Houdek b3d88c043d Merge pull request #4894 from pmatos/feature/improve-TestFiltering
Support configuration-specific known failures in asm_tests
2025-09-19 11:52:21 -07:00
Paulo Matos edd36dac91 Support configuration-specific known failures in asm_tests
This allows Known_Failures files to contain either:
- Partial paths (existing): Test_X87_F64/SomeTest.asm (affects all configs)
- Full paths (new): jit_1/Test_64Bit_X87_F64/SomeTest.asm (specific config)

For example: Mark only the jit_1 (-n 1) configuration as failing:
jit_1/Test_64Bit_X87_F64/Memcopy_int_F64.asm
2025-09-19 16:58:26 +02:00
Tony Wasserka 6de5fb2885 Merge pull request #4891 from Sonicadvance1/fix_sigill_report
FEX: Fixes SIGILL reporting
2025-09-18 23:35:25 +02:00
Ryan Houdek 3b2ebafd83 unittests/FEXLinuxTests: Adds simple sigill test
Just checks that RIP, siginfo rip, REG_ERR and REG_TRAPNO are correct.
2025-09-18 11:26:15 -07:00
Ryan Houdek dd4f508b77 FEX: Fixes SIGILL reporting
Two bug fixes here.
- We weren't reporting a correctl ErrorRegister. Needs to contain
  PF_INSTR at least.
- We weren't setting siginfo->si_addr in the case of `FaultToTopAndGeneratedException`

Fixes Mafia 2 (Classic)
2025-09-18 11:26:15 -07:00
Tony Wasserka 3058a825b7 Merge pull request #4889 from Sonicadvance1/fix_x87_state_saverestore
SignalDelegator: Correctly save and restore x87 state.
2025-09-17 11:53:42 +02:00
LC bdbeef3b9a Merge pull request #4884 from Sonicadvance1/fex_decodedinst_size
Frontend: Improve DecodeInst size from 128 bytes to 80
2025-09-16 10:05:33 +02:00
LC 8ead4a3c34 Merge pull request #4878 from Sonicadvance1/modrm_oob
unittests/ASM: Implement modrm OOB tests
2025-09-16 10:02:38 +02:00
LC 2d53a9b0dc Merge pull request #4885 from Sonicadvance1/remove_log_about_segments
OpcodeDispatcher: Removes a log about unknown segments
2025-09-16 10:01:27 +02:00
LC 8f47b46219 Merge pull request #4887 from Sonicadvance1/finish_v4l2
IoctlEmulation: Finish v4l2 implementation
2025-09-16 10:00:56 +02:00
Ryan Houdek 4f0d35bae0 unittests/FEXLinuxTests: Adds test for x87 save/restore testing 2025-09-15 13:20:58 -07:00
Ryan Houdek de6d907e27 SignalDelegator: Correctly save and restore x87 state.
We need to read the current x87 TOP location and rotate the values when
saving and restoring the context state on Linux. Wasn't visible in the
ASM tests since we read the FEX's CPU state directly there.
2025-09-15 13:20:43 -07:00
Ryan Houdek c9fc347de1 Merge pull request #4883 from Sonicadvance1/fexfexi
FEX: Rebrand FEXInterpreter as FEX
2025-09-15 13:19:58 -07:00
Ryan Houdek 85e67f92c5 FEX: Rebrand FEXInterpreter as FEX
Still creates a copy of FEXInterpreter from FEX for downstream projects
to have some time to get off the old name. Creating a symlink is kind of
a pain in cmake so just doing an install copy is easy.
2025-09-15 13:06:48 -07:00
Ryan Houdek 034a27ce5a Merge pull request #4888 from lioncash/fclass
IR: Remove unnecessary friend class declarations from OrderedNode
2025-09-15 08:52:47 -07:00
Lioncache 056abb5901 IR: Add missing includes
Fixes some indirect includes while we're at it.
2025-09-15 11:47:39 +02:00
Lioncache af28406dbc IR: Remove unnecessary friend class declarations from OrderedNode
These classes don't exist anymore, so we can get rid of these declarations.
2025-09-15 11:37:37 +02:00
Ryan Houdek 376d6ba72c Merge pull request #4886 from Sonicadvance1/futimesat
FEXLinuxTests: Adds a small sleep for consistency
2025-09-13 13:24:39 -07:00
Ryan Houdek 736a73453f FEXLinuxTests: Adds a small sleep for consistency
Sleep a second to make this test more consistent. Sometimes the
filesystem times are slightly different than wallclock times.
2025-09-13 12:51:01 -07:00
Ryan Houdek 6ced309c5c IoctlEmulation: Finish v4l2 implementation 2025-09-12 19:59:56 -07:00
Ryan Houdek a30593efd6 OpcodeDispatcher: Removes a log about unknown segments
This behaves like `UnimplementedOp`, giving a SIGILL if it actually gets
hit.
2025-09-12 17:34:16 -07:00
Ryan Houdek 3fa400bc55 Frontend: Improve DecodeInst size from 128 bytes to 80
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.

If we packed this would get to a further 62 bytes but probably not worth
it.
2025-09-12 16:07:01 -07:00
Ryan Houdek 90ea16325a Frontend: Move a couple of members from DecodeInst
LastEscapePrefix doesn't need to be in DecodeInst, and we can have a
couple of flags for ForceTSO/DecodedSIB/DecodedModRM.
2025-09-12 15:37:18 -07:00
LC be84a4332a Merge pull request #4875 from Sonicadvance1/remove_arguments
FEXInterpreter: Remove FEXLoader
2025-09-12 17:14:25 -04:00
LC cfeba859ac Merge pull request #4882 from Sonicadvance1/enable_tests_host
unittests/ASM: Enables some previously disabled tests
2025-09-12 17:12:57 -04:00
LC 38ed7c9bb1 Merge pull request #4881 from Sonicadvance1/fix_xlat
unittests/ASM: Fixes xlat test so it can be run
2025-09-12 17:11:04 -04:00
Ryan Houdek 315ee90837 unittests/ASM: Enables some previously disabled tests
The new Zen system in CI has a new enough kernel for these tests.

Just need to modify the tests to be safe for our signal handler.
2025-09-12 13:20:56 -07:00
Ryan Houdek 278574ce91 unittests/ASM: Fixes xlat test so it can be run
Just need to make sure to save and restore gs/fs in the test, otherwise
our signal handler gets /very/ upset.
2025-09-12 13:08:52 -07:00
Ryan Houdek 6c225d5469 Mark VEX as known failing in vixl sim 2025-09-12 12:42:18 -07:00
Ryan Houdek d6a466ac0b OpcodeDispatcher: Fixes 16-bit movbe to memory 2025-09-12 12:42:18 -07:00
Ryan Houdek 58526d4faa unittests/ASM: x87 table. reduced precision 2025-09-12 12:11:34 -07:00
Ryan Houdek 3c90a82e75 unittests/ASM: x87 table 2025-09-12 12:10:21 -07:00
Ryan Houdek 06a27deebb unittests/ASM: VEX Group table 2025-09-12 11:00:55 -07:00
Ryan Houdek f607f877c9 unittests/ASM: VEX table 2025-09-12 10:54:21 -07:00
Ryan Houdek 7273041314 FEXInterpreter: Remove FEXLoader
Doesn't /quite/ remove the ArgumentLoader because it is intertwined with
LinuxEmulation in an annoying way that will take another step to remove.
2025-09-12 10:24:58 -07:00
Ryan Houdek 997d041a84 Merge pull request #4880 from lioncash/context
Interface/Context: Remove unnecessary headers/forward declarations
2025-09-12 09:36:23 -07:00
LC 226233ddce Merge pull request #4874 from Sonicadvance1/testharnessrunner_args
TestHarnessRunner: Removes ArgumentLoader
2025-09-12 11:22:11 -04:00
LC 6c8edcbea5 Merge pull request #4869 from Sonicadvance1/tests_env_loader
unittests: Remove FEXLoader usage from unittests
2025-09-12 11:10:38 -04:00
Lioncache e0f27bb855 Interface/Context: Remove unnecessary headers/forward declarations 2025-09-12 04:13:26 -04:00
Ryan Houdek 4a14003f34 Merge pull request #4879 from lioncash/vdso
VDSO_Emulation: Remove duplicate span include
2025-09-12 00:56:33 -07:00
Lioncache 216a37348b VDSO_Emulation: Remove duplicate span include
Just an include that slipped through.
2025-09-12 03:43:12 -04:00
Tony Wasserka ef7b2a9d4f Merge pull request #4876 from Sonicadvance1/missed_format
OpcodeDispatcher: Fix misaligned clang-format
2025-09-12 08:47:32 +02:00
Tony Wasserka ffb1cf4c7c Merge pull request #4868 from Sonicadvance1/remove_cpack
CMake: Remove CPack, completely unused
2025-09-12 08:40:18 +02:00
Tony Wasserka 964165eabe Merge pull request #4871 from Sonicadvance1/move_fexloader
FEXLoader: Move to FEXInterpreter
2025-09-12 08:27:14 +02:00
Ryan Houdek dca2d74eb6 unittests/ASM: H0F3A table 2025-09-11 19:08:09 -07:00
Ryan Houdek e136ff52ab unittests/ASM: H0F38 table 2025-09-11 18:56:19 -07:00
Ryan Houdek 0df4944a7b unittests/ASM: DDD table 2025-09-11 18:56:19 -07:00
Ryan Houdek 2bb296df3b unittests/ASM: Primary Group table 2025-09-11 18:26:58 -07:00
Ryan Houdek 52380a3927 unittests/ASM: Secondary modrm table 2025-09-11 18:14:09 -07:00
Ryan Houdek 732f725a24 unittests/ASM: Secondary Group table 2025-09-11 18:14:09 -07:00
Ryan Houdek 8887b16299 unittests/ASM: Secondary OpSize table 2025-09-11 17:51:00 -07:00
Ryan Houdek 75c8b08d9b unittests/ASM: Secondary REPNE table 2025-09-11 17:32:14 -07:00
Ryan Houdek 9741dc214d unittests/ASM: Secondary REP table 2025-09-11 17:25:06 -07:00
Ryan Houdek 8b4b80c725 unittests/ASM: Secondary table 2025-09-11 16:54:07 -07:00
Ryan Houdek e3b6cccca1 unittests/ASM: Primary table 2025-09-11 16:53:58 -07:00
Ryan Houdek d28ef1859d unittests: modrm_oob include helper 2025-09-11 16:53:42 -07:00
Ryan Houdek 74b59f458d OpcodeDispatcher: Fix misaligned clang-format 2025-09-11 14:27:43 -07:00
Ryan Houdek 4ae04c2a9a TestHarnessRunner: Removes ArgumentLoader
Only use environment variables for setting arguments here.
2025-09-11 13:45:09 -07:00
LC 7bd0789402 Merge pull request #4870 from Sonicadvance1/remove_argloader_fexbash
FEXBash: Remove ArgLoader usage
2025-09-11 16:17:34 -04:00
LC eb9fb9e834 Merge pull request #4872 from Sonicadvance1/fexserver_remove_argloader
FEXServer: Remove usage of ArgLoader
2025-09-11 16:12:49 -04:00
LC c0ef0a7503 Merge pull request #4873 from Sonicadvance1/fexrootfsfetch_remove_argloader
FEXRootFSFetcher: Removes argloader usage
2025-09-11 16:12:33 -04:00
Ryan Houdek 61719115e5 Merge pull request #4817 from neobrain/refactor_code_cache_new_interfaces
CodeCache: Introduce new interfaces
2025-09-11 12:59:33 -07:00
Ryan Houdek 4a8cbe2ab8 FEXBash: Remove ArgLoader usage
ArgLoader is going to get removed. So delete this menial usage of it.
2025-09-11 12:21:35 -07:00
Ryan Houdek c966f44189 FEXRootFSFetcher: Removes argloader usage
Another case that ArgLoader was misused.
2025-09-11 12:20:28 -07:00
Ryan Houdek 193ef8b232 FEXServer: Remove usage of ArgLoader
This is entirely unused as FEXServer only ever reads the environment and
other configs.

There was /technically/ a weird conflict with FEXServer arguments but
that was an accident if it ever happened.
2025-09-11 12:16:47 -07:00
Ryan Houdek ad70a60d05 FEXLoader: Move to FEXInterpreter
NFC

FEXInterpreter is the supported path, rename things to reflect that.
2025-09-11 12:09:23 -07:00
Ryan Houdek 9dfd5b5526 unittests: Remove FEXLoader usage from unittests
Use FEXInterpreter directly and pass FEX options through environment
variables.
2025-09-11 12:03:56 -07:00
Ryan Houdek 09ec374eec CMake: Remove CPack, completely unused 2025-09-11 11:07:05 -07:00
Ryan Houdek 8c1e9eda12 Merge pull request #4867 from neobrain/refactor_cmake_tests
CMake: Merge BUILD_TESTS into BUILD_TESTING
2025-09-11 10:14:26 -07:00
Ryan Houdek 0eedd55dfd Merge pull request #4866 from neobrain/refactor_reformat
Update code formatting
2025-09-11 10:13:42 -07:00
Ryan Houdek 58c86d16f3 Merge pull request #4865 from neobrain/refactor_cmake_component
CMake: Enable runtime-only installation
2025-09-11 10:13:15 -07:00
Ryan Houdek 4a9170e981 Merge pull request #4837 from bylaws/x87opt
X87: Better cache intermediate results in the X87 slowpath
2025-09-11 10:12:36 -07:00
Tony Wasserka b5b8ff01d5 CodeCache: Move LibraryJITNaming and GDBSymbols checks to Core 2025-09-11 17:03:50 +02:00
Tony Wasserka 6c86ffdeb7 Core: Remove redundant nullptr check 2025-09-11 17:03:50 +02:00
Tony Wasserka 70a1d92d9f CodeCache: Drop ComputeCodeMapId member function 2025-09-11 17:03:50 +02:00
Tony Wasserka 0749477eb9 CodeCache: Introduce revamped interfaces 2025-09-11 17:03:50 +02:00
Tony Wasserka fb54341a1f CMake: Merge BUILD_TESTS into BUILD_TESTING
BUILD_TESTING is provided by CMake, so the separate toggle is just redundant.
2025-09-11 14:08:47 +02:00
Tony Wasserka db601d333b Core: Rename and move AOTIR.cpp and AOTIR.h
The new names better reflect the contents after recent/upcoming API changes.
2025-09-11 10:49:07 +02:00
Tony Wasserka af1c2cccac Rename AOTIR.cpp to CodeCache.cpp 2025-09-11 10:49:07 +02:00
Tony Wasserka bf1b76dcd2 LinuxSyscalls: Clean up MappedResource insertion
None of the call sites needed this to be a template, and the previous function
produced unreadable error messages.
2025-09-11 10:49:07 +02:00
Tony Wasserka 48787ab460 Context: Drop unneeded header includes 2025-09-11 10:49:07 +02:00
Tony Wasserka 7d4bf84304 FHU: Add GetFilename overload for std::string_view on WoA 2025-09-11 10:49:07 +02:00
Tony Wasserka 5cf5d3e3a2 docs: Drop reference to removed CreateIRCopy function 2025-09-11 10:49:07 +02:00
Tony Wasserka 8ddb229447 Add previous commit to git blame ignore file 2025-09-11 10:40:43 +02:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Tony Wasserka 7fbdd0607f CMake: Add targets missing from Runtime component 2025-09-11 10:31:06 +02:00
Tony Wasserka 0c0a1d8f12 CMake: Use consistent component naming 2025-09-11 10:31:06 +02:00
LC 30dd9ed267 Merge pull request #4863 from Sonicadvance1/more_cpuid
CPUID: Adds new CPU product names that exist
2025-09-10 23:09:33 -04:00
LC 5f0a1d55e5 Merge pull request #4864 from Sonicadvance1/fix_x87_reduced_load
FEXCore: Fixes OOB access on x87 reduced precision loads
2025-09-10 23:08:50 -04:00
Ryan Houdek 17bef26708 InstcountCI: Update 2025-09-10 19:24:50 -07:00
Ryan Houdek 58b5620f97 FEXCore: Fixes OOB access on x87 reduced precision loads
x87 80-bit loads, both BCD and regular tword were loading 128-bits of
data when reduced precision was enabled. This was an oversight from the
previous fix a while ago.

Adds a specific reduced precision test for this, and updates the current
test to ensure stores are still tested as well.
2025-09-10 19:19:55 -07:00
Ryan Houdek 327b62ea45 CPUID: Adds new CPU product names that exist 2025-09-10 11:51:04 -07:00
Ryan Houdek 91c6d693df Merge pull request #4860 from bylaws/susprip
JIT: Check for suspend interrupts at back-edges
2025-09-10 10:29:37 -07:00
Ryan Houdek aa3df3914b Merge pull request #4862 from lioncash/ir
IREmitter: Remove friend class declarations
2025-09-10 10:29:16 -07:00
Lioncache ad3a024b69 IREmitter: Remove friend class declarations
These no longer need specific access to internals.
2025-09-10 13:12:58 -04:00
Ryan Houdek b9ec94264e Merge pull request #4830 from neobrain/feature_drop_tso_automigration
Core: Drop support for TSO auto migration
2025-09-10 09:00:24 -07:00
Tony Wasserka fee9e91c4f Core: Drop support for TSO auto migration 2025-09-10 15:36:23 +02:00
Tony Wasserka 1dc560d45a InstCountCI: Explicitly disable TSO by default 2025-09-10 15:36:23 +02:00
Tony Wasserka 258e2b80f4 unittests/ASM: Explicitly disable TSO for vixl simulator runs
The simulator does not fully support this configuration and only works by
accident currently because of TSO auto-migration.
2025-09-10 15:36:23 +02:00
Ryan Houdek dcb889704d Merge pull request #4861 from bylaws/asrg
Dispatcher: Fix FABI_F32_I16_F80_PTR argument size
2025-09-09 20:14:06 -07:00
Billy Laws 93096c27f9 Dispatcher: Fix FABI_F32_I16_F80_PTR argument size
This takes an f80 as input and returns an f32. A copy-paste error had
this truncating the input float if !TMP_ABIARGS.
2025-09-10 03:14:30 +01:00
Ryan Houdek 3c6eba561f Merge pull request #4859 from bylaws/x87flushfix
OpcodeDispatcher: Only flush MMX registers on MMX -> x87 transitions
2025-09-09 14:57:28 -07:00
Billy Laws d2714f3338 JIT: Check for suspend interrupts at back-edges
Avoids most cases where an infinite loop could lead to suspend interrups
never getting checked.
2025-09-09 22:35:52 +01:00
Billy Laws b917db34c7 OpcodeDispatcher: Emit RIP metadata for RIP-setting instructions 2025-09-09 22:35:52 +01:00
Billy Laws c45abaaa6b Update InstCountCI 2025-09-09 21:34:59 +01:00
Billy Laws a2286cb00a OpDispatcher: Only use the MM register cache when in MMX state
Since the X87 pass now performs its own MM register caching, these must
be mutually exclusive.
2025-09-09 21:30:33 +01:00
Billy Laws 481787d45c OpDispatcher: Avoid using the register cache for AbridgedFTW
This is now cached internally in the X87 pass, and all operations
outside of that that use it are rare so can afford loading/storing
directly from context after flushing x87 regs.
2025-09-09 21:30:30 +01:00
Billy Laws 756002eef8 unittests: Add test for x87 mode switches wrongly flushing NZCV 2025-09-09 21:29:12 +01:00
Billy Laws 6789919fca X87StackOptimization: Don't flush to mem when entering slow path 2025-09-09 21:29:12 +01:00
Billy Laws d8c615ec65 OpcodeDispatcher: Only flush MMX registers on MMX -> x87 transitions
Flushing other regs is not necessary, and breaks any ConvertNZCVToX87 use
which relies previously saved NZCV values as the flag-setting NZCV op after
the save could trigger a flush of NZCV.
2025-09-09 21:29:12 +01:00
Billy Laws 6a1eb0e508 X87StackOptimization: Defer FTW valid bits calculation 2025-09-09 21:29:12 +01:00
Billy Laws 984b0260d5 X87StackOptimization: Cache stack values 2025-09-09 21:29:12 +01:00
Billy Laws 785e59e889 X87StackOptimization: Cache AbridgedFTW 2025-09-09 21:29:12 +01:00
Billy Laws 671be5b191 X87StackOptimization: Cache half-formed stack register context addresses 2025-09-09 21:29:12 +01:00
Billy Laws b4e556e63a IR: Introduce operation to form a scaled address from STATE 2025-09-09 21:29:12 +01:00
Billy Laws 2128943b57 X87StackOptimization: Use the x87 stack offset cache in SynchronizeStackValues 2025-09-09 21:29:12 +01:00
Billy Laws 6ddb1094b6 X87StackOptimization: Don't invalidate the x87 stack offset cache on push/pop 2025-09-09 21:29:12 +01:00
Billy Laws b9d93b19d4 X87StackOptimization: Defer writeback of the x87 stack pointer 2025-09-09 21:29:12 +01:00
Billy Laws 97a5da232d InstCountCI: Add reduced precision x87 game blocks 2025-09-09 21:29:12 +01:00
Ryan Houdek 0d77af5366 Merge pull request #4858 from bylaws/aluno
OpcodeDispatcher: Don't assert on invalid ALU op encoding
2025-09-09 11:28:29 -07:00
Ryan Houdek a3435f2d22 Merge pull request #4857 from bylaws/wincs
WOW64: Fix CsSeg initialization
2025-09-09 11:23:40 -07:00
Billy Laws 154dd46b6d OpcodeDispatcher: Don't assert on invalid ALU op encoding 2025-09-09 19:11:05 +01:00
Billy Laws 782952d55d WOW64: Fix CsSeg initialization 2025-09-09 19:10:07 +01:00
Ryan Houdek a6bd293ae8 Docs: Update for release FEX-2509 2025-09-09 09:21:22 -07:00
Ryan Houdek decd112a84 Merge pull request #4854 from lioncash/fend
Frontend/OpcodeDispatcher: Remove unused headers
2025-09-08 21:40:30 -07:00
Lioncache f9c15fa0a5 Frontend/OpcodeDispatcher: Remove unused headers
Removes assorted unused headers and forward declarations.
2025-09-09 00:28:05 -04:00
Ryan Houdek 1121f2a1fb Merge pull request #4853 from lioncash/allocator
FEXCore/Allocator: Remove unused headers
2025-09-08 20:25:25 -07:00
Lioncache a7989eb79f FEXCore/Allocator: Remove unused headers
Reveals some more indirect inclusions.
2025-09-08 22:53:09 -04:00
Ryan Houdek 80c199cefb Merge pull request #4852 from lioncash/loading
FileLoading: Add missing self-header
2025-09-08 19:48:01 -07:00
Ryan Houdek 317f92f4c8 Merge pull request #4851 from lioncash/ir
IR: Remove unnecessary forward declarations/headers
2025-09-08 19:47:51 -07:00
Ryan Houdek dd344710d4 Merge pull request #4850 from lioncash/vdso
VDSO_Emulator: Cull unnecessary includes
2025-09-08 19:16:56 -07:00
Lioncache e4e8683c7b FileLoading: Add missing self-header
Ensures that the exposed interface is visible to the implementation (i.e. gets rid of some -Wmissing-declaration warnings).
2025-09-08 22:13:42 -04:00
Lioncache b2407352a9 IR: Remove unnecessary forward declarations/headers
Pares the base IR header down to a modest size and also reveals a few indirect
reliances on the allocator facilities.
2025-09-08 22:03:25 -04:00
Lioncache e1b5e3e112 VDSO_Emulator: Cull unnecessary includes
Reveals a bunch of indirect includes inside the SignalDelegator.
2025-09-08 21:43:31 -04:00
Ryan Houdek 09793d295f Merge pull request #4849 from lioncash/config
Config: Remove unused Context.h include
2025-09-08 16:59:38 -07:00
Ryan Houdek e9f8c01e04 Merge pull request #4848 from lioncash/thread
InternalThreadState: Remove unused includes
2025-09-08 16:59:15 -07:00
Lioncache f69821d7db Config: Remove unused Context.h include
Removes quite a heavy include from the config system and specifies any
indirect inclusion that were relied on because of it.
2025-09-08 15:23:30 -04:00
Lioncache 1e03219ca3 Syscalls/Thread: Add missing forward declarations
Currently we were lucky that these were being seen already when being included.
2025-09-08 14:58:30 -04:00
Lioncache de8b4438e3 ThreadManager: Eliminate indirect inclusions
We had quite a bit of dependencies that were currently being indirectly relied upon.
We can forward declare the relevant ones and add the includes for the ones that are
explicit.
2025-09-08 14:58:27 -04:00
Ryan Houdek 2cfd3330f0 Merge pull request #4847 from lioncash/loader
CodeLoader: Tidy up interface
2025-09-08 11:50:30 -07:00
Ryan Houdek 1988b432ab Merge pull request #4846 from pmatos/fix/VSCond
Fix mapping for VS and VC condition codes
2025-09-08 11:29:35 -07:00
Lioncache 6b2363d3f6 InternalThreadState: Remove unused headers
Removes some includes in the core header and resolves indirect includes elsewhere.
2025-09-08 14:13:01 -04:00
Ryan Houdek e76af56c86 Merge pull request #4845 from lioncash/thunk
Core/Thunks: Remove unused IR header/forward declarations
2025-09-08 10:03:34 -07:00
Ryan Houdek b7173a3ae2 Merge pull request #4844 from lioncash/builtin
General: std::alignment_of -> alignof
2025-09-08 10:03:24 -07:00
Ryan Houdek 3f71517e1f Merge pull request #4843 from lioncash/util
Common: Remove duplicated StringUtil functionality
2025-09-08 10:03:14 -07:00
Lioncache 17fbe52cf6 CodeLoader: Mark GetStackPointer() as const
This isn't intended to modify anything.
2025-09-08 12:19:15 -04:00
Lioncache ecc526a9d8 CodeLoader: Make GetApplicationArguments() return by reference
We don't really need to make this potentially return a null pointer when we could always return a valid instance
that would just happen to be empty if unused.
2025-09-08 12:07:19 -04:00
Lioncache 436f4aa953 CodeLoader: Make GetExecveArguments() return by value
Same behavior, but more straightforward.
2025-09-08 12:07:19 -04:00
Lioncache fd10ce0600 CodeLoader: Return auxv results by value
Same behavior, but without the need to care about out values.
2025-09-08 12:07:14 -04:00
Lioncache d9cd40eb48 CodeLoader: Remove unused AddIR() member function
This isn't used in any implementations anymore.
2025-09-08 10:34:13 -04:00
Lioncache 9b9ab16b55 Common: Remove duplicated StringUtil functionality
We already have an equivalent header within FEXCore that's header only,
so we can adapt it to conform for both cases, allowing for removal of
one of them.
2025-09-08 09:44:25 -04:00
Paulo Matos 14256e7a90 Fix mapping for VS and VC condition codes
This doesn't seem to be generated in-code therefore it's not fixing any
existing bug, but fixes the mapping for future uses of the condition code.
2025-09-08 11:41:15 +02:00
Lioncache d36c387c1f Core/Thunks: Remove unused IR/forward declarations
Moves the IR include into one of the more specific headers, which avoids dumping the IR header into any core bits that use the interface.

Also uncovered a missing header guard.
2025-09-06 22:28:27 -04:00
Lioncache 7e61b172b2 General: std::alignment_of -> alignof
We can just use the compiler built-in in these cases. std::alignment_of
mainly has utility in metaprogramming.
2025-09-06 21:23:10 -04:00
Ryan Houdek bd0cae9298 Merge pull request #4825 from bylaws/unityomg
Frontend: Force acq/rel semantics for known Unity ringbuffer offsets
2025-09-06 17:09:13 -07:00
Ryan Houdek 6ae94581bb Merge pull request #4821 from bylaws/monof
Frontend: Fix tailcall handling when mono hacks are enabled
2025-09-06 16:51:08 -07:00
Ryan Houdek ea8ae2bf2c Merge pull request #4839 from lioncash/utils
Common Utils: Remove unused/unnecessary inclusions from headers
2025-09-06 15:50:00 -07:00
Ryan Houdek 694bc3312e Merge pull request #4842 from lioncash/python
config_generator: Remove superfluous format arguments
2025-09-06 15:49:50 -07:00
Ryan Houdek f565939437 Merge pull request #4841 from lioncash/env
Common: Remove EnvironmentLoader.cpp/.h
2025-09-06 15:49:42 -07:00
Lioncache 272d2747cd config_generator: Remove excess formatting arguments for EnumParser handling
Each instance had one more argument than necessary.
2025-09-06 11:18:38 -04:00
Lioncache 675956c92d config_generator: Indent ENVLOADER conditional bodies
Makes the indentation consistent with the rest of the generated file.
2025-09-06 10:57:31 -04:00
Lioncache 03410ed4a2 config_generator: Remove unnecessary format() in print_parse_jsonloader_options()
This format call essentially does nothing since it has no specifier.
We can also fix up the indentation of the else case too, for consistency.
2025-09-06 10:51:02 -04:00
Lioncache ac88e7fabf Common: Remove EnvironmentLoader.cpp/.h
This has been unused since the transition to the layered config system.
2025-09-06 10:23:53 -04:00
Tony Wasserka 71d42d7e73 Merge pull request #4840 from lioncash/internal
InternalThreadState: Make bool operator explicit
2025-09-06 09:20:08 +02:00
Lioncache 2613762ac2 InternalThreadState: Make bool operator explicit
We definitely don't want the boolean null test to be able to be implicitly converted
(e.g. to an int or whatever else).

For example a non-explicit bool operator allows for silly things like:

NonMovableUniquePtr<...> ptr;
// ...
auto k = 5 + ptr;

to build without issue, which we should really force the user to be explicit about if it's *really* a desired behavior.
2025-09-05 16:29:19 -04:00
Lioncache 831b21cd39 JitSymbols: Remove unused includes
Reduces header dependencies (and clarifies existing ones).
2025-09-05 14:27:49 -04:00
Lioncache 7869d80073 VolatileMetadata: Add missing include 2025-09-05 14:16:23 -04:00
Lioncache ea39d80e83 StringUtil: Move algorithm header to cpp file
The header doesn't make any direct use of it.
2025-09-05 14:16:23 -04:00
Lioncache 62e481a608 JSONPool: Remove unnecessary bounds check
Given the generic context here, we should just use std::data
2025-09-05 14:16:20 -04:00
Lioncache bac7148d69 JSONPool: Remove unnecessary include
This isn't used in this header.
2025-09-05 13:52:05 -04:00
Lioncache 62f16a3c4b FEXServerClient: Remove unnecessary includes from header
We can reduce includes, such as logging by specifying a concrete size for logging levels,
allowing the enum to be forward declared. We can also move FillHeader into the cpp file,
allowing the syscalls header to be removed.
2025-09-05 13:45:31 -04:00
Lioncache 248ecd54e4 FDUtils: Remove unused headers 2025-09-05 13:30:09 -04:00
Lioncache 1f59accb4e CPUInfo: Remove unused header 2025-09-05 13:28:54 -04:00
Lioncache dd7e4375cf Config: Add missing includes 2025-09-05 13:26:43 -04:00
LC 3df3999138 Merge pull request #4838 from lioncash/dumper
IRDumper: Add missing PrintArg for IndexNamedVectorConstant
2025-09-05 13:06:53 -04:00
Ryan Houdek 1ca6443c9d Merge pull request #4836 from neobrain/refactor_irscript_cleanup
json_ir_generator: Clean up generated assertion statements
2025-09-04 11:21:10 -07:00
Ryan Houdek a804cb5d07 Merge pull request #4833 from lioncash/ptr
CPUBackend: Remove unused variable in CheckCodeBufferUpdate
2025-09-04 11:20:42 -07:00
Ryan Houdek a0883c9615 Merge pull request #4832 from lioncash/threads
Utils/Threads: Minor cleanup
2025-09-04 11:20:24 -07:00
Ryan Houdek 03c968a8fe Merge pull request #4831 from lioncash/header
IntrusiveIRList/RegisterAllocationData: Remove unused headers
2025-09-04 11:20:14 -07:00
Ryan Houdek e34de22cf4 Merge pull request #4829 from lioncash/clz
VixlUtils: Use std::countl_zero where applicable
2025-09-04 11:19:54 -07:00
Ryan Houdek 3db1f1e01f Merge pull request #4826 from Sonicadvance1/classify_features
unittests/ASM: Adds support for classifying more requirements
2025-09-04 11:19:41 -07:00
Ryan Houdek 5e467c737f Merge pull request #4822 from bylaws/ecent
Dispatcher: Move the EC entrypoint check prior to the L1 lookup
2025-09-04 11:19:29 -07:00
Lioncache dfc42fa7c5 IRDumper: Add missing PrintArg for IndexNamedVectorConstant
Results in better named output for these.
2025-09-04 12:04:48 -04:00
LC f1016e82c7 Merge pull request #4834 from neobrain/refactor_async_simplify_ancbuffer
AsyncNet: Simplify declaration of ancillary buffer
2025-09-04 11:28:41 -04:00
Tony Wasserka 5e099a5dbd IR: Fix indentation 2025-09-04 15:39:40 +02:00
Tony Wasserka 8dceec9316 IR: Fix array emptiness check 2025-09-04 15:39:40 +02:00
Tony Wasserka 6f51092998 AsyncNet: Simplify declaration of ancillary buffer 2025-09-04 14:40:39 +02:00
Lioncache 4e83ca67f5 CPUBackend: std::move shared_ptr where applicable
Very minor, just avoids churning reference count increments and decrements.
2025-09-03 21:20:58 -04:00
Lioncache 76e6c6f8ea CPUBackend: Remove unused shared_ptr in CheckCodeBufferUpdate()
Just gets rid of an unused variable.
2025-09-03 21:19:16 -04:00
Ryan Houdek 343a529e63 Merge pull request #4824 from bylaws/tf2
OpDispatcher: Force a recheck of TF after POPF
2025-09-03 15:24:40 -07:00
Ryan Houdek c6ccc68607 Merge pull request #4823 from bylaws/mbooo
Frontend: Always explore fallthrough branches with multiblock
2025-09-03 15:24:29 -07:00
Ryan Houdek 8009f4a6ee Merge pull request #4820 from bylaws/neg
AtomicOps: Reimplement AtomicNeg using 8.1 CAS atomics
2025-09-03 15:23:40 -07:00
Tony Wasserka e126cecb16 Merge pull request #4818 from lioncash/at
General: Replace usages of .at(0) with .data() where applicable
2025-09-03 22:37:50 +02:00
Tony Wasserka 5e19202540 Merge pull request #4816 from lioncash/float
F80Fallbacks: Remove unnecessary memset
2025-09-03 22:36:59 +02:00
Billy Laws 06466bac0e Dispatcher: Drop the L1 lookup
This is almost never hit outside of exceptional cases like signals/SMC,
for which other overheads dominate.
2025-09-03 18:26:47 +01:00
Billy Laws 156edcacb2 Dispatcher: Move the EC entrypoint check prior to the L1 lookup
Profiling shows this is by far the hottest path, as every DX call will
hit this (as arm64ec functions in a vtable). L2 hits are much rarer and
L1 almost never hits since that is looked up inline in most cases.
2025-09-03 18:26:47 +01:00
Lioncache bf4b8d246a Threads: Simplify copying in SetInternalPointers
We can just use assignment here, which lets us also
remove <cstring> since this is the only usage of memcpy.
2025-09-03 11:01:12 -04:00
Lioncache abc37ec35e Threads: Mark default handler functions as static
These aren't used outside of the translation unit.
2025-09-03 11:01:12 -04:00
Lioncache dd735464a2 Threads: Remove unnecessary includes
We have quite a few unnecessary includes here left over from code movement,
so we can clean those out.

Also fix an indirect include in the header.
2025-09-03 11:00:57 -04:00
Lioncache 7041bdb144 IntrusiveIRList: Remove unused headers
Gets rid of unnecessary dependencies and clarifies previously indirect dependencies.
2025-09-03 10:39:45 -04:00
Lioncache 19e8c3fcb6 RegisterAllocationData: Remove unused headers
Gets rid of unnecessary header dependencies.
2025-09-03 10:31:41 -04:00
LC 06c5135d4e Merge pull request #4815 from lioncash/abi
Dispatcher: Reduce header dependencies
2025-09-03 09:45:36 -04:00
LC f66066cbd8 Merge pull request #4819 from lioncash/constprop
ConstProp: Remove file
2025-09-03 09:45:01 -04:00
Lioncache a144622a06 VixlUtils: Remove IsPowerOf2()
std::has_single_bit already does this.
2025-09-03 05:46:29 -04:00
Lioncache 05c814b8c1 VixlUtils: Use std::countl_zero where applicable
Makes for a little less code.
2025-09-03 05:45:02 -04:00
Ryan Houdek e0af51fb06 unittests/ASM: Classify unittests with new flags
Necessary for running on old hardware.
2025-09-02 18:37:59 -07:00
Ryan Houdek 7b749558df HostRunner: Add support for non-fsgsbase hardware 2025-09-02 18:37:59 -07:00
Ryan Houdek 1797c3c617 unittests/ASM: Add support for more feature flags 2025-09-02 18:15:31 -07:00
Billy Laws 28a1c28cb4 Frontend: Track the raw opcode and use for mono tailcall detection
This path was nonfunctional prior to this, as the jump is a group-opcode
which wouldn't be picked up (as Op is rewritten in NormalOp).
2025-09-02 22:42:02 +01:00
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Billy Laws 7c9c97f2b9 Update InstCountCI 2025-09-02 22:23:41 +01:00
Billy Laws d72e121fd4 AtomicOps: Reimplement AtomicNeg using 8.1 CAS atomics 2025-09-02 22:23:14 +01:00
Billy Laws 003fc3dab5 Update InstCountCI 2025-09-02 22:21:56 +01:00
Billy Laws 2c883d7cdc OpDispatcher: Force a recheck of TF after POPF
The dispatcher/block linker will handle this, but if the instruction
following a POPF flag doesn't otherwise trigger one of those the
interrupt would be missed.
2025-09-02 22:21:42 +01:00
Billy Laws fcba49768c Frontend: Force acq/rel semantics for known Unity ringbuffer offsets
Unity games crash with TSO disabled due to the SPSC GfxDevice
ThreadedStreamBuffer read/write pointer updates missing acq/rel
semantics. Rather than attempting to use heuristics to match cases like
this, which could be overzealous and hit more accesses than necessary, just
target the problem directly as this is consist across 32/64 bit Unity versions
for at least the past 10 years. Gate this behind the existing Unity mono
hacks to avoid false-positives in non-Unity games.
2025-09-02 22:05:22 +01:00
Billy Laws 959af9c3af Frontend: Always explore fallthrough branches with multiblock 2025-09-02 22:02:53 +01:00
Lioncache 1edacf05e9 ConstProp: Remove file
This was left in the tree, but the functionality was removed a while ago, so we can delete this.
2025-09-02 14:50:22 -04:00
Lioncache 1f756fb386 General: Replace usages of .at(0) with .data() where applicable
This is generally more straightforward in cases where bounds checking isn't necessary.
2025-09-02 13:03:37 -04:00
Lioncache 5ebdc01783 F80Fallbacks: Remove unnecessary memset
X80SoftFloat instances already initialize zeroed out in the default constructor.
2025-09-02 12:25:03 -04:00
Lioncache aa29201e00 Dispatcher: Reduce header dependencies
Gets rid of some unnecessary dependencies and also resolves some indirect dependencies.
2025-09-02 12:10:05 -04:00
LC b98acdf97a Merge pull request #4814 from neobrain/refactor_code_cache_remove_legacy
CodeCache: Remove legacy interfaces
2025-09-02 07:41:57 -04:00
LC 7aeb4788ed Merge pull request #4813 from Sonicadvance1/movsxd_test
unittests: Adds missing movsxd edge case
2025-09-02 07:40:24 -04:00
Tony Wasserka 7b1db40d4a Config: Remove legacy code caching interfaces 2025-09-02 12:06:57 +02:00
Tony Wasserka dcaa90a855 CodeCache: Remove legacy interfaces 2025-09-02 12:06:52 +02:00
Ryan Houdek 0a74efe3f8 unittests: Adds missing movsxd edge case
nasm doesn't let you encode all of these but we had already supported
them in FEX.

It's a bit of a weird edge case behaviour, since it's supposed to be
sign extended a 32-bit value in to a 64-bit register but with prefixes
(or lack of) it can be a move with zero-extend or a 16-bit insert.

Make sure we're actually testing these edge cases.
2025-09-02 01:16:42 -07:00
LC 5a73134c7c Merge pull request #4812 from Sonicadvance1/disable_gvisor
gvisor_tests: Update expectations for kernel 6.16
2025-09-02 04:15:37 -04:00
Ryan Houdek 43d535cf8d gvisor_tests: Update expectations for kernel 6.16 2025-09-01 16:13:42 -07:00
Ryan Houdek 4a71edf7bc Merge pull request #4811 from pmatos/fix/SPEC_GCC-C-execute-ieee-fp-cmp-8l
Fix IEEE 754 unordered comparison detection in x87 floating-point operations
2025-09-01 14:39:59 -07:00
Ryan Houdek ae90040ed2 Merge pull request #4810 from lioncash/file
FileManagement: Minor interface cleanup
2025-09-01 14:39:48 -07:00
Ryan Houdek fbf982f001 Merge pull request #4809 from lioncash/ir
IR: Remove some unimplemented prototypes
2025-09-01 14:39:39 -07:00
Ryan Houdek 839e3ac775 Merge pull request #4808 from lioncash/sig
Dispatcher: Centralize SignalDelegator config creation
2025-09-01 14:39:30 -07:00
Paulo Matos 65d230f32b instcountci: Fix IEEE 754 unordered comparison detection in x87 floating-point operations 2025-09-01 21:43:23 +02:00
Paulo Matos 0882b0db0c asm_tests: Fix IEEE 754 unordered comparison detection in x87 floating-point operations 2025-09-01 21:42:01 +02:00
Paulo Matos 3d92c1fb32 Fix IEEE 754 unordered comparison detection in x87 floating-point operations
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.

Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
2025-09-01 18:26:48 +02:00
Lioncache d4e679432c FileManagement: Add missing includes
Gets rid of a few indirect dependencies.
2025-08-29 13:42:17 -04:00
Lioncache 54c8cb9fd7 FileManagement: Mark helper functions as const where applicable
These don't modify internal class state, so we can mark them as such.
2025-08-29 13:41:58 -04:00
Lioncache 221ee1a122 IREmitter: Remove unimplemented prototypes
Get/SetPackedRFLAG() functions were moved to the opcode dispatcher, but these were
accidentally left behind.
2025-08-29 12:00:53 -04:00
Lioncache 2b29f90c92 IREmitter: Remove unused <array> header
Just one less include to worry about
2025-08-29 11:57:03 -04:00
Lioncache d7ec0e570f IR: Make IsFragmentExit() internally linked
This isn't used outside of the IR emitter. Plus, IsBlockExit() is already a more general interface to use,
since it handles the fragment exit case as well.
2025-08-29 11:51:49 -04:00
Lioncache 86c2144943 Core: Simplify define in InitCore()
Instead of inverting off a separate define, we can combine both checks together.
2025-08-29 11:30:48 -04:00
Lioncache ee85230db2 Dispatcher: Remove unused IntCallbackReturnAddress member
Turns out that this isn't used anymore. Found due to an unused field warning
when moving all public members into the private section.
2025-08-29 11:19:10 -04:00
Lioncache fde99a8dc1 Dispatcher: Centralize SignalDelegator config creation
Since all of the information comes from the Dispatcher, we can have the dispatcher
provide that information. This way we can also eliminate a bunch of now-redundant
public interface members and simplify the config setup within InitCoreImpl().

Conveniently, this also allows making all members of the dispatcher non-public.
2025-08-29 11:19:07 -04:00
Lioncache 2e692a1146 SignalDelegator: Move struct outside of SignalDelegator
The name itself is already qualified with SignalDelegator, so this can reasonably be outside the class itself.
This also allows for forward declarations of the config struct
2025-08-29 10:23:03 -04:00
Lioncache 73f36fb5b3 SignalDelegator: Move includes into cpp file
Reduces header dependencies
2025-08-29 10:07:56 -04:00
Ryan Houdek f866ad52dc Merge pull request #4807 from Sonicadvance1/remove_visibility
Remove visibility on functions that is no longer necessary
2025-08-28 18:34:37 -07:00
LC 9566910574 Merge pull request #4806 from Sonicadvance1/avx_constexpr
X86Tables: Convert AVX tables to constexpr
2025-08-27 19:18:41 -04:00
Ryan Houdek 4188d4d10b Add hidden visibility to FEXLoader/LinuxEmulation/Common 2025-08-27 16:14:59 -07:00
Ryan Houdek d75152cb74 IR: Remove visibility on functions that is no longer necessary
These used to be required because we leaked IR semantics across the API
boundary. Since we no longer do, we can remove the visibility from
these.
2025-08-27 15:34:57 -07:00
Ryan Houdek add54b8089 X86Tables: Convert AVX tables to constexpr
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
2025-08-27 14:55:10 -07:00
LC 7431495f13 Merge pull request #4797 from Sonicadvance1/more_brk_fixes
More BRK fixes
2025-08-27 16:49:55 -04:00
Ryan Houdek 5ed5566106 X86Tables/AVX128: Stop installing state functions a second time 2025-08-27 13:09:53 -07:00
Ryan Houdek 2d9a4412a4 X86Tables: Have AVX128 table always have PCLMUL 2025-08-27 13:08:27 -07:00
Ryan Houdek 982abd5719 X86Tables: Have AVX256 table always have PCLMUL 2025-08-27 13:06:43 -07:00
Ryan Houdek f09db0cbe5 Merge pull request #4805 from Sonicadvance1/hammer_hammer
FEXCore: Move most instruction tables to be constexpr
2025-08-27 12:57:21 -07:00
Ryan Houdek 80cacc462e X86Tables: Move x87 tables to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 1a2b2f8870 X86Tables: Move Second ModRM table to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek e8c576047f X86Tables: Moves Base ops to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 995d3152eb X86Tables: Convert Secondary tables to constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek ddc40cc273 X86Tables: Converts PrimaryGroup table to constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 5ee311b831 X86Tables: Converts H0F3A table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek ea90296e22 X86Tables: Converts H0F38 table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek 58c1a8d148 X86Tables: Convert DDDNow table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek 87a19c7938 FEXCore: Move SecondaryGroupTables Arch specific ops to a table
No runtime-installation necessary.
2025-08-27 12:24:24 -07:00
Ryan Houdek 5b0709de2f Syscalls: Removes an old workaround for MAP_FIXED_NOREPLACE
We don't support these old kernels, so remove this workaround.
2025-08-27 11:52:10 -07:00
Ryan Houdek 87db37c65e Syscalls: Allocate brk granules at page size
This matches kernel behaviour for how brk works. I initially implemented
this as a way to speed up some early applications that relied heavily on
brk. We're long past that time and we should match Linux behaviour
instead.

This fixes issues where applications (FEXLoader really) can see we've
allocated a larger granule and breaks self-hosting. There's no reason to
be smarter around this, if applications want more perf then they'll
switch to a better allocator than brk.

Also need to make sure that after the brk region has been reserved, to
unmap its initial mapping to allow future mmap syscalls to overwrite it.
We need to reserve it early to ensure the region initially exists and we
don't accidentally map other things in that space. Then it is up to the
guest if they don't want to overwrite it.
2025-08-27 11:51:52 -07:00
Ryan Houdek 100db13f99 Syscalls: Make sure to munmap on full brk zero
Failed to munmap in this case.
2025-08-27 11:51:52 -07:00
Ryan Houdek d2061b27a2 Syscalls: Rename DataSpaceMaxSize
Implied the maximum size that brk could grow to. Which isn't correct, it
is the current maximum size mapped which is page size, versus the
current brk offset which is byte ranged.
2025-08-27 11:51:52 -07:00
Ryan Houdek 002dfb9b84 Syscalls: Remove unused DataSpaceStartingSize 2025-08-27 11:51:52 -07:00
Ryan Houdek 0ab1241f09 X86Tables: Switch OpDispatch pointer over to a union
A wild use case of union over a variant because we don't want to
increase the encoding size from 128-bit to 256-bit (because of padding).
The type of operation is encoded with the table operation type, so a
variant is unnecessary and we get to keep the 128-bit encoding.

This allows us to have "recursive" x86 table descriptions. But in
reality this is going to only be one layer deep. As this will allow the
Frontend decoder to select instruction encodings based on arch bitness
once the tables are generated correctly.
2025-08-26 15:36:37 -07:00
Ryan Houdek aa63656ff1 X86Tables: Removes nullptr on DispatchPtr
Let it value initialize to zero which will be necessary when it switches
to a union.
2025-08-26 15:36:33 -07:00
LC 6ba32da73a Merge pull request #4804 from Sonicadvance1/portable_tests
Github: Force FEX portable config
2025-08-26 12:52:22 -04:00
Ryan Houdek 1129e71069 Github: Force FEX portable config
Noticed some runners had generated a .fex-emu/Config.json, run the tests
in portable mode to ensure they don't pick up a spurious config option.

FEXLinuxTests still ran non-portable because it's required for the thunk
tests.
2025-08-25 13:32:51 -07:00
Ryan Houdek 3b32fd5c49 Merge pull request #4794 from neobrain/fix_reformat_script
Scripts: Fix use of removed helper script
2025-08-25 12:20:33 -07:00
Ryan Houdek 7df1ed271f Merge pull request #4799 from neobrain/refactor_ranges
VolatileMetadata: Clean up implementation using ranges::views::split
2025-08-25 11:45:44 -07:00
Ryan Houdek 67e9b40bab Merge pull request #4803 from neobrain/refactor_irdumper_cleanup
IRDumper: Clean up formatting using fmt
2025-08-25 11:43:36 -07:00
Ryan Houdek 62b80903f0 Merge pull request #4801 from neobrain/refactor_maybe_unused
Reduce use of maybe_unused attributes
2025-08-25 11:43:09 -07:00
Ryan Houdek 993d917478 Merge pull request #4800 from neobrain/fix_assert_logargs
LogManager: Evaluate non-predicate arguments for assertions lazily
2025-08-25 11:41:05 -07:00
Ryan Houdek ac20880a14 Merge pull request #4802 from neobrain/fix_fedora_rootfs
FEXConfig: Fix RootFS override in muvm-based setups
2025-08-25 11:40:26 -07:00
Tony Wasserka 99cfe05ee5 IRDumper: Clean up formatting using fmt 2025-08-25 17:05:10 +02:00
Tony Wasserka 7c187e82b7 FEXConfig: Don't set an empty RootFS if none is selected 2025-08-25 16:30:58 +02:00
Tony Wasserka 2a760660a9 Merge pull request #4795 from pmatos/formatCI
Do not use diff_from_common_commit
2025-08-25 14:03:24 +02:00
Tony Wasserka 4b11826a7d LogManager: Evaluate non-predicate arguments for assertions lazily
This allows assertions like "some_iterator == end()" to dereference the
unexpected iterator value when formatting the log message on failure.
2025-08-25 10:54:39 +02:00
Tony Wasserka 887f586874 IRDumper: Remove unneeded maybe_unused attribute 2025-08-25 10:36:01 +02:00
Tony Wasserka db94c04179 FEXCore: Use inline functions instead of maybe_unused static ones 2025-08-25 10:36:01 +02:00
Tony Wasserka 51e64c69f3 LogManager: Unconditionally evaluate assertion conditions
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
2025-08-25 10:36:01 +02:00
Tony Wasserka 97c46843d2 LogManager: Evaluate non-predicate arguments for assertions lazily
This allows assertions like "some_iterator == end()" to dereference the
unexpected iterator value when formatting the log message on failure.
2025-08-25 10:21:14 +02:00
Tony Wasserka 16310c91c9 VolatileMetadata: Clean up implementation using ranges::views::split 2025-08-25 09:29:27 +02:00
Tony Wasserka 220e421c22 External: Add range-v3 submodule 2025-08-25 08:47:56 +02:00
LC a156752fae Merge pull request #4796 from Sonicadvance1/brk_fixes
ELFCodeLoader: Fixes BRK allocation with non-PT_DYN ELF files
2025-08-23 20:03:54 -04:00
Ryan Houdek 5d1c574dc5 ELFCodeLoader: Fixes BRK allocation non PT_DYN ELF files
When loading an ELF file (usually PT_EXEC), we would have BSS regions
that ended up in the high pages of the address space. We would then
attempt mapping BRK after whatever the highest address ending up being.

This was problematic because we used `MAP_FIXED` which means that the
BRK region could cross in to 64-bit address space, overwrite VDSO,
overwrite the stack, vsyscall, maybe a couple of other things.

Instead of letting that happen, reorder some of the logic so that in
`LoadElfFile` will ensure the full `BRK_SIZE` is allocated (or return
zero) and then we can ensure we're never overwriting other various
memory regions.

This doesn't fix two fundamental issues that FEX has with brk:
* If BRK gets mapped below the ELF, that should mean it isn't mapped at all
  * Linux-isms, doubt this breaks new software
* We still map the full 8MB, which isn't correct
  * I'm working on fixing this issue, which is why I encountered this.
2025-08-22 20:30:04 -07:00
Ryan Houdek 6310c20217 Merge pull request #4773 from Sonicadvance1/extended_volatile_metadata
Implement support for additional provided volatile metadata
2025-08-22 15:15:50 -07:00
Ryan Houdek 0dfc30fdd8 APITests: Adds extendedvolatile 2025-08-22 15:03:56 -07:00
Ryan Houdek f050571696 Implement support for additional provided volatile metadata
Only wired up for wow64 and arm64ec. Gives more granular control over
TSO enabling and disabling. Matches arm64ec volatile metadata except
with one more additional feature that whole modules can be disabled at a
time.

Once we know the mapped size of files in Linux then we'll be able to do
the same thing there, but there's not a full mechanism wired up for that
yet.
2025-08-22 15:03:56 -07:00
Ryan Houdek 6db5bc3721 FEXCore/Utils/IntervalList: Adds a couple functions 2025-08-22 15:03:56 -07:00
LC 4f84c281ef Merge pull request #4791 from bylaws/norace
ARM64EC: Cooperatively suspend if needed when leaving the JIT
2025-08-22 12:21:09 -04:00
LC 0ca438e081 Merge pull request #4790 from bylaws/nomorewine
WinAPI: Force WOW64 allocations to start >4GB
2025-08-22 12:17:28 -04:00
LC 2472aee473 Merge pull request #4786 from Sonicadvance1/noise_fcw
Dispatcher: Most minor of optimizations for f64 x87
2025-08-22 12:16:43 -04:00
LC 714856f74c Merge pull request #4788 from Sonicadvance1/modify_ldt_x64
LinuxSyscalls: Support modify_ldt on x64
2025-08-22 12:15:17 -04:00
Paulo Matos e98932df92 Do not use diff_from_common_commit 2025-08-22 15:10:17 +02:00
Tony Wasserka 95623dac63 Scripts: Fix use of removed helper script
This was accidentally changed back to clang-format.py in
3b322d8f37.
2025-08-22 11:05:01 +02:00
Billy Laws 827015a8e3 ARM64EC: Cooperatively suspend if needed when leaving the JIT 2025-08-21 17:24:56 +01:00
Billy Laws a379818386 WinAPI: Force WOW64 allocations to start >4GB
Avoids stealing address space from the guest.
2025-08-21 17:19:46 +01:00
Ryan Houdek 4516e4e9f6 LinuxSyscalls: Support modify_ldt on x64
This nearly gets FEX's TestHarnessRunner to be self-hosting inside of
FEX. The only thing blocking it currently is that our SBRK emulation
reserves the whole region, when it should be "soft-reserved" and mmap
with MAP_FIXED_NOREPLACE can override it. Plus an assert in
OpcodeDispatcher preventing any 32-bit code from running from a 64-bit
process.

In the most simple terms, gdt and ldt are setup to be unique per thread,
and modify_ldt then modifies that thread's ldt entry. On thread
creation, these values get inherited as a copy.

This allows installation of 32-bit code entries, which with the previous
PRs merged allows the code to attempt jumping to that 32-bit code entry.
It then will immediately explode with an assert in our OpcodeDispatcher.
We can't allow 32-bit code jumping yet until our OpDispatcher/X86Tables
allows runtime selection of 64-bit and 32-bit code entries which is
still a ways away.

With the assert removed and the SBRK code handling hacked out,
/technically/ the TestHarnessRunner can run some code, albeit anything
that changes behaviour between bitness is completely incorrect.
2025-08-20 13:22:56 -07:00
Ryan Houdek 83e44961fc Merge pull request #4784 from Sonicadvance1/more_ldt_gdt
FEXCore: Move LDT/GDT management to the frontend
2025-08-20 13:22:08 -07:00
Ryan Houdek bc2e959897 Merge pull request #4742 from pmatos/fninit_IE
Clear IE flag on fninit
2025-08-19 12:57:25 -07:00
Ryan Houdek 1b0b4ba416 InstCountCI: Add frontend handling of GDT 2025-08-18 14:03:51 -07:00
Ryan Houdek 25a0752e58 FEXCore: Move LDT/GDT management to the frontend
This will allow the frontend to manage LDT and GDT so that Linux will be
able to install 32-bit code segments.
2025-08-18 14:03:51 -07:00
Paulo Matos b0a7e0bbcc instcountci: Clear IE flag on fninit 2025-08-18 15:49:10 +02:00
Paulo Matos af561abc71 asm_tests: Clear IE flag on fninit 2025-08-18 15:40:46 +02:00
Paulo Matos 3bf323aa2b Clear IE flag on fninit 2025-08-18 15:40:46 +02:00
Ryan Houdek c9a33c638a Dispatcher: Most minor of optimizations for f64 x87
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
2025-08-17 19:50:21 -07:00
Ryan Houdek 8d1042859d CoreState: Use a named constant for LDT/GDT
Just to make this easier, NFC.
2025-08-15 19:09:33 -07:00
LC 9072571fc1 Merge pull request #4783 from Sonicadvance1/save_restore_segments
LinuxSyscalls: Save and restore segment registers on signal
2025-08-14 15:34:10 -04:00
Ryan Houdek d8748b8216 LinuxSyscalls: Save and restore segment registers on signal
Effectively NFC today while we don't allow installing LDT entries, but
is required to have signal handlers execute in the correct bitness if
32-bit code segments are enabled.

Nothing too crazy, just that four segments are stuffed in to the
`REG_CSGSFS` data value, and then CS/SS gets restored on restore.
FS and GS are ignored in this instance because edge case behaviour where
those don't get reset on signal entry, so if they get changed then it is
on the userspace application to fix it.

Gets another change out of my stashes.
2025-08-14 12:22:28 -07:00
Ryan Houdek 2f930b201a Merge pull request #4782 from bylaws/oo0
OpcodeDispatcher: Simplify CALLOp return addr calc
2025-08-14 11:23:21 -07:00
Ryan Houdek 26b4195f1f Merge pull request #4779 from asLody/environ
Windows: Support EnvLoader via _environ
2025-08-14 11:22:23 -07:00
Billy Laws 5c078d13c0 Update InstCountCI 2025-08-14 15:35:23 +01:00
LC 2825ac282a Merge pull request #4781 from Sonicadvance1/implement_call_ret_far
OpcodeDispatcher: Implement support for CALLF/RETF
2025-08-14 00:15:31 -04:00
Lody cd31a79f41 Windows: Support EnvLoader via _environ 2025-08-14 10:41:11 +08:00
Ryan Houdek 52eb9f46d2 InstcountCI: Add callf/retf 2025-08-13 16:33:23 -07:00
Billy Laws 08c430dcfb OpcodeDispatcher: Simplify CALLOp return addr calc 2025-08-14 00:25:23 +01:00
Ryan Houdek b486aeeac5 unittests/ASM: Unittests for callf/retf
Not too crazy, just make sure to check RSP locations on both sides.
2025-08-13 16:20:18 -07:00
Ryan Houdek 13da6102b3 OpcodeDispatcher: Implement support for CALLF/RETF
Similar to the previous far jmp, if the CS changes operating mode then
things will still explode with other FEX asserts. But this gets another
change out of my stashes.

These instructions go hand-in-hand obviously so they get implemented as
a pair.
2025-08-13 16:18:12 -07:00
LC 1e1c4e017e Merge pull request #4780 from Sonicadvance1/implement_jump_far
OpcodeDispatcher: Implement support for far jump
2025-08-13 15:58:57 -04:00
Ryan Houdek ceaaab5a86 unittests/ASM: Test jmpf without cs change 2025-08-13 12:26:34 -07:00
Ryan Houdek 2b7d03d4f9 OpcodeDispatcher: Implement support for far jump
If anything actually attempts to use this to change CS then it'll very
quickly hit some other asserts, but it gets one more change off my list.
2025-08-13 12:19:16 -07:00
LC c8810557c1 Merge pull request #4777 from Sonicadvance1/move_gdt_up
FEXCore: Split GDT and LDT prep work
2025-08-13 12:06:41 -04:00
Ryan Houdek 2a4bfe49f5 Merge pull request #4743 from pmatos/fclex
Implementation of FCNLEX to clear IE flag
2025-08-12 12:35:15 -07:00
Ryan Houdek b71879e90e Merge pull request #4776 from bylaws/correct
Windows: Misc race and correctness fixes
2025-08-12 12:07:53 -07:00
Billy Laws 63ae262752 Merge pull request #4778 from FEX-Emu/fix_3dnow_doc
HostFeatures: Fix insufficient documentation to reproduce 3DNow! bug
2025-08-12 15:06:36 +01:00
Tony Wasserka 2c3a4fef26 HostFeatures: Fix insufficient documentation to reproduce 3DNow! bug
This ensures we don't forget about it if bug diagnosis leads nowhere.
2025-08-12 15:55:45 +02:00
Ryan Houdek e696e32b95 InstcountCI: Update 2025-08-11 20:48:47 -07:00
Ryan Houdek 3f9df49e3e FEXCore: Split GDT and LDT prep work
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.

Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
2025-08-11 20:45:51 -07:00
Billy Laws 1ac3d3c5a6 Windows: Extend user handle access masks in callbacks where necessary
While most applications will just create ALL_ACCESS handles which
support the additional operations FEX needs over the regular windows syscall
(usually just QUERY_INFORMATION to lookup the thread object), it is
valid for them to pass in the minimal set required for each operation
and windows/wine will reject any other operations (e.g. querying the TEB
base on a handle with only the SUSPEND right). Windows permits
duplicating handles to extend their permissions given the process itself
has such permissions, so do that as necessary for the operations FEX
needs.
2025-08-12 03:19:20 +01:00
LC b3d1aade71 Merge pull request #4775 from Sonicadvance1/fix_32bit_cmpxchg
FEXCore: Fixes 32-bit zero-extend semantics on cmpxchg
2025-08-11 21:53:14 -04:00
Ryan Houdek c300c723f7 InstcountCI: Update 2025-08-11 17:50:18 -07:00
Ryan Houdek d0ec069dcb FEXCore: Fixes 32-bit zero-extend semantics on cmpxchg
We were missing the semantics when the result was a success.
2025-08-11 17:49:17 -07:00
LC 6811406b24 Merge pull request #4774 from Sonicadvance1/fix_o16_enterleave
FEXCore: Fixes 16-bit operand size Enter and Leave instructions
2025-08-11 20:15:44 -04:00
Ryan Houdek 0879994aff FEXCore: Fixes 16-bit operand size Leave
Weren't testing this edge case just like enter
2025-08-11 17:01:49 -07:00
Billy Laws b2b5ccf69c Windows: Validate any user handle permissions for syscall callbacks
In the terminate case, if we don't return early then we will delete the
FEX thread object but the actual terminate call that follows will
fail, leaving the thread in a bad state.
2025-08-12 00:51:00 +01:00
Billy Laws 6f5d4fbf30 Windows: Avoid potential deadlocks is a thread is terminated while in the JIT
Try to suspend the thread before termination to ensure no important JIT locks
are held that could be left locked if the thread is compiling code etc.
2025-08-12 00:51:00 +01:00
Billy Laws 8ad1ea499d Windows: Avoid a double-free if a thread is terminated twice
Unlikely to show up in any real code, but if an attempt is made to terminate
a terminated thread we should handle this gracefully.
2025-08-12 00:51:00 +01:00
Billy Laws 4df17abea7 Windows: Avoid potential race between thread init and termination
With the prior locking, we could try to terminate a thread while it
was initializing and perhaps holding global locks leading to deadlock.
2025-08-12 00:51:00 +01:00
Billy Laws 972c8a97dc ARM64EC: Disable syscall callbacks when handling OvercommitTracker AVs
We don't need to be informed of our own mappings, avoids a deadlock that
occurs if the ThreadCreationMutex lifetime is extended when creating a
thread.
2025-08-12 00:50:51 +01:00
Ryan Houdek 135d17fbd3 FEXCore: Fixes 16-bit operand size Enter
We weren't handling this edge case. Noticed it while looking at some
other test failures.
2025-08-11 16:50:30 -07:00
LC 095b2a8926 Merge pull request #4772 from Sonicadvance1/fix_3dnow_test
unittests/3DNow: Fix femms test
2025-08-09 17:19:17 -04:00
Alyssa Anne Rosenzweig 297677c2c6 Merge pull request #4771 from alyssarosenzweig/opts
Local JIT translation time wins
2025-08-09 09:44:23 -04:00
Alyssa Rosenzweig 2a00f23459 JIT: replace JumpTargets map with vector
Hashmaps are super expensive and there's no reason not to use a vector - we
already have compact block IDs so we don't benefit from the sparseness. Huge win
for very little effort.

Spotted when profiling FEX. CondJump() in the JIT was almost 4% of our time (?!)
and all because of map slowness. Easy fix.

Difference at 95.0% confidence
	-0.0196494 +/- 0.00194956
	-3.92827% +/- 0.389753%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Alyssa Rosenzweig 68ad916a91 Addressing: remove closures from SelectAddressMode
this is hot and closures make the assembly harder to reason about.

No difference at 95%.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Alyssa Rosenzweig 0f3883152b RegisterAllocationPass: call GetRAArgs less
..and be consistent about signedness.

clang doesn't seem able to do this itself.

Difference at 95.0% confidence, n=100
	-0.272326% +/- 0.242353%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-08 15:46:35 -04:00
Ryan Houdek 60baee4020 unittests/3DNow: Fix femms test
We are onyl wanting to test the FTW, which is 16-bit but we were reading
a whole 32-bits. The upper 16-bits was ending up in the reserved
section, which is undefined as to what it contains. In FEX it'll contain
0 but on a Zen 4 CPU it contains 0xFFFF.

Only read the bits we care about.
2025-08-08 12:08:52 -07:00
Ryan Houdek 5a7ff4780f Merge pull request #4770 from bylaws/arm64ec
InvalidationTracker: Disable mono hacks when multiblock is disabled
2025-08-07 13:59:33 -07:00
Billy Laws dd2845401c InvalidationTracker: Disable mono hacks when multiblock is disabled 2025-08-07 21:27:04 +01:00
Ryan Houdek a47c947065 Merge pull request #4768 from bylaws/memover
Avoid OOB reads for mem sources in some float conversion ops
2025-08-07 09:55:20 -07:00
Billy Laws 203ac51681 Update InstCountCI 2025-08-07 14:36:12 +01:00
Billy Laws a7986993bf InstCountCI: Add tests for AVX float conversion ops with mem src 2025-08-07 14:36:12 +01:00
Billy Laws ce59d40570 Vector: Avoid OOB reads in (v)cvt(t)s{s,d}2s{s,d} with mem src 2025-08-07 14:36:12 +01:00
Billy Laws 19f33e22d7 AVX128: Avoid OOB reads in vcvtsi2s{s,d} with mem src 2025-08-07 14:36:12 +01:00
Ryan Houdek 7d818e18be Merge pull request #4767 from bylaws/steamy
Windows: Ignore MEM_RESET(_UNDO) allocation protection notifications
2025-08-06 18:33:35 -07:00
Billy Laws 52af0aceea Windows: Ignore MEM_RESET(_UNDO) allocation protection notifications
When these are set the allocation permissions aren't actually updated.
2025-08-07 02:05:35 +01:00
Billy Laws cd1cf6f615 ARM64EC: Drop redundant ThreadCreationMutex locks 2025-08-07 02:04:48 +01:00
Ryan Houdek 2f925aeb22 Merge pull request #4764 from alyssarosenzweig/opt/drop-constprop-2
Drop ConstProp
2025-08-06 18:03:06 -07:00
Alyssa Rosenzweig f99b42da5e InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:42:36 -04:00
Alyssa Rosenzweig 40beef061f ConstProp: drop pass
now obsolete!

Results for the whole series are excellent:

Difference at 95.0% confidence
	-3.97603% +/- 0.254656%

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f0a434c167 IR: inline as we go
Augment IREmitter to inline constants as we generate code, rather than
needing a later clean up pass. This replaces the last function of ConstProp.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 92b66dbf17 IR: describe immediate inlining in json
This drops some register class validation since InlineConstants don't have a
register class.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f336261d6b IR: introduce & use more add helpers
so we can stash all the opts.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig c6324b83ea IR: introduce & use constant add helpers
more ergonomic and gives us a place to stash more logic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 54873992ab OpcodeDispatcher: optimize storecontext(0) in dispatcher
It's actually slightly less code to do it here than ConstProp, lol.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 38d568a807 OpcodeDispatcher: rework SelectCC
inline as we go, while trimming the internal interfaces.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig dcef1212c6 OpcodeDispatcher: add and use NZCVSelect01 helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 2030b70ff8 OpcodeDispatcher: do Select 0/1 inline in dispatcher
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig 41a188eaec OpcodeDispatcher: inline entrypoint offsets manually
it's literally less code with some helpers.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Alyssa Rosenzweig f3ddaf455c OpcodeDispatcher: drop dead constructor
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-06 20:40:49 -04:00
Ryan Houdek aa979e4bc7 Merge pull request #4728 from bylaws/monohacks
Implement mono specific hacks
2025-08-06 15:51:35 -07:00
Ryan Houdek baddc56f01 Merge pull request #4766 from bylaws/3dnow
Disable 3DNow by default on WOW64 FEX
2025-08-06 15:30:27 -07:00
Billy Laws 9836e43d10 Update InstCountCI 2025-08-06 22:39:44 +01:00
Billy Laws c8cecd9f46 Frontend: Adjust forward branch distance limit
Now executable page tracking is implemented, this limit is technically
unnecessary and effectively never hit in normal code. Keep a reasonable
limit however to avoid accidentally inlining tail calls and exploring
dead branches in obfuscated code.
2025-08-06 22:39:17 +01:00
Billy Laws 38c59a44de Frontend: Disable multiblock across calls in mono JIT code 2025-08-06 22:39:17 +01:00
Billy Laws fddc86c2b1 Windows: Opportunistically detect and replace the mono backpatcher 2025-08-06 22:39:17 +01:00
Billy Laws 0dddc9b5cb InvalidationTrack: Support disabling SMC detection 2025-08-06 22:39:17 +01:00
Billy Laws 8b14bd4e87 FEXCore: Fix broken RemoveCustomIREntrypoint
Unused, but would have crashed prior due to providing a nullptr thread.
2025-08-06 22:39:17 +01:00
Billy Laws 5c77969e83 Frontend: Force full SMC detection for mono jump thunk callsites
See IsBranchMonoTailcall
2025-08-06 22:39:17 +01:00
Billy Laws e049596252 FEXCore: Use MonoBackpatcherWrite for XCHG ops in the mono backpatcher block 2025-08-06 22:39:17 +01:00
Billy Laws afa1327242 FEXCore: Implement a write+code invalidate IR op for mono SMC 2025-08-06 22:39:17 +01:00
Billy Laws da3b7f5a41 FEXCore: Add an option to enable future mono-specific hacks 2025-08-06 22:39:17 +01:00
Billy Laws 100f61ee24 FEXCore: Allow the frontend to force full SMC detection for a block 2025-08-06 22:39:17 +01:00
Billy Laws dcd2794ff5 FEXCore: Make ThreadRemoveCodeEntryFromJit invalidate over all threads
Avoids redundant CompileBlock hits and generally easier to reason about.
2025-08-06 22:39:17 +01:00
Billy Laws 8d1d3fe12b Windows: Implement SyscallHandler code range invalidation 2025-08-06 22:39:17 +01:00
Billy Laws 87ae0f058a LinuxEmulation: Implement SyscallHandler code range invalidation 2025-08-06 22:39:17 +01:00
Billy Laws 7e1be6625f SyscallHandler: Add a method for invalidating a guest code range 2025-08-06 22:39:17 +01:00
Billy Laws 0c6fcd1678 FEXCore: Add function to recover the current block entrypoint 2025-08-06 22:39:17 +01:00
Billy Laws 37ac46273d Windows: Print extra debug info on image map 2025-08-06 22:39:17 +01:00
Billy Laws 1e21416ccb Disable 3DNow by default on WOW64 FEX 2025-08-06 22:30:38 +01:00
Ryan Houdek e2353959c1 Merge pull request #4755 from ChanthMiao/fix/wine32-v4l2
LinuxEmulation: add basic supports for v4l2 driver of wine32.
2025-08-05 21:36:13 -07:00
LC 61d3d5b068 Merge pull request #4763 from Sonicadvance1/fix_hostrunner_bug
HostRunner: Fixes bug with saving CSFSGS
2025-08-05 23:23:53 -04:00
Ryan Houdek 572e5ed9c6 HostRunner: Fixes bug with saving CSFSGS
We were failing to save the entire 64-bit register, cutting off the top
32-bits which contain SS and CS. This ...happened to not cause issues
but is a bug that I encountered with far jump handling.
2025-08-05 19:45:53 -07:00
Ryan Houdek 4dbfa2febb Merge pull request #4741 from pmatos/fninit_tests
asm_tests: Tests for fninit
2025-08-05 19:29:58 -07:00
Changwei Miao 0293d2027d LinuxEmulation: add basic supports for v4l2 driver of wine32.
This patch does cover up all v4l2 syscalls in i386. It just works
well with the poor v4l2 driver of wine32.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-08-06 04:49:49 +08:00
LC c28b94890c Merge pull request #4761 from neobrain/fix_debug_strace_build
Syscalls: Fix DEBUG_STRACE build
2025-08-05 11:17:19 -04:00
LC 63a910d2fc Merge pull request #4760 from Sonicadvance1/vex_log
Frontend: Remove log about VEX map_select
2025-08-05 11:15:53 -04:00
Tony Wasserka fbef5c6a6a Syscalls: Fix DEBUG_STRACE build 2025-08-05 15:51:34 +02:00
Ryan Houdek 674b6e9f43 Frontend: Remove log about VEX map_select
During multiblock code discovery this fires a lot and it isn't
interesting. Just remove the log, it'll SIGILL correctly if it actually
hits.
2025-08-04 15:46:36 -07:00
Ryan Houdek 7897b6ad55 Merge pull request #4759 from alyssarosenzweig/ra/simplify
RegisterAllocationPass: simplify next-use logic
2025-08-04 14:59:13 -07:00
Alyssa Rosenzweig 6769b54e1c unittests: add blake3 test
this provokes RA spilling and hit an assertion fail on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 17:06:15 -04:00
Alyssa Rosenzweig 734ab4429f RegisterAllocationPass: fix SRA spilling corner
I hate this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 17:06:15 -04:00
Ryan Houdek 5db6a64f3f Merge pull request #4758 from alyssarosenzweig/opt/x87ftw-2
OpcodeDispatcher: optimize SetX87FTW
2025-08-04 13:14:57 -07:00
Alyssa Rosenzweig dcc458ea94 RegisterAllocationPass: simplify next-use logic
I doubt this will fix the regression but it might make it easier to identify.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:31:08 -04:00
Alyssa Rosenzweig 7354e7e3c9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:13:10 -04:00
Alyssa Rosenzweig 07ef765f7e OpcodeDispatcher: optimize SetX87FTW
In b8dd5d95b ("OpcodeDispatcher: optimize X87FTWTag"), we optimized
X87FTWTag using an efficient Morton interleave operation. Here, we do the
inverse, optimizing SetX87FTW using an efficient Morton deinterleave
operation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:13:10 -04:00
Alyssa Rosenzweig c58ad8b593 IR: add AndShift op
will use it for next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 14:09:20 -04:00
Alyssa Rosenzweig 31091c1053 RegisterAllocationPass: remove some indentation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-08-04 08:16:28 -04:00
Paulo Matos c95f9f5e6a instcountci: Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:32 +02:00
Paulo Matos 93c3eb064d asm_tests: Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:21 +02:00
Paulo Matos b13162c5bf asm_tests: Division by zero does not set IE flag 2025-07-30 10:50:20 +02:00
Paulo Matos ee1725684a Implementation of FCNLEX to clear IE flag 2025-07-30 10:50:16 +02:00
Paulo Matos a07b234659 asm_tests: Tests for fninit
NFC.
These tests existed for 32bits but were missing for 64bits and reduced precision.
Adjust 32bit test to match new tests.
2025-07-30 09:54:08 +02:00
623 changed files with 85916 additions and 58235 deletions

No files matched your search

+3
View File
@@ -7,3 +7,6 @@ FEXCore/Source/Interface/Core/X86Tables/*
# Inline headers with list-like content that can't be processed individually
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/SyscallsNames.inl
Source/Tools/LinuxEmulation/LinuxSyscalls/x*/Ioctl/*.inl
# Include files in unittests
unittests/*ASM/Includes/*.inc
+2
View File
@@ -20,3 +20,5 @@
# Whole-tree reformat with clang-format-19
5267cde60e7642852d18f20ae8568643bb5293d5
# Minor reformat with clang-format-19
9fdd96af61c969cb5732471223f00eda64b7a069
+4 -1
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
build_plus_test:
@@ -33,7 +34,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
@@ -136,6 +136,9 @@ jobs:
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
env:
# These tests require non-portable install due to thunks.
FEX_PORTABLE: 0
run: cmake --build . --config $BUILD_TYPE --target fex_linux_tests_all
- name: FEXLinuxTests Results move
+1 -1
View File
@@ -20,6 +20,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
glibc_fault_test:
@@ -40,7 +41,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
+1 -1
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
hostrunner_tests:
@@ -33,7 +34,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
-1
View File
@@ -33,7 +33,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
+1 -2
View File
@@ -48,7 +48,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
@@ -78,7 +77,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
-1
View File
@@ -60,7 +60,6 @@ jobs:
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
# Using --diff_from_common_commit option available in clang-format-19
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
+1 -1
View File
@@ -13,6 +13,7 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_PORTABLE: 1
jobs:
vixl_simulator:
@@ -34,7 +35,6 @@ jobs:
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
+2 -2
View File
@@ -46,12 +46,12 @@ jobs:
- name: Configure CMake arm64ec
shell: bash
working-directory: ${{runner.workspace}}/build_arm64ec
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Configure CMake wow64
shell: bash
working-directory: ${{runner.workspace}}/build_wow64
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=/usr
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=/usr -DBUILD_TESTING=False -DCMAKE_INSTALL_PREFIX=/usr
- name: Build arm64ec
working-directory: ${{runner.workspace}}/build_arm64ec
+3
View File
@@ -46,3 +46,6 @@
[submodule "External/tracy"]
path = External/tracy
url = https://github.com/wolfpld/tracy
[submodule "External/range-v3"]
path = External/range-v3
url = https://github.com/ericniebler/range-v3.git
+19 -55
View File
@@ -4,7 +4,6 @@ project(FEX C CXX ASM)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(BUILD_FEXCONFIG "Build FEXConfig" TRUE)
@@ -304,7 +303,8 @@ set (CMAKE_LINKER_FLAGS_RELEASE "${CMAKE_LINKER_FLAGS_RELEASE} -fomit-frame-poin
include_directories(External/robin-map/include/)
if (BUILD_TESTS OR ENABLE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
include(CTest)
if (BUILD_TESTING OR ENABLE_VIXL_DISASSEMBLER OR ENABLE_VIXL_SIMULATOR)
add_subdirectory(External/vixl/)
include_directories(SYSTEM External/vixl/src/)
endif()
@@ -319,7 +319,7 @@ if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
find_package(Python 3.9 REQUIRED COMPONENTS Interpreter)
set(BUILD_SHARED_LIBS OFF)
@@ -335,7 +335,7 @@ endif()
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
if (BUILD_TESTING)
find_package(Catch2 3 QUIET)
if (NOT Catch2_FOUND)
add_subdirectory(External/Catch2/)
@@ -345,6 +345,9 @@ if (BUILD_TESTS)
endif()
include(Catch)
else ()
# Override any previously generated test list to avoid running stale test binaries
file(GENERATE OUTPUT CTestTestfile.cmake CONTENT "# No tests since BUILD_TESTING is disabled")
endif()
find_package(fmt QUIET)
@@ -354,6 +357,12 @@ if (NOT fmt_FOUND)
add_subdirectory(External/fmt/)
endif()
find_package(range-v3 QUIET)
if (NOT range-v3_FOUND)
add_subdirectory(External/range-v3/)
target_compile_definitions(range-v3 INTERFACE RANGES_DISABLE_DEPRECATED_WARNINGS)
endif()
add_subdirectory(External/tiny-json/)
include_directories(External/tiny-json/)
@@ -449,13 +458,8 @@ endif()
add_compile_options(-Wall)
include(CTest)
if (BUILD_TESTS)
if (BUILD_TESTING)
message(STATUS "Unit tests are enabled")
if (NOT BUILD_TESTING)
# CMake checks this variable before generating CTestTestfile.cmake
message(SEND_ERROR "Unit tests require BUILD_TESTING to be enabled")
endif()
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
if (TEST_JOB_COUNT)
@@ -486,10 +490,11 @@ file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.js
# Any application configuration json file gets installed
foreach(CONFIG_SRC ${CONFIG_SOURCES})
install(FILES ${CONFIG_SRC}
DESTINATION ${DATA_DIRECTORY}/)
DESTINATION ${DATA_DIRECTORY}/
COMPONENT Runtime)
endforeach()
if (BUILD_TESTS)
if (BUILD_TESTING)
add_subdirectory(unittests/)
endif()
@@ -550,6 +555,7 @@ if (BUILD_THUNKS)
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)"
DEPENDS guest-libs
COMPONENT Runtime
)
install(
@@ -559,6 +565,7 @@ if (BUILD_THUNKS)
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32
)"
DEPENDS guest-libs-32
COMPONENT Runtime
)
add_custom_target(uninstall_guest-libs
@@ -600,46 +607,3 @@ if (OVERRIDE_VERSION STREQUAL "detect")
else()
set(GIT_DESCRIBE_STRING "FEX-${OVERRIDE_VERSION}")
endif()
# Parse the version here
# Change something like `FEX-2106.1-76-<hash>` in to a list
string(REPLACE "-" ";" DESCRIBE_LIST ${GIT_DESCRIBE_STRING})
# Extract the `2106.1` element
list(GET DESCRIBE_LIST 1 DESCRIBE_LIST)
# Change `2106.1` in to a list
string(REPLACE "." ";" DESCRIBE_LIST ${DESCRIBE_LIST})
# Calculate list size
list(LENGTH DESCRIBE_LIST LIST_SIZE)
# Pull out the major version
list(GET DESCRIBE_LIST 0 FEX_VERSION_MAJOR)
# Minor version only exists if there is a .1 at the end
# eg: 2106 versus 2106.1
if (LIST_SIZE GREATER 1)
list(GET DESCRIBE_LIST 1 FEX_VERSION_MINOR)
endif()
# Package creation
set (CPACK_GENERATOR "DEB")
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_PACKAGE_FILE_NAME "${CPACK_PACKAGE_NAME}-${GIT_DESCRIBE_STRING}_${CMAKE_SYSTEM_PROCESSOR}")
set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.com>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "${CPACK_DEBIAN_PACKAGE_CONFLICTS}, qemu-user-static")
endif()
include (CPack)
+6 -6
View File
@@ -174,7 +174,7 @@ public:
// Logical immediate
void and_(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
and_(s, rd, rn, n, immr, imms);
}
@@ -185,7 +185,7 @@ public:
void ands(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
ands(s, rd, rn, n, immr, imms);
}
@@ -196,14 +196,14 @@ public:
void orr(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
orr(s, rd, rn, n, immr, imms);
}
void eor(ARMEmitter::Size s, ARMEmitter::Register rd, ARMEmitter::Register rn, uint64_t Imm) {
uint32_t n, immr, imms;
[[maybe_unused]] const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
const auto IsImm = IsImmLogical(Imm, RegSizeInBits(s), &n, &imms, &immr);
LOGMAN_THROW_A_FMT(IsImm, "Couldn't encode immediate to logical op");
eor(s, rd, rn, n, immr, imms);
}
@@ -333,7 +333,7 @@ public:
bfi(s, rd, Reg::zr, lsb, width);
}
void bfxil(ARMEmitter::Size s, Register rd, Register rn, uint32_t lsb, uint32_t width) {
[[maybe_unused]] const auto reg_size_bits = RegSizeInBits(s);
const auto reg_size_bits = RegSizeInBits(s);
const auto lsb_p_width = lsb + width;
LOGMAN_THROW_A_FMT(width >= 1, "bfxil needs width >= 1");
@@ -977,7 +977,7 @@ private:
}
void xbfiz_helper(bool is_signed, ARMEmitter::Size s, Register rd, Register rn, uint32_t lsb, uint32_t width) {
[[maybe_unused]] const auto lsb_p_width = lsb + width;
const auto lsb_p_width = lsb + width;
const auto reg_size_bits = RegSizeInBits(s);
LOGMAN_THROW_A_FMT(lsb_p_width <= reg_size_bits, "lsb + width ({}) must be <= {}. lsb={}, width={}", lsb_p_width, reg_size_bits, lsb, width);
+1 -2
View File
@@ -2244,8 +2244,7 @@ public:
template<IsQOrDRegister T>
void movi(SubRegSize size, T rd, uint64_t Imm, uint16_t Shift = 0) {
LOGMAN_THROW_A_FMT(size == SubRegSize::i8Bit || size == SubRegSize::i16Bit || size == SubRegSize::i32Bit ||
size == SubRegSize::i64Bit,
LOGMAN_THROW_A_FMT(size == SubRegSize::i8Bit || size == SubRegSize::i16Bit || size == SubRegSize::i32Bit || size == SubRegSize::i64Bit,
"Unsupported movi size");
uint32_t cmode;
+1
View File
@@ -12,6 +12,7 @@
#include <CodeEmitter/Registers.h>
#include <array>
#include <bit>
#include <cstdint>
#include <utility>
#include <type_traits>
+23 -17
View File
@@ -1541,7 +1541,7 @@ public:
void sqincp(SubRegSize size, XRegister rdn, PRegister pm) {
SVEIncDecPredicateCountScalar(0, 1, 0b10, 0b00, size, rdn, pm);
}
void sqincp(SubRegSize size, XRegister rdn, PRegister pm, [[maybe_unused]] WRegister wn) {
void sqincp(SubRegSize size, XRegister rdn, PRegister pm, WRegister wn) {
LOGMAN_THROW_A_FMT(rdn.Idx() == wn.Idx(), "rdn and wn must be the same");
SVEIncDecPredicateCountScalar(0, 1, 0b00, 0b00, size, rdn, pm);
}
@@ -1554,7 +1554,7 @@ public:
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm) {
SVEIncDecPredicateCountScalar(0, 1, 0b10, 0b10, size, rdn, pm);
}
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm, [[maybe_unused]] WRegister wn) {
void sqdecp(SubRegSize size, XRegister rdn, PRegister pm, WRegister wn) {
LOGMAN_THROW_A_FMT(rdn.Idx() == wn.Idx(), "rdn and wn must be the same");
SVEIncDecPredicateCountScalar(0, 1, 0b00, 0b10, size, rdn, pm);
}
@@ -3296,7 +3296,7 @@ private:
const auto log2_size_bytes = FEXCore::ilog2(size_bytes);
// We can index up to 512-bit registers with dup
[[maybe_unused]] const auto max_index = (64U >> log2_size_bytes) - 1;
const auto max_index = (64U >> log2_size_bytes) - 1;
LOGMAN_THROW_A_FMT(Index <= max_index, "dup index ({}) too large. Must be within [0, {}].", Index, max_index);
// imm2:tsz make up a 7 bit wide field, with each increasing element size
@@ -3326,7 +3326,7 @@ private:
uint32_t shift = 0;
if (!is_uint8_imm) {
[[maybe_unused]] const bool is_uint16_imm = (imm >> 16) == 0;
const bool is_uint16_imm = (imm >> 16) == 0;
LOGMAN_THROW_A_FMT(is_uint16_imm, "Immediate ({}) must be a 16-bit value within [256, 65280]", imm);
LOGMAN_THROW_A_FMT((imm % 256) == 0, "Immediate ({}) must be a multiple of 256", imm);
@@ -4152,7 +4152,7 @@ private:
const auto& op_data = mem_op.MetaType.ScalarVectorType;
const bool is_scaled = op_data.scale != 0;
[[maybe_unused]] const auto msize_value = FEXCore::ToUnderlying(msize);
const auto msize_value = FEXCore::ToUnderlying(msize);
LOGMAN_THROW_A_FMT(op_data.scale == 0 || op_data.scale == msize_value, "scale may only be 0 or {}", msize_value);
@@ -4266,7 +4266,7 @@ private:
const auto msize_value = FEXCore::ToUnderlying(msize);
const auto msize_bytes = 1U << msize_value;
[[maybe_unused]] const auto imm_limit = (32U << msize_value) - msize_bytes;
const auto imm_limit = (32U << msize_value) - msize_bytes;
const auto imm = mem_op.MetaType.VectorImmType.Imm;
const auto imm_to_encode = imm >> msize_value;
@@ -4332,8 +4332,8 @@ private:
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_A_FMT((imm % num_regs) == 0, "Offset must be a multiple of {}", num_regs);
[[maybe_unused]] const auto min_offset = -8 * num_regs;
[[maybe_unused]] const auto max_offset = 7 * num_regs;
const auto min_offset = -8 * num_regs;
const auto max_offset = 7 * num_regs;
LOGMAN_THROW_A_FMT(imm >= min_offset && imm <= max_offset,
"Invalid load/store offset ({}). Offset must be a multiple of {} and be within [{}, {}]", imm, num_regs, min_offset,
max_offset);
@@ -4440,8 +4440,8 @@ private:
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
const auto esize = static_cast<int>(16 << ssz);
[[maybe_unused]] const auto max_imm = (esize << 3) - esize;
[[maybe_unused]] const auto min_imm = -(max_imm + esize);
const auto max_imm = (esize << 3) - esize;
const auto min_imm = -(max_imm + esize);
LOGMAN_THROW_A_FMT((imm % esize) == 0, "imm ({}) must be a multiple of {}", imm, esize);
LOGMAN_THROW_A_FMT(imm >= min_imm && imm <= max_imm, "imm ({}) must be within [{}, {}]", imm, min_imm, max_imm);
@@ -4485,7 +4485,7 @@ private:
const auto msize_value = FEXCore::ToUnderlying(msize);
const auto data_size_bytes = 1U << msize_value;
[[maybe_unused]] const auto max_imm = (64U << msize_value) - data_size_bytes;
const auto max_imm = (64U << msize_value) - data_size_bytes;
LOGMAN_THROW_A_FMT((imm % data_size_bytes) == 0 && imm <= max_imm, "imm must be a multiple of {} and be within [0, {}]",
data_size_bytes, max_imm);
@@ -4861,7 +4861,7 @@ private:
"64-bit variants may only use Zm between z0-z15");
const auto Underlying = FEXCore::ToUnderlying(size);
[[maybe_unused]] const uint32_t IndexMax = (16 / (1U << Underlying)) - 1;
const uint32_t IndexMax = (16 / (1U << Underlying)) - 1;
LOGMAN_THROW_A_FMT(index <= IndexMax, "Index must be within 0-{}", IndexMax);
// Can be bit 20 or 19 depending on whether or not the element size is 64-bit.
@@ -5117,14 +5117,15 @@ private:
requires (std::is_same_v<T, float> || std::is_same_v<T, double>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
// Determines if a floating-point value is capable of being converted
// into an 8-bit immediate. See pseudocode definition of VFPExpandImm
// in ARM A-profile reference manual for a general overview of how this was derived.
template<typename T>
requires (std::is_same_v<T, float> || std::is_same_v<T, double>)
[[nodiscard, maybe_unused]]
[[nodiscard]]
static bool IsValidFPValueForImm8(T value) {
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
const uint64_t bits = std::bit_cast<FloatToEquivalentUInt<T>>(value);
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
static constexpr std::array mantissa_masks {
@@ -5162,12 +5163,15 @@ private:
return true;
}
#endif
protected:
static uint32_t FP32ToImm8(float value) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint32_t>(value);
const auto bits = std::bit_cast<uint32_t>(value);
const auto sign = (bits & 0x80000000) >> 24;
const auto expb2 = (bits & 0x20000000) >> 23;
const auto b5_to_0 = (bits >> 19) & 0x3F;
@@ -5176,9 +5180,11 @@ protected:
}
static uint32_t FP64ToImm8(double value) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value), "Value ({}) cannot be encoded into an 8-bit immediate", value);
#endif
const auto bits = FEXCore::BitCast<uint64_t>(value);
const auto bits = std::bit_cast<uint64_t>(value);
const auto sign = (bits & 0x80000000'00000000) >> 56;
const auto expb2 = (bits & 0x20000000'00000000) >> 55;
const auto b5_to_0 = (bits >> 48) & 0x3F;
@@ -5202,7 +5208,7 @@ private:
uint32_t shift = 0;
if (!is_int8_imm) {
const int32_t imm16_limit = 32768;
[[maybe_unused]] const bool is_int16_imm = -imm16_limit <= imm && imm < imm16_limit;
const bool is_int16_imm = -imm16_limit <= imm && imm < imm16_limit;
LOGMAN_THROW_A_FMT(is_int16_imm, "Immediate ({}) must be a 16-bit value within [-32768, 32512]", imm);
LOGMAN_THROW_A_FMT((imm % 256) == 0, "Immediate ({}) must be a multiple of 256", imm);
+2 -2
View File
@@ -30,7 +30,7 @@ public:
const uint32_t SizeImm = FEXCore::ToUnderlying(size);
const uint32_t IndexShift = SizeImm + 1;
const uint32_t ElementSize = 1U << SizeImm;
[[maybe_unused]] const uint32_t MaxIndex = 128U / (ElementSize * 8);
const uint32_t MaxIndex = 128U / (ElementSize * 8);
LOGMAN_THROW_A_FMT(Index < MaxIndex, "Index too large. Index={}, Max Index: {}", Index, MaxIndex);
@@ -1381,7 +1381,7 @@ private:
void ASIMDScalarXIndexedElement(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i8Bit, "Scalar size must not be 8-bit");
[[maybe_unused]] const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
LOGMAN_THROW_A_FMT(index < invalid_bound, "Index ({}) must be within [0-{}]", index, invalid_bound - 1);
uint32_t Instr = 0b0101'1111'0000'0000'0000'0000'0000'0000;
+6 -58
View File
@@ -41,7 +41,6 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
[[maybe_unused]] constexpr auto kDRegSize = 64;
constexpr auto kWRegSize = 32;
constexpr auto kXRegSize = 64;
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) || (width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
@@ -129,8 +128,8 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// Compute the repeat distance d, and set up a bitmask covering the basic
// unit of repetition (i.e. a word with the bottom d bits set). Also, in all
// of these cases the N bit of the output will be zero.
clz_a = CountLeadingZeros(a, kXRegSize);
int clz_c = CountLeadingZeros(c, kXRegSize);
clz_a = std::countl_zero(a);
int clz_c = std::countl_zero(c);
d = clz_a - clz_c;
mask = ((UINT64_C(1) << d) - 1);
out_n = 0;
@@ -151,7 +150,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// of set bits in our word, meaning that we have the trivial case of
// d == 64 and only one 'repetition'. Set up all the same variables as in
// the general case above, and set the N bit in the output.
clz_a = CountLeadingZeros(a, kXRegSize);
clz_a = std::countl_zero(a);
d = 64;
mask = ~UINT64_C(0);
out_n = 1;
@@ -159,7 +158,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
}
// If the repeat period d is not a power of two, it can't be encoded.
if (!IsPowerOf2(d)) {
if (!std::has_single_bit(uint32_t(d))) {
return false;
}
@@ -179,7 +178,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
static const uint64_t multipliers[] = {
0x0000000000000001UL, 0x0000000100000001UL, 0x0001000100010001UL, 0x0101010101010101UL, 0x1111111111111111UL, 0x5555555555555555UL,
};
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
uint64_t multiplier = multipliers[std::countl_zero(uint64_t(d)) - 57];
uint64_t candidate = (b - a) * multiplier;
if (value != candidate) {
@@ -194,7 +193,7 @@ static bool IsImmLogical(uint64_t value, unsigned width, unsigned* n = nullptr,
// Count the set bits in our basic stretch. The special case of clz(0) == -1
// makes the answer come out right for stretches that reach the very top of
// the word (e.g. numbers like 0xffffc00000000000).
int clz_b = (b == 0) ? -1 : CountLeadingZeros(b, kXRegSize);
int clz_b = (b == 0) ? -1 : std::countl_zero(b);
int s = clz_a - clz_b;
// Decide how many bits to rotate right by, to put the low bit of that basic
@@ -285,11 +284,6 @@ INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
private:
template<typename V>
static inline bool IsPowerOf2(V value) {
return (value != 0) && ((value & (value - 1)) == 0);
}
// Some compilers dislike negating unsigned integers,
// so we provide an equivalent.
template<typename T>
@@ -302,50 +296,4 @@ static inline uint64_t LowestSetBit(uint64_t value) {
return value & UnsignedNegate(value);
}
template<typename V>
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
#if COMPILER_HAS_BUILTIN_CLZ
if (width == 32) {
return (value == 0) ? 32 : __builtin_clz(static_cast<unsigned>(value));
} else if (width == 64) {
return (value == 0) ? 64 : __builtin_clzll(value);
}
#endif
return CountLeadingZerosFallBack(value, width);
}
static inline int CountLeadingZerosFallBack(uint64_t value, int width) {
LOGMAN_THROW_A_FMT(IsPowerOf2(width) && (width <= 64), "Invalid width");
if (value == 0) {
return width;
}
int count = 0;
value = value << (64 - width);
if ((value & UINT64_C(0xffffffff00000000)) == 0) {
count += 32;
value = value << 32;
}
if ((value & UINT64_C(0xffff000000000000)) == 0) {
count += 16;
value = value << 16;
}
if ((value & UINT64_C(0xff00000000000000)) == 0) {
count += 8;
value = value << 8;
}
if ((value & UINT64_C(0xf000000000000000)) == 0) {
count += 4;
value = value << 4;
}
if ((value & UINT64_C(0xc000000000000000)) == 0) {
count += 2;
value = value << 2;
}
if ((value & UINT64_C(0x8000000000000000)) == 0) {
count += 1;
}
count += (value == 0);
return count;
}
public:
+4 -2
View File
@@ -4,7 +4,8 @@ file(GLOB GEN_CONFIG_SOURCES CONFIGURE_DEPENDS *.json.in)
# Any application configuration json file gets installed
foreach(CONFIG_SRC ${CONFIG_SOURCES})
install(FILES ${CONFIG_SRC}
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
DESTINATION ${DATA_DIRECTORY}/AppConfig/
COMPONENT Runtime)
endforeach()
# Any configuration file json file that needs to be generated
@@ -21,5 +22,6 @@ foreach(GEN_CONFIG_SRC ${GEN_CONFIG_SOURCES})
# Then install the configured json
install(
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
DESTINATION ${DATA_DIRECTORY}/AppConfig/
COMPONENT Runtime)
endforeach()
-3
View File
@@ -1,3 +0,0 @@
x86 and x86-64 Linux emulator
FEX allows you to run x86 applications on ARM64 Linux devices. It offers broad compatibility with both 32-bit and 64-bit binaries, and it can be used alongside Wine/Proton to play Windows games.
-18
View File
@@ -1,18 +0,0 @@
#!/bin/sh
set -e
update_binfmt() {
# Check for update-binfmts
command -v update-binfmts >/dev/null || return 0
# Setup binfmt_misc
update-binfmts --import FEX-x86
update-binfmts --import FEX-x86_64
}
# Install FEXInterpreter hardlink
# Needs to be done before setting up binfmt_misc
ln -f /usr/bin/FEXLoader /usr/bin/FEXInterpreter
if [ $(uname -m) = 'aarch64' ]; then
update_binfmt
fi
-17
View File
@@ -1,17 +0,0 @@
#!/bin/sh
set -e
update_binfmt() {
# Check for update-binfmts
command -v update-binfmts >/dev/null || return 0
# Uninstall
update-binfmts --unimport FEX-x86
update-binfmts --unimport FEX-x86_64
}
if [ $(uname -m) = 'aarch64' ]; then
update_binfmt
fi
# Remove FEXInterpreter hardlink
unlink /usr/bin/FEXInterpreter
-1
View File
@@ -1 +0,0 @@
activate-noawait ldconfig
+1 -1
View File
@@ -14,7 +14,7 @@ RUN mkdir build
ARG CC=clang-13
ARG CXX=clang++-13
RUN cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DUSE_LINKER=lld -DENABLE_LTO=True -DBUILD_TESTS=False -DENABLE_ASSERTIONS=False -G Ninja .
RUN cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DUSE_LINKER=lld -DENABLE_LTO=True -DBUILD_TESTING=False -DENABLE_ASSERTIONS=False -G Ninja .
RUN ninja
WORKDIR /FEX/build
+4 -2
View File
@@ -10,7 +10,8 @@ function(GenBinFmt Name)
# Then install the configured binfmt
install(
FILES ${CMAKE_BINARY_DIR}/Data/binfmts/${FMT_NAME}
DESTINATION ${CMAKE_INSTALL_PREFIX}/share/binfmts/)
DESTINATION ${CMAKE_INSTALL_PREFIX}/share/binfmts/
COMPONENT Runtime)
endfunction()
if (NOT USE_LEGACY_BINFMTMISC)
@@ -19,7 +20,8 @@ if (NOT USE_LEGACY_BINFMTMISC)
install(
FILES ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86.conf ${CMAKE_BINARY_DIR}/Data/binfmts/FEX-x86_64.conf
DESTINATION ${CMAKE_INSTALL_PREFIX}/lib/binfmt.d/)
DESTINATION ${CMAKE_INSTALL_PREFIX}/lib/binfmt.d/
COMPONENT Runtime)
else()
GenBinFmt(FEX-x86.in)
GenBinFmt(FEX-x86_64.in)
+1 -1
View File
@@ -1 +1 @@
:FEX-x86:M:0:\x7fELF\x01\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x03\x00:\xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff:@CMAKE_INSTALL_PREFIX@/bin/FEXInterpreter:POCF
:FEX-x86:M:0:\x7fELF\x01\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x03\x00:\xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff:@CMAKE_INSTALL_PREFIX@/bin/FEX:POCF
+1 -1
View File
@@ -1,5 +1,5 @@
package fex
interpreter @CMAKE_INSTALL_PREFIX@/bin/FEXInterpreter
interpreter @CMAKE_INSTALL_PREFIX@/bin/FEX
magic \x7fELF\x01\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x03\x00
offset 0
mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff
+1 -1
View File
@@ -1 +1 @@
:FEX-x86_64:M:0:\x7fELF\x02\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x3e\x00:\xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff:@CMAKE_INSTALL_PREFIX@/bin/FEXInterpreter:POCF
:FEX-x86_64:M:0:\x7fELF\x02\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x3e\x00:\xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff:@CMAKE_INSTALL_PREFIX@/bin/FEX:POCF
+1 -1
View File
@@ -1,5 +1,5 @@
package fex
interpreter @CMAKE_INSTALL_PREFIX@/bin/FEXInterpreter
interpreter @CMAKE_INSTALL_PREFIX@/bin/FEX
magic \x7fELF\x02\x01\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x3e\x00
offset 0
mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xff\xff
+3 -3
View File
@@ -2,8 +2,8 @@
let
toolchain = pkgs.fetchzip {
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250305/llvm-mingw-20250305-ucrt-ubuntu-20.04-aarch64.tar.xz";
sha256 = "sha256-cA03/ab9O61eO9+S2JzIXD4V0HzTXK5/AYyxW2d73Po=";
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250920/llvm-mingw-20250920-ucrt-ubuntu-22.04-aarch64.tar.xz";
sha256 = "sha256-LaojKjC8KzY+soW5u6eoDoXE3qtYk9Ejr7M3enTqRAE=";
};
cmakeToolchainFile = pkgs.substitute {
@@ -45,7 +45,7 @@ pkgs.mkShell {
fi
'';
# E.g. cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False
# E.g. cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTING=False
FEX_CMAKE_TOOLCHAIN_ARM64EC = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_CMAKE_TOOLCHAIN_WOW64 = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_MESON_CROSSFILE = "--cross-file ${mesonCrossFile}";
+1 -1
View File
@@ -18,4 +18,4 @@ then
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_WOW64 -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
cmake $FEX_CMAKE_TOOLCHAIN_WOW64 -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTING=False $@
+1 -1
View File
@@ -18,4 +18,4 @@ then
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTING=False $@
+1 -1
View File
@@ -14,4 +14,4 @@ fi
rm -rf unittests/FEXLinuxTests
set -o xtrace
cmake . $FEX_CMAKE_TOOLCHAINS -DBUILD_TESTS=ON -DBUILD_FEX_LINUX_TESTS=ON
cmake . $FEX_CMAKE_TOOLCHAINS -DBUILD_TESTING=ON -DBUILD_FEX_LINUX_TESTS=ON
-1
View File
@@ -214,7 +214,6 @@ class ClangFormatHelper(FormatHelper):
self.clang_fmt_path,
"--binary=clang-format-19",
"--diff",
"--diff_from_common_commit",
]
if args.start_rev and args.end_rev:
+1 -1
Vendored Submodule
+1
Submodule External/range-v3 added at ca1388fb9d.
+1 -1
+1 -1
View File
@@ -78,6 +78,6 @@ install (DIRECTORY include/FEXCore ${CMAKE_BINARY_DIR}/include/FEXCore
DESTINATION include
COMPONENT Development)
if (BUILD_TESTS)
if (BUILD_TESTING)
add_subdirectory(unittests/)
endif()
+11 -170
View File
@@ -118,41 +118,6 @@ def print_man_env_option(name, desc, default, no_json_key):
output_man.write("\\fBdefault:\\fR {0}\n".format(default))
output_man.write(".Pp\n\n")
def print_man_options(options):
output_man.write(".Sh OPTIONS\n")
output_man.write(".Bl -tag -width -indent\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
default = op_vals["Default"]
value_type = op_vals["Type"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (value_type == "str" or value_type == "strarray" or value_type == "strenum"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_option(
short,
long,
op_vals["Desc"],
default
)
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
output_man.write(", ".join(f"{enum_op_val}" for [_, enum_op_val] in Enums.items()))
output_man.write("\n.sp\n")
output_man.write(".El\n")
def print_man_environment(options):
output_man.write(".Sh ENVIRONMENT\n")
output_man.write(".Bl -tag -width -indent\n")
@@ -194,7 +159,7 @@ def print_man_environment_tail():
"By default FEX will look in {$HOME, $XDG_CONFIG_HOME}/.fex-emu/",
"This will override the full path",
"If FEX_PORTABLE is declared then relative paths are also supported",
"For FEXInterpreter: Relative to the FEXInterpreter binary",
"For FEX: Relative to the FEX binary",
"For WINE: Relative to %LOCALAPPDATA%"
],
"''", True)
@@ -208,7 +173,7 @@ def print_man_environment_tail():
"One must be careful with this option as it will override any applications that load with execve as well"
"If you need to support applications that execve then use FEX_APP_CONFIG_LOCATION instead"
"If FEX_PORTABLE is declared then relative paths are also supported",
"For FEXInterpreter: Relative to the FEXInterpreter binary",
"For FEX: Relative to the FEX binary",
"For WINE: Relative to %LOCALAPPDATA%"
],
"''", True)
@@ -227,8 +192,8 @@ def print_man_environment_tail():
"PORTABLE",
[
"Allows FEX to run without installation. Global locations for configuration and binfmt_misc are ignored.",
"For FEXInterpreter on Linux:",
"These files are instead read from <FEXInterpreterPath>/fex-emu/ by default.",
"For FEX on Linux:",
"These files are instead read from <FEXPath>/fex-emu/ by default.",
"For Arm64ec/Wow64 WINE builds:",
"These files are instead read from $LOCALAPPDATA/fex-emu/ by default.",
"For further customization, see FEX_APP_CONFIG_LOCATION and FEX_APP_DATA_LOCATION."
@@ -240,20 +205,12 @@ def print_man_header():
.Dt FEX
.Os Linux
.Sh NAME
.Nm FEXLoader
.Nm FEXInterpreter
.Nm FEX
.Nm FEXBash
.Nd Fast x86-64 and x86 emulation.
.Sh SYNOPSIS
.Nm
.Op options
.Op Ar --
.Ar Application
<args> ...
.Pp
.Nm FEXInterpreter
.Ar Application
<args> ...
.Ar <args> ...
.Pp
.Nm FEXBash
.Ar <args> ...
@@ -361,82 +318,6 @@ def print_config_option(type, group_name, json_name, default_value, short, choic
output_argloader.write("\n");
def print_argloader_options(options):
output_argloader.write("#ifdef BEFORE_PARSE\n")
output_argloader.write("#undef BEFORE_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray" or op_vals["Type"] == "strenum"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = "\"" + op_vals["TextDefault"] + "\""
short = None
choices = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if ("Choices" in op_vals):
choices = op_vals["Choices"]
print_config_option(
op_vals["Type"],
op_group,
op_key,
default,
short,
choices,
op_vals["Desc"])
output_argloader.write("\n")
output_argloader.write("#endif\n")
def print_parse_argloader_options(options):
output_argloader.write("#ifdef AFTER_PARSE\n")
output_argloader.write("#undef AFTER_PARSE\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
output_argloader.write("if (Options.is_set_by_user(\"{0}\")) {{\n".format(op_key))
value_type = op_vals["Type"]
NeedsString = False
conversion_func = "fextl::fmt::format(\"{}\", "
if ("ArgumentHandler" in op_vals):
NeedsString = True
conversion_func = "FEXCore::Config::Handler::{0}(".format(op_vals["ArgumentHandler"])
if (value_type == "str"):
NeedsString = True
conversion_func = "std::move("
if (value_type == "bool"):
# boolean values need a decimal specifier. Otherwise fmt prints strings.
conversion_func = "fextl::fmt::format(\"{:d}\", "
if (value_type == "strenum"):
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key, op_key))
elif (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
output_argloader.write("\tfor (auto iter = Array.begin(); iter != Array.end(); ++iter) {\n")
output_argloader.write("\t\tAppendStrArrayValue(FEXCore::Config::ConfigOption::CONFIG_{0}, *iter);\n".format(op_key.upper()))
output_argloader.write("\t}\n")
else:
if (NeedsString):
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
else:
output_argloader.write("\t{0} UserValue = Options.get(\"{1}\");\n".format(value_type, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{0}, {1}UserValue));\n".format(op_key.upper(), conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def print_parse_envloader_options(options):
output_argloader.write("#ifdef ENVLOADER\n")
output_argloader.write("#undef ENVLOADER\n")
@@ -447,13 +328,13 @@ def print_parse_envloader_options(options):
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key, op_key))
output_argloader.write("\tValue = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key))
output_argloader.write("}\n")
if ("ArgumentHandler" in op_vals):
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("\tValue = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
@@ -467,15 +348,15 @@ def print_parse_jsonloader_options(options):
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (KeyName == \"{0}\") {{\n".format(op_key))
output_argloader.write("\tSet(KeyOption, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View));\n".format(op_key, op_key, op_key))
output_argloader.write("\tSet(KeyOption, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View));\n".format(op_key, op_key))
output_argloader.write("}\n")
elif (value_type == "strarray"):
output_argloader.write("else if (KeyName == \"{0}\") {{\n".format(op_key))
output_argloader.write("\tAppendStrArrayValue(KeyOption, ConfigString);\n")
output_argloader.write("}\n")
assert op_key is not None, "No options found in JSONLOADER"
output_argloader.write("else {{\n".format(op_key))
output_argloader.write("Set(KeyOption, ConfigString);\n")
output_argloader.write("else {\n")
output_argloader.write("\tSet(KeyOption, ConfigString);\n")
output_argloader.write("}\n")
output_argloader.write("#endif\n")
@@ -517,41 +398,6 @@ def print_parse_enum_options(options):
output_argloader.write("#endif\n")
def check_for_duplicate_options(options):
short_map = []
long_map = []
# Spin through all the items and see if we have a duplicate option
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
short = None
long = op_key.lower()
long_invert = None
if ("ShortArg" in op_vals):
short = op_vals["ShortArg"]
if (op_vals["Type"] == "bool"):
long_invert = "no-" + long
# Check for short key duplication
if (short != None):
if (short in short_map):
raise Exception("Short config '{0}' for option '{1}' has duplicate entry!".format(short, op_key))
else:
short_map.append(short)
# Check for long key duplication
if (long in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long))
else:
long_map.append(long)
# Check for long key duplication
if (long_invert != None):
if (long_invert in long_map):
raise Exception("Long config '{0}' has duplicate entry!".format(long_invert))
else:
long_map.append(long_invert)
if (len(sys.argv) < 5):
sys.exit()
@@ -568,8 +414,6 @@ json_object = json.loads(json_text)
options = json_object["Options"]
unnamed_options = json_object["UnnamedOptions"]
check_for_duplicate_options(options)
# Generate config include file
output_file = open(output_filename, "w")
print_header()
@@ -581,7 +425,6 @@ output_file.close()
# Generate man file
output_man = open(output_man_page, "w")
print_man_header()
print_man_options(options)
print_man_environment(options)
print_man_tail()
@@ -589,8 +432,6 @@ output_man.close()
# Generate argument loader code
output_argloader = open(output_argumentloader_filename, "w")
print_argloader_options(options);
print_parse_argloader_options(options);
# Generate environment loader code
print_parse_envloader_options(options);
+93 -72
View File
@@ -58,9 +58,10 @@ class OpDefinition:
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
Arguments: list
EmitValidation: list
Desc: list
Inline: list[str]
Arguments: list[OpArgument]
EmitValidation: list[str]
Desc: list[str]
def __init__(self):
self.Name = None
@@ -91,19 +92,14 @@ class OpDefinition:
attrs = vars(self)
print(", ".join("%s: %s" % item for item in attrs.items()))
IRTypesToCXX = {}
CXXTypeToIR = {}
IROps = []
IRTypesToCXX: dict[str, IRType] = {}
CXXTypeToIR: dict[str, IRType] = {}
IROps: list[OpDefinition] = []
IROpNameMap = {}
IROpNameSet: set[str] = set()
def is_ssa_type(type):
if (type == "SSA" or
type == "GPR" or
type == "GPRPair" or
type == "FPR"):
return True
return False
def is_ssa_type(op_type: str):
return op_type in {"SSA", "GPR", "GPRPair", "FPR"}
def parse_irtypes(irtypes):
for op_key, op_val in irtypes.items():
@@ -218,11 +214,8 @@ def parse_ops(ops):
OpArg.DefaultInitializer = DefaultInit[1][:-1]
# If SSA type then we can generate validation for this op
if (OpArg.IsSSA and
(OpArg.Type == "GPR" or
OpArg.Type == "GPRPair" or
OpArg.Type == "FPR")):
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == InvalidClass || WalkFindRegClass({ArgName}) == {OpArg.Type}Class")
if OpArg.IsSSA and OpArg.Type in {"GPR", "GPRPair", "FPR"}:
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == RegClass::Invalid || WalkFindRegClass({ArgName}) == RegClass::{OpArg.Type}")
OpArg.Name = ArgName
OpArg.NameWithPrefix = NameWithPrefix
@@ -278,6 +271,12 @@ def parse_ops(ops):
if "TiedSource" in op_val:
OpDef.TiedSource = op_val["TiedSource"]
# Pad Inline out to the argument count
OpDef.Inline = [''] * len(OpDef.Arguments)
if "Inline" in op_val:
Value = op_val["Inline"]
OpDef.Inline[0:len(Value)] = Value
# Do some fixups of the data here
if len(OpDef.EmitValidation) != 0:
for i in range(len(OpDef.EmitValidation)):
@@ -289,22 +288,29 @@ def parse_ops(ops):
#OpDef.print()
# Error on duplicate op
if OpDef.Name in IROpNameMap:
if OpDef.Name in IROpNameSet:
ExitError("Duplicate Op defined! {}".format(OpDef.Name))
IROps.append(OpDef)
IROpNameMap[OpDef.Name] = 1
IROpNameSet.add(OpDef.Name)
# Print out enum values
def print_enums():
def print_enums(enums):
output_file.write("#ifdef IROP_ENUM\n")
output_file.write("enum IROps : uint16_t {\n")
for op in IROps:
output_file.write("\tOP_{},\n" .format(op.Name.upper()))
output_file.write("};\n")
for name, members in enums.items():
output_file.write(f"enum {name} {{\n")
for member in members:
if member:
output_file.write(f"\t{member}\n")
else:
output_file.write("\n")
output_file.write("};\n\n")
output_file.write("#undef IROP_ENUM\n")
output_file.write("#endif\n\n")
@@ -397,16 +403,16 @@ def print_ir_sizes():
// Make sure our array maps directly to the IROps enum
static_assert(IRSizes[IROps::OP_LAST] == -1ULL);
[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);
[[nodiscard]] inline size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const]] FEXCore::IR::RegClass GetRegClass(IROps Op);
[[nodiscard, gnu::const]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const]] int8_t TiedSource(IROps Op);
#undef IROP_SIZES
#endif
@@ -415,30 +421,29 @@ def print_ir_sizes():
def print_ir_reg_classes():
output_file.write("#ifdef IROP_REG_CLASSES_IMPL\n")
output_file.write("constexpr std::array<FEXCore::IR::RegisterClassType, IROps::OP_LAST + 1> IRRegClasses = {\n")
output_file.write("constexpr std::array<FEXCore::IR::RegClass, IROps::OP_LAST + 1> IRRegClasses = {\n")
for op in IROps:
if op.Name == "Last":
output_file.write("\tFEXCore::IR::InvalidClass,\n")
output_file.write("\tRegClass::Invalid,\n")
else:
Class = "Invalid"
if op.HasDest and op.DestType == None:
if op.HasDest and op.DestType is None:
ExitError("IR op {} has destination with no destination class".format(op.Name))
if op.HasDest and op.DestType == "SSA": # Special case SSA type
output_file.write("\tFEXCore::IR::ComplexClass,\n")
output_file.write("\tRegClass::Complex,\n")
elif op.HasDest:
output_file.write("\tFEXCore::IR::{}Class,\n".format(op.DestType))
output_file.write("\tRegClass::{},\n".format(op.DestType))
else:
# No destination so it has an invalid destination class
output_file.write("\tFEXCore::IR::InvalidClass, // No destination\n")
output_file.write("\tRegClass::Invalid, // No destination\n")
output_file.write("};\n\n")
output_file.write("// Make sure our array maps directly to the IROps enum\n")
output_file.write("static_assert(IRRegClasses[IROps::OP_LAST] == FEXCore::IR::InvalidClass);\n\n")
output_file.write("static_assert(IRRegClasses[IROps::OP_LAST] == RegClass::Invalid);\n\n")
output_file.write("FEXCore::IR::RegisterClassType GetRegClass(IROps Op) { return IRRegClasses[Op]; }\n\n")
output_file.write("FEXCore::IR::RegClass GetRegClass(IROps Op) { return IRRegClasses[Op]; }\n\n")
output_file.write("#undef IROP_REG_CLASSES_IMPL\n")
output_file.write("#endif\n\n")
@@ -561,9 +566,7 @@ def print_ir_arg_printer():
SSAArgNum = 0
FirstArg = True
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
for arg in op.Arguments:
# No point printing temporaries that we can't recover
if arg.Temporary:
continue
@@ -588,13 +591,13 @@ def print_ir_arg_printer():
output_file.write("#endif\n")
def print_validation(op):
if op.EmitValidation != None:
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
if len(op.EmitValidation) != 0:
output_file.write("#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
Sanitized = Validation.replace("\"", "\\\"")
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
output_file.write("\t\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("#endif\n")
# Print out IR allocator helpers
def print_ir_allocator_helpers():
@@ -664,7 +667,7 @@ def print_ir_allocator_helpers():
output_file.write("\t\treturn HeaderOp->Op;\n")
output_file.write("\t}\n\n")
output_file.write("\tFEXCore::IR::RegisterClassType GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\tFEXCore::IR::RegClass GetOpRegClass(const OrderedNode *Op) const {\n")
output_file.write("\t\treturn GetRegClass(GetOpType(Op));\n")
output_file.write("\t}\n\n")
@@ -678,22 +681,21 @@ def print_ir_allocator_helpers():
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
# Output SSA args first
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
elif arg.IsSSA:
# SSA value
output_file.write("OrderedNodeWrapper {}".format(arg.Name))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
if arg.DefaultInitializer != None:
if arg.DefaultInitializer:
output_file.write(" = {}".format(arg.DefaultInitializer))
if not LastArg:
@@ -751,20 +753,19 @@ def print_ir_allocator_helpers():
if op.SSAArgNum:
output_file.write("\tIRPair<IROp_{}> _{}(" .format(op.Name, op.Name))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
elif arg.IsSSA:
output_file.write("OrderedNode *{}".format(arg.Name))
else:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} {}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name))
if arg.DefaultInitializer != None:
if arg.DefaultInitializer:
output_file.write(" = {}".format(arg.DefaultInitializer))
if not LastArg:
@@ -773,9 +774,29 @@ def print_ir_allocator_helpers():
output_file.write(") {\n")
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
idx = 0
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
# Inline an immediate if we can
inline = op.Inline[idx]
idx += 1
if inline != '':
Sized = "Size" in [x.Name for x in op.Arguments]
P = ["Size" if Sized else "OpSize::i64Bit", arg.Name]
# A few cases need extra info plumbed.
if inline == "SubtractZero":
P += ["Src2"]
elif inline == "Mem":
P += ["OffsetType", "OffsetScale"]
elif inline == "Memtso":
P += ["OffsetType", "OffsetScale", "true /* TSO */"]
inline = "Mem"
output_file.write(f"\t\t{arg.Name} = Inline{inline}({', '.join(P)});\n")
output_file.write(f"\t\t{arg.Name}->AddUse();\n")
# Insert validation here. This is skipped for the
# OrderedNodeWrapper version because validation can depend on
@@ -785,16 +806,15 @@ def print_ir_allocator_helpers():
print_validation(op)
output_file.write(f"\t\treturn _{op.Name}(")
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
for i, arg in enumerate(op.Arguments):
LastArg = i == len(op.Arguments) - 1
output_file.write(arg.Name)
if arg.IsSSA:
output_file.write("->Wrapped(ListDataBegin)")
if not LastArg:
output_file.write(", ")
output_file.write(");\n");
output_file.write("\t}\n\n");
output_file.write(");\n")
output_file.write("\t}\n\n")
output_file.write("#undef IROP_ALLOCATE_HELPERS\n")
output_file.write("#endif\n")
@@ -825,8 +845,8 @@ def print_ir_dispatcher_dispatch():
output_dispatch_file.write("#endif\n")
if (len(sys.argv) < 4):
ExitError()
if len(sys.argv) < 4:
ExitError("Insufficient parameters passed to script")
output_filename = sys.argv[2]
output_dispatcher_filename = sys.argv[3]
@@ -838,6 +858,7 @@ json_file.close()
json_object = json.loads(json_text)
json_object = {k.upper(): v for k, v in json_object.items()}
enums = json_object["ENUMS"]
ops = json_object["OPS"]
irtypes = json_object["IRTYPES"]
defines = json_object["DEFINES"]
@@ -847,7 +868,7 @@ parse_ops(ops)
output_file = open(output_filename, "w")
print_enums()
print_enums(enums)
print_ir_structs(defines)
print_ir_sizes()
print_ir_reg_classes()
+2 -7
View File
@@ -18,14 +18,12 @@ set (SRCS
Common/JitSymbols.cpp
Interface/Context/Context.cpp
Interface/Core/LookupCache.cpp
Interface/Core/CodeCache.cpp
Interface/Core/Core.cpp
Interface/Core/CPUBackend.cpp
Interface/Core/Addressing.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
Interface/Core/OpcodeDispatcher/AVX_128.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
@@ -33,7 +31,6 @@ set (SRCS
Interface/Core/OpcodeDispatcher/X87.cpp
Interface/Core/OpcodeDispatcher/X87F64.cpp
Interface/Core/OpcodeDispatcher.cpp
Interface/Core/X86Tables.cpp
Interface/Core/X86HelperGen.cpp
Interface/Core/ArchHelpers/Arm64Emitter.cpp
Interface/Core/Dispatcher/Dispatcher.cpp
@@ -61,11 +58,9 @@ set (SRCS
Interface/Core/X86Tables/VEXTables.cpp
Interface/Core/X86Tables/X87Tables.cpp
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
@@ -207,7 +202,7 @@ add_custom_target(CONFIG_INC
DEPENDS "${OUTPUT_MAN_NAME_COMPRESS}")
# Install the compressed man page
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} DESTINATION ${MAN_DIR}/man1)
install(FILES ${OUTPUT_MAN_NAME_COMPRESS} COMPONENT Runtime DESTINATION ${MAN_DIR}/man1)
# Add in diagnostic colours if the option is available.
# Ninja code generator will kill colours if this isn't here
+1 -3
View File
@@ -2,12 +2,10 @@
#pragma once
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <chrono>
#include <cstddef>
#include <cstdint>
#include <cstdio>
#include <memory>
#include <string_view>
namespace FEXCore {
+87 -9
View File
@@ -4,9 +4,9 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/BitUtils.h>
#include "cephes_128bit.h"
#include <bit>
#include <cmath>
#include <cstring>
#include <stdint.h>
@@ -294,7 +294,11 @@ struct FEX_PACKED X80SoftFloat {
FCMP(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(state, lhs, rhs);
*lt = extF80_lt(state, lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
// Use IEEE 754 semantics: unordered if neither <, =, nor > is true
// This is more reliable than custom NaN detection
bool gt = !(*eq) && !(*lt) && extF80_le(state, rhs, lhs);
*nan = !(*eq) && !(*lt) && !gt;
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
@@ -497,12 +501,53 @@ struct FEX_PACKED X80SoftFloat {
float ToF32(softfloat_state* state) const {
const float32_t Result = extF80_to_f32(state, *this);
return FEXCore::BitCast<float>(Result);
return std::bit_cast<float>(Result);
}
bool IsSignalingNaN() const {
return (Exponent == 0x7FFF) && (Significand & 0x8000000000000000ULL) && !(Significand & 0x4000000000000000ULL) && // Bit 62 clear (signaling)
(Significand & 0x3FFFFFFFFFFFFFFFULL);
}
bool IsQuietNaN() const {
return (Exponent == 0x7FFF) && (Significand & 0x8000000000000000ULL) && (Significand & 0x4000000000000000ULL); // Bit 62 set (quiet)
}
// Helper to detect if this is any NaN
bool IsNaN() const {
return IsSignalingNaN() || IsQuietNaN();
}
// X87 value to F64 while preserving signaling nan property
double ToF64_PreserveNan(softfloat_state* state) const {
if (IsSignalingNaN()) {
// we keep it as a signaling nan in ieee754 in 64bits
uint64_t sign_bit = Sign ? 0x8000000000000000ULL : 0;
uint64_t exp_bits = 0x7FF0000000000000ULL;
uint64_t x87_frac = Significand & 0x3FFFFFFFFFFFFFFFULL;
uint64_t ieee_frac = (x87_frac >> 11) & 0x0007FFFFFFFFFFFFULL;
if (ieee_frac == 0) {
ieee_frac = 1;
}
ieee_frac &= ~0x0008000000000000ULL;
uint64_t result_bits = sign_bit | exp_bits | ieee_frac;
return std::bit_cast<double>(result_bits);
} else if (IsQuietNaN()) {
const float64_t Result = extF80_to_f64(state, *this);
uint64_t result_bits = std::bit_cast<uint64_t>(Result);
result_bits |= 0x0008000000000000ULL;
return std::bit_cast<double>(result_bits);
} else {
const float64_t Result = extF80_to_f64(state, *this);
return std::bit_cast<double>(Result);
}
}
double ToF64(softfloat_state* state) const {
const float64_t Result = extF80_to_f64(state, *this);
return FEXCore::BitCast<double>(Result);
return std::bit_cast<double>(Result);
}
FEXCore::VectorRegType ToVector() const {
@@ -514,7 +559,7 @@ struct FEX_PACKED X80SoftFloat {
BIGFLOAT ToFMax(softfloat_state* state) const {
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(state, *this);
return FEXCore::BitCast<BIGFLOAT>(Result);
return std::bit_cast<BIGFLOAT>(Result);
#else
BIGFLOAT result {};
memcpy(&result, this, sizeof(result));
@@ -573,18 +618,51 @@ struct FEX_PACKED X80SoftFloat {
}
X80SoftFloat(softfloat_state* state, const float rhs) {
*this = f32_to_extF80(state, FEXCore::BitCast<float32_t>(rhs));
*this = f32_to_extF80(state, std::bit_cast<float32_t>(rhs));
}
X80SoftFloat(softfloat_state* state, const double rhs) {
*this = f64_to_extF80(state, FEXCore::BitCast<float64_t>(rhs));
*this = f64_to_extF80(state, std::bit_cast<float64_t>(rhs));
}
// Create X80SoftFloat from double while preserving NaN signaling properties
static X80SoftFloat FromF64_PreserveNaN(softfloat_state* state, double value) {
uint64_t bits = std::bit_cast<uint64_t>(value);
// Check if it's a nan
if ((bits & 0x7FF0000000000000ULL) == 0x7FF0000000000000ULL && (bits & 0x000FFFFFFFFFFFFFULL) != 0) {
X80SoftFloat result;
result.Sign = (bits >> 63) & 1;
result.Exponent = 0x7FFF;
bool is_signaling = !(bits & 0x0008000000000000ULL);
uint64_t ieee_payload = bits & 0x0007FFFFFFFFFFFFULL;
// set bit 63 required for x87
result.Significand = 0x8000000000000000ULL;
if (is_signaling) { // clear bit 62 for signaling nan
result.Significand &= ~0x4000000000000000ULL;
} else { // clear bit 62 for quiet nan
result.Significand |= 0x4000000000000000ULL;
}
// ieee754 51-bit payload -> x87 62-bit payload
result.Significand |= (ieee_payload << 11) & 0x3FFFFFFFFFFFFFFFULL;
return result;
}
// For non-NaN values, use standard conversion
return X80SoftFloat(state, value);
}
X80SoftFloat(softfloat_state* state, BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(state, FEXCore::BitCast<float128_t>(rhs));
*this = f128_to_extF80(state, std::bit_cast<float128_t>(rhs));
#else
*this = FEXCore::BitCast<long double>(rhs);
*this = std::bit_cast<long double>(rhs);
#endif
}
+12 -59
View File
@@ -2,74 +2,27 @@
#pragma once
#include <FEXCore/fextl/string.h>
#include <cstdint>
#include <concepts>
#include <string_view>
#include <optional>
namespace FEXCore::StrConv {
[[maybe_unused]]
static bool Conv(std::string_view Value, bool* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
template<std::integral T>
bool Conv(std::string_view Value, T* Result) {
if constexpr (std::is_signed_v<T>) {
*Result = static_cast<T>(std::strtoll(Value.data(), nullptr, 0));
} else {
*Result = static_cast<T>(std::strtoull(Value.data(), nullptr, 0));
}
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint8_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
template<typename T, typename = std::enable_if_t<std::is_enum_v<T>, T>>
bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::strtoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int8_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint16_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int16_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint32_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int32_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint64_t* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int64_t* Result) {
*Result = std::strtoll(Value.data(), nullptr, 0);
return true;
}
template<typename T, typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]]
static bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, fextl::string* Result) {
inline bool Conv(std::string_view Value, fextl::string* Result) {
*Result = Value;
return true;
}
+44 -62
View File
@@ -4,7 +4,6 @@
"Multiblock": {
"Type": "bool",
"Default": "true",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation",
"Can cause long JIT compilation times and stutter"
@@ -13,22 +12,10 @@
"MaxInst": {
"Type": "int32",
"Default": "5000",
"ShortArg": "n",
"Desc": [
"Maximum number of instruction to store in a block"
]
},
"CacheObjectCodeCompilation": {
"Type": "uint32",
"Default": "FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE",
"TextDefault": "none",
"Choices": [ "none", "read", "readwrite" ],
"ArgumentHandler": "CacheObjectCodeHandler",
"Desc": [
"Cache JIT object code to drive.",
"Allows JIT code to be shared between applications"
]
},
"HostFeatures": {
"Type": "strenum",
"Default": "FEXCore::Config::HostFeatures::OFF",
@@ -70,7 +57,11 @@
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi",
"ENABLEWFXT": "enablewfxt",
"DISABLEWFXT": "disablewfxt"
"DISABLEWFXT": "disablewfxt",
"ENABLE3DNOW": "enable3dnow",
"DISABLE3DNOW": "disable3dnow",
"ENABLESSE4A": "enablesse4a",
"DISABLESSE4A": "disablesse4a"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -92,7 +83,9 @@
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it",
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it"
"\t{enable,disable}wfxt: Will force enable or disable wfxt even if the host doesn't support it",
"\t{enable,disable}3dnow: Will force enable or disable 3DNow! even if the host doesn't support it",
"\t{enable,disable}sse4a: Will force enable or disable SSE4a even if the host doesn't support it"
]
},
"SmallTSCScale": {
@@ -107,7 +100,6 @@
"RootFS": {
"Type": "str",
"Default": "",
"ShortArg": "R",
"Desc": [
"Which Root filesystem prefix to use",
"This can be a filesystem path",
@@ -122,7 +114,6 @@
"ThunkHostLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_FULL_LIBDIR@/fex-emu/HostThunks",
"ShortArg": "t",
"Desc": [
"Folder to find the host-side thunking libraries."
]
@@ -130,7 +121,6 @@
"ThunkGuestLibs": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks",
"ShortArg": "j",
"Desc": [
"Folder to find the guest-side thunking libraries."
]
@@ -138,7 +128,6 @@
"ThunkConfig": {
"Type": "str",
"Default": "",
"ShortArg": "k",
"Desc": [
"A json file specifying where to overlay the thunks.",
"This can be a filesystem path",
@@ -153,7 +142,6 @@
"Env": {
"Type": "strarray",
"Default": "",
"ShortArg": "E",
"Desc": [
"Adds an environment variable to the emulated environment."
]
@@ -161,7 +149,6 @@
"HostEnv": {
"Type": "strarray",
"Default": "",
"ShortArg": "H",
"Desc": [
"Adds an environment variable to the host environment.",
"This can be useful for setting environment variables that thunks can pick up.",
@@ -180,7 +167,6 @@
"SingleStep": {
"Type": "bool",
"Default": "false",
"ShortArg": "S",
"Desc": [
"Single stepping configuration."
]
@@ -188,7 +174,6 @@
"GdbServer": {
"Type": "bool",
"Default": "false",
"ShortArg": "G",
"Desc": [
"Enables the GDB server."
]
@@ -222,7 +207,6 @@
"DumpGPRs": {
"Type": "bool",
"Default": "false",
"ShortArg": "g",
"Desc": [
"When the test harness ends, print the GPR state."
]
@@ -230,7 +214,6 @@
"O0": {
"Type": "bool",
"Default": "false",
"ShortArg": "O0",
"Desc": [
"Disables optimizations passes for debugging."
]
@@ -320,7 +303,6 @@
"SilentLog": {
"Type": "bool",
"Default": "true",
"ShortArg": "s",
"Desc": [
"Disables logging"
]
@@ -328,7 +310,6 @@
"OutputLog": {
"Type": "str",
"Default": "server",
"ShortArg": "o",
"Desc": [
"File to write FEX output to.",
"[stdout, stderr, server, <Filename>]"
@@ -403,14 +384,6 @@
"This is required to ensure a split-lock doesn't tear inside the process"
]
},
"TSOAutoMigration": {
"Type": "bool",
"Default": "true",
"Desc": [
"Automatically enables TSO when shared memory is used.",
"Should work without issues in most cases."
]
},
"VolatileMetadata": {
"Type": "bool",
"Default": "true",
@@ -426,6 +399,14 @@
"Emulates X87 floating point using 64-bit precision. This reduces emulation accuracy and may result in rendering bugs."
]
},
"X87StrictReducedPrecision": {
"Type": "bool",
"Default": "false",
"Desc": [
"Enables stricter X87 floating point behavior when X87ReducedPrecision is enabled.",
"Adds additional checks and implementations like NaN propagation for better compatibility."
]
},
"ABILocalFlags": {
"Type": "bool",
"Default": "false",
@@ -473,32 +454,16 @@
"Desc": [
"Contrains the startup sleep to only apply to processes that match this name."
]
},
"MonoHacks": {
"Type": "bool",
"Default": "true",
"Desc": [
"Permits a hook-based SMC approach and smaller JIT blocks when mono is detected."
]
}
},
"Misc": {
"AOTIRCapture": {
"Type": "bool",
"Default": "false",
"Desc": [
"Captures IR and generates an AOT IR cache.",
"Captures both the loaded executable and libraries it loads."
]
},
"AOTIRGenerate": {
"Type": "bool",
"Default": "false",
"Desc": [
"Scans file for executable code and generates an AOT IR cache.",
"Does not run the executable."
]
},
"AOTIRLoad": {
"Type": "bool",
"Default": "false",
"Desc": [
"Loads an AOT IR cache for the loaded executable."
]
},
"ServerSocketPath": {
"Type": "str",
"Default": "",
@@ -512,15 +477,32 @@
"Desc": [
"Disables inline syscalls in order to support seccomp handling"
]
},
"ExtendedVolatileMetadata": {
"Type": "str",
"Default": "",
"Desc": [
"Configuration provided volatile metadata. Only implemented for WoW64/arm64ec.",
"Limited in its use but can be handy.",
"Extends on top of what Microsoft has for volatile metadata, but also supported for WoW64.",
"Colon delimited modules, then semi-colon delimited instructions, then comma delimited ranges",
"Default disables TSO in the module, unless instructions overlap the range",
"<module>;<offset begin>-<offset-end>,...;<instruction offset to force TSO>,...:<another>",
"examples:",
" * Disable TSO for a full module: Just provide the module name:",
" `hl2_linux`",
" * Disable TSO for a part of the module:",
" `hl2_linux;<offset begin>-<offset-end>`",
" * Disable TSO for a part of the module, but enable TSO for some instructions within the module",
" `hl2_linux;<offset begin>-<offset-end>;<instruction offset>,<instruction offset>`",
" * Disable TSO for multiple modules",
" `hl2_linux:libsdl2.so`"
]
}
}
},
"UnnamedOptions": {
"Misc": {
"IS_INTERPRETER": {
"Type": "bool",
"Default": "false"
},
"INTERPRETER_INSTALLED": {
"Type": "bool",
"Default": "false"
+3 -8
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: MIT
#include "Interface/Context/Context.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/CoreState.h>
@@ -8,18 +9,12 @@
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Core/Thunks.h>
#include "FEXCore/Debug/InternalThreadState.h"
#include <string.h>
#include <utility>
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
X86Tables::InitializeInfoTables(Mode);
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext(const FEXCore::HostFeatures& Features) {
return fextl::make_unique<FEXCore::Context::ContextImpl>(Features);
}
+66 -82
View File
@@ -2,61 +2,49 @@
#pragma once
#include "Common/JitSymbols.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/AOTIR.h"
#include <Interface/IR/IntrusiveIRList.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <stdint.h>
#include <atomic>
#include <cstddef>
#include <cstdint>
#include <mutex>
#include <optional>
#include <shared_mutex>
namespace FEXCore {
class CodeLoader;
class SignalDelegator;
class ThunkHandler;
namespace CodeSerialize {
class CodeObjectSerializeService;
}
namespace Core {
struct DebugData;
struct InternalThreadState;
} // namespace Core
namespace CPU {
class Arm64JITCore;
class Dispatcher;
} // namespace CPU
namespace HLE {
struct SyscallArguments;
class SyscallHandler;
class SourcecodeResolver;
struct SourcecodeMap;
class SyscallHandler;
} // namespace HLE
} // namespace FEXCore
namespace FEXCore::IR {
struct IRListCopy;
class IRListView;
namespace Validation {
class IRValidation;
}
} // namespace FEXCore::IR
namespace FEXCore::Context {
struct FEX_PACKED ExitFunctionLinkData {
uint64_t HostCode;
@@ -76,7 +64,23 @@ struct CustomIRResult {
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class ContextImpl final : public FEXCore::Context::Context, CPU::CodeBufferManager {
class CodeCache : public AbstractCodeCache {
public:
CodeCache(ContextImpl&);
~CodeCache();
ContextImpl& CTX;
bool IsGeneratingCache = false;
void LoadData(Core::InternalThreadState&, std::byte* MappedCacheFile, const ExecutableFileSectionInfo&) override;
bool SaveData(Core::InternalThreadState&, int TargetFD, const ExecutableFileSectionInfo&, uint64_t SerializedBaseAddress) override;
void InitiateCacheGeneration() override {
IsGeneratingCache = true;
}
};
class ContextImpl final : public FEXCore::Context::Context, public CPU::CodeBufferManager {
public:
// Context base class implementation.
bool InitCore() override;
@@ -90,6 +94,7 @@ public:
bool IsAddressInCurrentBlock(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, uint64_t Size) override;
bool IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* Thread) override;
uint64_t GetGuestBlockEntry(FEXCore::Core::InternalThreadState* Thread) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, const uint64_t* HostGPRs, uint64_t PSTATE) override;
@@ -145,24 +150,8 @@ public:
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::IR::AOTIRCacheEntry* LoadAOTIRCacheEntry(const fextl::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry* Entry) override;
void SetAOTIRLoader(AOTIRLoaderCBFn CacheReader) override {
IRCaptureCache.SetAOTIRLoader(std::move(CacheReader));
}
void SetAOTIRWriter(AOTIRWriterCBFn CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(std::move(CacheWriter));
}
void SetAOTIRRenamer(AOTIRRenamerCBFn CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(std::move(CacheRenamer));
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(AOTIRCodeFileWriterFn Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
CodeCache& GetCodeCache() override {
return CodeCache;
}
void OnCodeBufferAllocated(CPU::CodeBuffer&) override;
@@ -173,8 +162,6 @@ public:
return CodeInvalidationMutex;
}
void MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) override;
bool IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const override;
@@ -189,14 +176,13 @@ public:
void RemoveForceTSOInformation(uint64_t Address, uint64_t Size) override;
void MarkMonoDetected() override {
MonoDetected = true;
}
void MarkMonoBackpatcherBlock(uint64_t BlockEntry) override;
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
friend class FEXCore::IR::Validation::IRValidation;
struct {
uint64_t VirtualMemSize {1ULL << 36};
uint64_t TSCScale = 0;
@@ -209,13 +195,9 @@ public:
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(TSOAutoMigration, TSOAUTOMIGRATION);
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRGenerate, AOTIRGENERATE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
@@ -224,12 +206,13 @@ public:
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(x87StrictReducedPrecision, X87STRICTREDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
FEX_CONFIG_OPT(StrictInProcessSplitLocks, STRICTINPROCESSSPLITLOCKS);
FEX_CONFIG_OPT(MonoHacks, MONOHACKS);
} Config;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
@@ -243,29 +226,23 @@ public:
FEXCore::HLE::SourcecodeResolver* SourcecodeResolver {};
FEXCore::ThunkHandler* ThunkHandler {};
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CodeCache CodeCache;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl(const FEXCore::HostFeatures& Features);
~ContextImpl();
static bool ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP);
// NOTE: Other threads sharing the same CodeBuffer may reference
// invalidated data ranges through their L1/L2 caches. This is
// not currently a problem since FEX does not repurpose the
// invalidated CodeBuffer memory range currently.
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
// This is used as a replacement for the SMC writes in the mono callsite backpatcher that avoids atomic operations
// (safe as the invalidation mutex is locked) and manually invalidates the modified range. Allowing SMC to be detected
// even if faulting is disabled.
static void MonoBackpatcherWrite(FEXCore::Core::CpuStateFrame* Frame, uint8_t Size, uint64_t Address, uint64_t Value);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
void RemoveCustomIREntrypoint(FEXCore::Core::InternalThreadState* Thread, uintptr_t Entrypoint);
struct GenerateIRResult {
std::optional<IR::IRListView> IRView;
@@ -324,6 +301,10 @@ public:
return ExitOnHLT;
}
bool AreMonoHacksActive() const {
return Config.MonoHacks && MonoDetected;
}
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -336,12 +317,16 @@ protected:
VectorAtomicTSOEmulationEnabled = true;
MemcpyAtomicTSOEmulationEnabled = true;
} else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
// Atomic vector TSO emulation only enabled if TSO emulation is enabled and also vector TSO is enabled.
VectorAtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled && Config.VectorTSOEnabled;
// Atomic memcpy TSO emulation only enabled if TSO emulation is enabled and also memcpy TSO is enabled.
MemcpyAtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled && Config.MemcpySetTSOEnabled;
AtomicTSOEmulationEnabled = Config.TSOEnabled;
VectorAtomicTSOEmulationEnabled = Config.TSOEnabled && Config.VectorTSOEnabled;
MemcpyAtomicTSOEmulationEnabled = Config.TSOEnabled && Config.MemcpySetTSOEnabled;
}
}
void UpdateX87PrecisionConfig() {
// If strict reduced precision is enabled, automatically enable reduced precision
if (Config.x87StrictReducedPrecision() && !Config.x87ReducedPrecision()) {
FEXCore::Config::Set(FEXCore::Config::CONFIG_X87REDUCEDPRECISION, "1");
}
}
@@ -355,10 +340,6 @@ private:
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
bool VectorAtomicTSOEmulationEnabled = false;
@@ -371,11 +352,14 @@ private:
std::atomic<bool> HasCustomIRHandlers {};
struct CustomIRHandlerEntry final {
CustomIREntrypointHandler Handler;
void *Creator;
void *Data;
void* Creator;
void* Data;
};
fextl::unordered_map<uint64_t, CustomIRHandlerEntry> CustomIRHandlers;
IntervalList<uint64_t> ForceTSOValidRanges; // The ranges for which ForceTSOInstructions has populated data
fextl::set<uint64_t> ForceTSOInstructions;
bool MonoDetected = false;
std::atomic<uint64_t> MonoBackpatcherBlock;
};
} // namespace FEXCore::Context
+30 -42
View File
@@ -7,12 +7,11 @@
namespace FEXCore::IR {
Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage) {
Ref LoadEffectiveAddress(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage) {
Ref Tmp = A.Base;
if (A.Offset) {
Ref Offset = IREmit->Constant(A.Offset);
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, Offset) : Offset;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Offset) : IREmit->Constant(A.Offset);
}
if (A.Index) {
@@ -25,7 +24,7 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
Tmp = IREmit->_Lshl(GPRSize, A.Index, IREmit->Constant(Log2));
}
} else {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Index) : A.Index;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Index) : A.Index;
}
}
@@ -46,27 +45,23 @@ Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, b
}
if (A.Segment && AddSegmentBase) {
Tmp = Tmp ? IREmit->_Add(GPRSize, Tmp, A.Segment) : A.Segment;
Tmp = Tmp ? IREmit->Add(GPRSize, Tmp, A.Segment) : A.Segment;
}
return Tmp ?: IREmit->Constant(0);
}
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
IR::OpSize AccessSize) {
auto SoftwareAddressCalculation = [IREmit, &A, GPRSize]() -> AddressMode {
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
};
AddressMode SelectAddressMode(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO,
bool Vector, IR::OpSize AccessSize) {
const auto Is32Bit = GPRSize == OpSize::i32Bit;
const auto GPRSizeMatchesAddrSize = A.AddrSize == GPRSize;
const auto OffsetIndexToLargeFor32Bit = Is32Bit && (A.Offset <= -16384 || A.Offset >= 16384);
if (!GPRSizeMatchesAddrSize || OffsetIndexToLargeFor32Bit) {
// If address size doesn't match GPR size then no optimizations can occur.
return SoftwareAddressCalculation();
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
}
// Loadstore rules:
@@ -100,7 +95,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
const bool OffsetIsSIMM9 = A.Offset && A.Offset >= -256 && A.Offset <= 255;
const bool OffsetIsUnsignedScaled = A.Offset > 0 && (A.Offset & (AccessSizeAsImm - 1)) == 0 && (A.Offset / AccessSizeAsImm) <= 4095;
auto InlineImmOffsetLoadstore = [IREmit, &GPRSize](AddressMode A) -> AddressMode {
if ((AtomicTSO && !Vector && HostSupportsTSOImm9 && OffsetIsSIMM9) || (!AtomicTSO && (OffsetIsSIMM9 || OffsetIsUnsignedScaled))) {
// Peel off the offset
AddressMode B = A;
B.Offset = 0;
@@ -108,35 +103,25 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexType = MemOffsetType::SXTX,
.IndexScale = 1,
};
};
auto ScaledRegisterLoadstore = [IREmit, GPRSize](AddressMode A) -> AddressMode {
if (A.Index && A.Segment) {
A.Base = IREmit->_Add(GPRSize, A.Base, A.Segment);
} else if (A.Segment) {
A.Index = A.Segment;
A.IndexScale = 1;
}
return A;
};
}
if (AtomicTSO) {
if (!Vector) {
if (HostSupportsTSOImm9 && OffsetIsSIMM9) {
return InlineImmOffsetLoadstore(A);
}
} else {
// TODO: LRCPC3 support for vector Imm9.
}
} else {
if (OffsetIsSIMM9 || OffsetIsUnsignedScaled) {
return InlineImmOffsetLoadstore(A);
} else if (!Is32Bit && A.Base && (A.Index || A.Segment) && !A.Offset && (A.IndexScale == 1 || A.IndexScale == AccessSizeAsImm)) {
return ScaledRegisterLoadstore(A);
// TODO: LRCPC3 support for vector Imm9.
} else if (!Is32Bit && A.Base && (A.Index || A.Segment) && !A.Offset && (A.IndexScale == 1 || A.IndexScale == AccessSizeAsImm)) {
AddressMode B = A;
// ScaledRegisterLoadstore
if (B.Index && B.Segment) {
B.Base = IREmit->Add(GPRSize, B.Base, B.Segment);
} else if (B.Segment) {
B.Index = B.Segment;
B.IndexScale = 1;
}
return B;
}
if (Vector || !AtomicTSO) {
@@ -151,7 +136,7 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
return {
.Base = LoadEffectiveAddress(IREmit, B, GPRSize, true /* AddSegmentBase */, false),
.Index = IREmit->Constant(A.Offset),
.IndexType = MEM_OFFSET_SXTX,
.IndexType = MemOffsetType::SXTX,
.IndexScale = 1,
};
}
@@ -159,7 +144,10 @@ AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSi
}
// Fallback on software address calculation
return SoftwareAddressCalculation();
return {
.Base = LoadEffectiveAddress(IREmit, A, GPRSize, true),
.Index = IREmit->Invalid(),
};
}
+7 -6
View File
@@ -11,17 +11,18 @@ struct AddressMode {
Ref Segment {nullptr};
Ref Base {nullptr};
Ref Index {nullptr};
MemOffsetType IndexType = MEM_OFFSET_SXTX;
uint8_t IndexScale = 1;
int64_t Offset = 0;
MemOffsetType IndexType = MemOffsetType::SXTX;
uint8_t IndexScale = 1;
// Size in bytes for the address calculation. 8 for an arm64 hardware mode.
IR::OpSize AddrSize;
bool NonTSO;
};
Ref LoadEffectiveAddress(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage = false);
AddressMode SelectAddressMode(IREmitter* IREmit, AddressMode A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO, bool Vector,
IR::OpSize AccessSize);
Ref LoadEffectiveAddress(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool AddSegmentBase, bool AllowUpperGarbage = false);
AddressMode SelectAddressMode(IREmitter* IREmit, const AddressMode& A, IR::OpSize GPRSize, bool HostSupportsTSOImm9, bool AtomicTSO,
bool Vector, IR::OpSize AccessSize);
}; // namespace FEXCore::IR
} // namespace FEXCore::IR
@@ -1,10 +1,10 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Core/X86Enums.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -712,7 +712,7 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
// Now handle PF/AF
if (PFAFSpillMask) {
auto PFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.pf_raw);
[[maybe_unused]] auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
auto AFOffset = offsetof(FEXCore::Core::CpuStateFrame, State.af_raw);
LOGMAN_THROW_A_FMT(PFAFSpillMask == PFAFMask, "PF/AF not spilled together");
LOGMAN_THROW_A_FMT(AFOffset == PFOffset + 4, "PF/AF are together");
@@ -1,30 +1,31 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#ifdef VIXL_DISASSEMBLER
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/vector.h>
#endif
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#include <aarch64/simulator-constants-aarch64.h>
#endif
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#include <cstddef>
#include <cstdint>
#include <optional>
#include <span>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::X86State {
enum X86Reg : uint32_t;
}
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
+13 -7
View File
@@ -1,14 +1,17 @@
// SPDX-License-Identifier: MIT
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/Utils/PrctlUtils.h>
#include <cstdint>
#include "LookupCache.h"
#ifndef _WIN32
#include <linux/prctl.h>
#include <sys/prctl.h>
#endif
@@ -317,7 +320,7 @@ namespace CPU {
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer = CodeBuffers.StartLargerCodeBuffer();
RegisterForSignalHandler(PrevCodeBuffer);
RegisterForSignalHandler(std::move(PrevCodeBuffer));
return CurrentCodeBuffer.get();
}
@@ -326,14 +329,13 @@ namespace CPU {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Keep a reference to the old code buffer to delay deallocation
SignalHandlerCodeBuffers.push_back(CodeBuffer);
SignalHandlerCodeBuffers.push_back(std::move(CodeBuffer));
} else {
SignalHandlerCodeBuffers.clear();
}
}
fextl::shared_ptr<CodeBuffer> CPUBackend::CheckCodeBufferUpdate() {
fextl::shared_ptr<CodeBuffer> OldCodeBuffer;
auto NewCodeBuffer = CodeBuffers.GetLatest();
if (CurrentCodeBuffer != NewCodeBuffer) {
RegisterForSignalHandler(CurrentCodeBuffer);
@@ -358,6 +360,10 @@ namespace CPU {
LogMan::Msg::EFmt("Failed to mprotect last page of code buffer.");
}
#ifndef _WIN32
prctl(PR_SET_VMA, PR_SET_VMA_ANON_NAME, Ptr, Size, "FEXMemJIT");
#endif
LookupCache = fextl::make_unique<GuestToHostMap>();
}
+5 -12
View File
@@ -17,6 +17,10 @@ $end_info$
#include <cstdint>
namespace FEXCore::CPU {
union Relocation;
}
namespace FEXCore {
namespace IR {
@@ -157,18 +161,7 @@ namespace CPU {
virtual CompiledCode CompileCode(uint64_t Entry, uint64_t Size, bool SingleInst, const FEXCore::IR::IRListView* IR,
FEXCore::Core::DebugData* DebugData, bool CheckTF) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
*
* @param Entry - RIP of the entry
* @param SerializationData - Serialization data referring to the object cache for `Entry`
*
* @return An executable function pointer relocated from the cache object
*/
[[nodiscard]]
virtual void* RelocateJITObjectCode(uint64_t /* Entry */, const CodeSerialize::CodeObjectFileSection* /* SerializationData */) {
return nullptr;
}
virtual fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() = 0;
virtual void ClearCache() {}
+113 -96
View File
@@ -43,12 +43,15 @@ namespace ProductNames {
static const char ARM_A715[] = "Cortex-A715";
static const char ARM_A720[] = "Cortex-A720";
static const char ARM_A725[] = "Cortex-A725";
static const char ARM_C1Pro[] = "C1-Pro";
static const char ARM_C1Premium[] = "C1-Premium";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X1C[] = "Cortex-X1C";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_X3[] = "Cortex-X3";
static const char ARM_X4[] = "Cortex-X4";
static const char ARM_X925[] = "Cortex-X925";
static const char ARM_C1Ultra[] = "C1-Ultra";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
static const char ARM_N3[] = "Neoverse N3";
@@ -59,6 +62,7 @@ namespace ProductNames {
static const char ARM_A65[] = "Cortex-A65";
static const char ARM_A510[] = "Cortex-A510";
static const char ARM_A520[] = "Cortex-A520";
static const char ARM_C1Nano[] = "C1-Nano";
static const char ARM_Kryo200[] = "Kryo 2xx";
static const char ARM_Kryo300[] = "Kryo 3xx";
@@ -70,6 +74,7 @@ namespace ProductNames {
static const char ARM_Denver[] = "Nvidia Denver";
static const char ARM_Carmel[] = "Nvidia Carmel";
static const char ARM_Olympus[] = "Nvidia Olympus";
static const char ARM_Firestorm_M1[] = "Apple Firestorm (M1)";
static const char ARM_Icestorm_M1[] = "Apple Icestorm (M1)";
@@ -85,6 +90,9 @@ namespace ProductNames {
static const char ARM_Blizzard_M2Max[] = "Apple Blizzard (M2 Max)";
static const char ARM_ORYON_1[] = "Oryon-1";
static const char ARM_Ampere_1[] = "AmpereOne";
static const char ARM_Ampere_1A[] = "AmpereOneA";
static const char ARM_Ampere_1B[] = "AmpereOneB";
#else
#endif
} // namespace ProductNames
@@ -170,7 +178,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 58> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 66> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
@@ -181,38 +189,46 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x025, 1, ProductNames::ARM_Firestorm_M1Pro}, // Apple Firestorm (M1 Pro)
{0x61, 0x023, 1, ProductNames::ARM_Firestorm_M1}, // Apple Firestorm (M1)
{0x41, 0xd85, 1, ProductNames::ARM_X925}, // X925
{0x41, 0xd87, 1, ProductNames::ARM_A725}, // A725
{0x41, 0xd84, 1, ProductNames::ARM_V3}, // V3
{0x41, 0xd83, 1, ProductNames::ARM_V3AE}, // V3AE
{0x41, 0xd8e, 1, ProductNames::ARM_N3}, // N3
{0x41, 0xd82, 1, ProductNames::ARM_X4}, // X4
{0x41, 0xd81, 1, ProductNames::ARM_A720}, // A720
{0x41, 0xd4e, 1, ProductNames::ARM_X3}, // X3
{0x41, 0xd4d, 1, ProductNames::ARM_A715}, // A715
{0x41, 0xd4f, 1, ProductNames::ARM_V2}, // V2
{0x41, 0xd4b, 1, ProductNames::ARM_A78C}, // A78C
{0x41, 0xd4a, 1, ProductNames::ARM_E1}, // E1
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd48, 1, ProductNames::ARM_X2}, // X2
{0x41, 0xd47, 1, ProductNames::ARM_A710}, // A710
{0x41, 0xd4C, 1, ProductNames::ARM_X1C}, // X1C
{0x41, 0xd44, 1, ProductNames::ARM_X1}, // X1
{0x41, 0xd42, 1, ProductNames::ARM_A78AE}, // A78AE
{0x41, 0xd41, 1, ProductNames::ARM_A78}, // A78
{0x41, 0xd40, 1, ProductNames::ARM_V1}, // V1
{0x41, 0xd0e, 1, ProductNames::ARM_A76AE}, // A76AE
{0x41, 0xd0d, 1, ProductNames::ARM_A77}, // A77
{0x41, 0xd0c, 1, ProductNames::ARM_N1}, // N1
{0x41, 0xd0b, 1, ProductNames::ARM_A76}, // A76
{0x51, 0x804, 1, ProductNames::ARM_Kryo400}, // Kryo 4xx Gold (A76 based)
{0x41, 0xd0a, 1, ProductNames::ARM_A75}, // A75
{0x51, 0x802, 1, ProductNames::ARM_Kryo300}, // Kryo 3xx Gold (A75 based)
{0x41, 0xd09, 1, ProductNames::ARM_A73}, // A73
{0x51, 0x800, 1, ProductNames::ARM_Kryo200}, // Kryo 2xx Gold (A73 based)
{0x41, 0xd08, 1, ProductNames::ARM_A72}, // A72
{0x41, 0xd8c, 1, ProductNames::ARM_C1Ultra}, // C1-Ultra
{0x41, 0xd90, 1, ProductNames::ARM_C1Premium}, // C1-Premium
{0x41, 0xd8b, 1, ProductNames::ARM_C1Pro}, // C1-Pro
{0x41, 0xd85, 1, ProductNames::ARM_X925}, // X925
{0x41, 0xd87, 1, ProductNames::ARM_A725}, // A725
{0x41, 0xd84, 1, ProductNames::ARM_V3}, // V3
{0x41, 0xd83, 1, ProductNames::ARM_V3AE}, // V3AE
{0x41, 0xd8e, 1, ProductNames::ARM_N3}, // N3
{0x41, 0xd82, 1, ProductNames::ARM_X4}, // X4
{0x41, 0xd81, 1, ProductNames::ARM_A720}, // A720
{0x41, 0xd4e, 1, ProductNames::ARM_X3}, // X3
{0x41, 0xd4d, 1, ProductNames::ARM_A715}, // A715
{0x41, 0xd4f, 1, ProductNames::ARM_V2}, // V2
{0x41, 0xd4b, 1, ProductNames::ARM_A78C}, // A78C
{0x41, 0xd4a, 1, ProductNames::ARM_E1}, // E1
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd48, 1, ProductNames::ARM_X2}, // X2
{0x41, 0xd47, 1, ProductNames::ARM_A710}, // A710
{0x41, 0xd4C, 1, ProductNames::ARM_X1C}, // X1C
{0x41, 0xd44, 1, ProductNames::ARM_X1}, // X1
{0x41, 0xd42, 1, ProductNames::ARM_A78AE}, // A78AE
{0x41, 0xd41, 1, ProductNames::ARM_A78}, // A78
{0x41, 0xd40, 1, ProductNames::ARM_V1}, // V1
{0x41, 0xd0e, 1, ProductNames::ARM_A76AE}, // A76AE
{0x41, 0xd0d, 1, ProductNames::ARM_A77}, // A77
{0x41, 0xd0c, 1, ProductNames::ARM_N1}, // N1
{0x41, 0xd0b, 1, ProductNames::ARM_A76}, // A76
{0x51, 0x804, 1, ProductNames::ARM_Kryo400}, // Kryo 4xx Gold (A76 based)
{0x41, 0xd0a, 1, ProductNames::ARM_A75}, // A75
{0x51, 0x802, 1, ProductNames::ARM_Kryo300}, // Kryo 3xx Gold (A75 based)
{0x41, 0xd09, 1, ProductNames::ARM_A73}, // A73
{0x51, 0x800, 1, ProductNames::ARM_Kryo200}, // Kryo 2xx Gold (A73 based)
{0x41, 0xd08, 1, ProductNames::ARM_A72}, // A72
{0x4e, 0x004, 1, ProductNames::ARM_Carmel}, // Carmel
{0xc0, 0xac3, 1, ProductNames::ARM_Ampere_1}, // AmpereOne
{0xc0, 0xac4, 1, ProductNames::ARM_Ampere_1A}, // AmpereOneA
{0xc0, 0xac5, 1, ProductNames::ARM_Ampere_1B}, // AmpereOneB
{0x4e, 0x010, 1, ProductNames::ARM_Olympus}, // Olympus
{0x4e, 0x004, 1, ProductNames::ARM_Carmel}, // Carmel
// Denver rated above A57 to match TX2 weirdness
{0x4e, 0x003, 1, ProductNames::ARM_Denver}, // Denver
@@ -227,6 +243,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x61, 0x024, 0, ProductNames::ARM_Icestorm_M1Pro}, // Apple Icestorm (M1 Pro)
{0x61, 0x022, 0, ProductNames::ARM_Icestorm_M1}, // Apple Icestorm (M1)
{0x41, 0xd8a, 1, ProductNames::ARM_C1Nano}, // C1-Nano
{0x41, 0xd80, 0, ProductNames::ARM_A520}, // A520
{0x41, 0xd46, 0, ProductNames::ARM_A510}, // A510
{0x41, 0xd06, 0, ProductNames::ARM_A65}, // A65
@@ -892,71 +909,71 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
Res.eax = FAMILY_IDENTIFIER;
Res.ecx = (1 << 0) | // LAHF/SAHF
(1 << 1) | // 0 = Single core product, 1 = multi core product
(0 << 2) | // SVM
(1 << 3) | // Extended APIC register space
(0 << 4) | // LOCK MOV CR0 means MOV CR8
(1 << 5) | // ABM instructions
(0 << 6) | // SSE4a
(0 << 7) | // Misaligned SSE mode
(1 << 8) | // PREFETCHW
(0 << 9) | // OS visible workaround support
(0 << 10) | // Instruction based sampling support
(0 << 11) | // XOP
(0 << 12) | // SKINIT
(0 << 13) | // Watchdog timer support
(0 << 14) | // Reserved
(0 << 15) | // Lightweight profiling support
(0 << 16) | // FMA4
(1 << 17) | // Translation cache extension
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(0 << 20) | // Reserved
(0 << 21) | // XOP-TBM
(0 << 22) | // Topology extensions support
(0 << 23) | // Core performance counter extensions
(0 << 24) | // NB performance counter extensions
(0 << 25) | // Reserved
(0 << 26) | // Data breakpoints extensions
(0 << 27) | // Performance TSC
(0 << 28) | // L2 perf counter extensions
(0 << 29) | // MONITORX
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.ecx = (1 << 0) | // LAHF/SAHF
(1 << 1) | // 0 = Single core product, 1 = multi core product
(0 << 2) | // SVM
(1 << 3) | // Extended APIC register space
(0 << 4) | // LOCK MOV CR0 means MOV CR8
(1 << 5) | // ABM instructions
(CTX->HostFeatures.SupportsSSE4a << 6) | // SSE4a
(0 << 7) | // Misaligned SSE mode
(1 << 8) | // PREFETCHW
(0 << 9) | // OS visible workaround support
(0 << 10) | // Instruction based sampling support
(0 << 11) | // XOP
(0 << 12) | // SKINIT
(0 << 13) | // Watchdog timer support
(0 << 14) | // Reserved
(0 << 15) | // Lightweight profiling support
(0 << 16) | // FMA4
(1 << 17) | // Translation cache extension
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(0 << 20) | // Reserved
(0 << 21) | // XOP-TBM
(0 << 22) | // Topology extensions support
(0 << 23) | // Core performance counter extensions
(0 << 24) | // NB performance counter extensions
(0 << 25) | // Reserved
(0 << 26) | // Data breakpoints extensions
(0 << 27) | // Performance TSC
(0 << 28) | // L2 perf counter extensions
(0 << 29) | // MONITORX
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual mode extensions
(1 << 2) | // Debugging extensions
(1 << 3) | // Page size extensions
(1 << 4) | // TSC
(1 << 5) | // MSR support
(1 << 6) | // PAE
(1 << 7) | // Machine Check Exception
(1 << 8) | // CMPXCHG8B
(1 << 9) | // APIC
(0 << 10) | // Reserved
(1 << 11) | // SYSCALL/SYSRET
(1 << 12) | // MTRR
(1 << 13) | // Page global extension
(1 << 14) | // Machine Check architecture
(1 << 15) | // CMOV
(1 << 16) | // Page attribute table
(1 << 17) | // Page-size extensions
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(1 << 20) | // NX
(0 << 21) | // Reserved
(1 << 22) | // MMXExt
(1 << 23) | // MMX
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(1 << 30) | // 3DNow! Extensions
(1 << 31); // 3DNow!
Res.edx = (1 << 0) | // FPU
(1 << 1) | // Virtual mode extensions
(1 << 2) | // Debugging extensions
(1 << 3) | // Page size extensions
(1 << 4) | // TSC
(1 << 5) | // MSR support
(1 << 6) | // PAE
(1 << 7) | // Machine Check Exception
(1 << 8) | // CMPXCHG8B
(1 << 9) | // APIC
(0 << 10) | // Reserved
(1 << 11) | // SYSCALL/SYSRET
(1 << 12) | // MTRR
(1 << 13) | // Page global extension
(1 << 14) | // Machine Check architecture
(1 << 15) | // CMOV
(1 << 16) | // Page attribute table
(1 << 17) | // Page-size extensions
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(1 << 20) | // NX
(0 << 21) | // Reserved
(1 << 22) | // MMXExt
(1 << 23) | // MMX
(1 << 24) | // FXSAVE/FXRSTOR
(1 << 25) | // FXSAVE/FXRSTOR Optimizations
(0 << 26) | // 1 gigabit pages
(SUPPORTS_RDTSCP << 27) | // RDTSCP
(0 << 28) | // Reserved
(1 << 29) | // Long Mode
(CTX->HostFeatures.Supports3DNow << 30) | // 3DNow! Extensions
(CTX->HostFeatures.Supports3DNow << 31); // 3DNow!
return Res;
}
@@ -0,0 +1,27 @@
// SPDX-License-Identifier: MIT
#include <Interface/Context/Context.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
namespace FEXCore {
ExecutableFileInfo::~ExecutableFileInfo() = default;
} // namespace FEXCore
namespace FEXCore::Context {
CodeCache::CodeCache(ContextImpl& CTX_)
: CTX(CTX_) {}
CodeCache::~CodeCache() = default;
void CodeCache::LoadData(Core::InternalThreadState& Thread, std::byte* MappedCacheFile, const ExecutableFileSectionInfo& GuestRIPLookup) {
// TODO
}
bool CodeCache::SaveData(Core::InternalThreadState& Thread, int fd, const ExecutableFileSectionInfo& SourceBinary, uint64_t SerializedBaseAddress) {
// TODO
return true;
}
} // namespace FEXCore::Context
+95 -178
View File
@@ -14,11 +14,11 @@ $end_info$
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/Frontend.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include <Interface/GDBJIT/GDBJIT.h>
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
@@ -57,20 +57,13 @@ $end_info$
#include <algorithm>
#include <array>
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <fcntl.h>
#include <functional>
#include <mutex>
#include <queue>
#include <shared_mutex>
#include <signal.h>
#include <stdio.h>
#include <string_view>
#include <sys/stat.h>
#include <type_traits>
#include <unistd.h>
#include <unordered_map>
#include <utility>
#include <xxhash.h>
@@ -78,10 +71,7 @@ namespace FEXCore::Context {
ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
: HostFeatures {Features}
, CPUID {this}
, IRCaptureCache {this} {
if (Config.CacheObjectCodeCompilation() != FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
CodeObjectCacheService = fextl::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
}
, CodeCache {*this} {
if (!Config.Is64BitMode()) {
// When operating in 32-bit mode, the virtual memory we care about is only the lower 32-bits.
Config.VirtualMemSize = 1ULL << 32;
@@ -103,14 +93,8 @@ ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
// Track atomic TSO emulation configuration.
UpdateAtomicTSOEmulationConfig();
}
ContextImpl::~ContextImpl() {
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
}
}
// Ensure X87 precision constraints are respected.
UpdateX87PrecisionConfig();
}
struct GetFrameBlockInfoResult {
@@ -139,6 +123,11 @@ bool ContextImpl::IsCurrentBlockSingleInst(FEXCore::Core::InternalThreadState* T
return InlineTail && InlineTail->SingleInst;
}
uint64_t ContextImpl::GetGuestBlockEntry(FEXCore::Core::InternalThreadState* Thread) {
auto [_, InlineTail] = GetFrameBlockInfo(Thread->CurrentFrame);
return InlineTail ? InlineTail->RIP : 0;
}
uint64_t ContextImpl::RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) {
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
@@ -349,36 +338,9 @@ bool ContextImpl::InitCore() {
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
// Set up the SignalDelegator config since core is initialized.
FEXCore::SignalDelegator::SignalDelegatorConfig SignalConfig {
.DispatcherBegin = Dispatcher->Start,
.DispatcherEnd = Dispatcher->End,
SignalDelegation->SetConfig(Dispatcher->MakeSignalDelegatorConfig());
.AbsoluteLoopTopAddress = Dispatcher->AbsoluteLoopTopAddress,
.AbsoluteLoopTopAddressFillSRA = Dispatcher->AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = Dispatcher->SignalHandlerReturnAddressRT,
.PauseReturnInstruction = Dispatcher->PauseReturnInstruction,
.ThreadPauseHandlerAddressSpillSRA = Dispatcher->ThreadPauseHandlerAddressSpillSRA,
.ThreadPauseHandlerAddress = Dispatcher->ThreadPauseHandlerAddress,
// Stop handlers.
.ThreadStopHandlerAddressSpillSRA = Dispatcher->ThreadStopHandlerAddressSpillSRA,
.ThreadStopHandlerAddress = Dispatcher->ThreadStopHandlerAddress,
// SRA information.
.SRAGPRCount = Dispatcher->GetSRAGPRCount(),
.SRAFPRCount = Dispatcher->GetSRAFPRCount(),
};
Dispatcher->GetSRAGPRMapping(SignalConfig.SRAGPRMapping);
Dispatcher->GetSRAFPRMapping(SignalConfig.SRAFPRMapping);
// Give this configuration to the SignalDelegator.
SignalDelegation->SetConfig(SignalConfig);
#ifndef _WIN32
#elif !defined(_M_ARM_64EC)
#if defined(_WIN32) && !defined(_M_ARM_64EC)
// WOW64 always needs the interrupt fault check to be enabled.
Config.NeedsPendingInterruptFaultCheck = true;
#endif
@@ -398,12 +360,6 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
// If it is the parent thread that died then just leave
// TODO: This doesn't make sense when the parent thread doesn't outlive its children
}
@@ -441,22 +397,6 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, const FEXC
Thread->CurrentFrame->State.gregs[X86State::REG_RSP] = StackPointer;
Thread->CurrentFrame->State.rip = InitialRIP;
// Set up default code segment.
// Default code segment indexes match the numbers that the Linux kernel uses.
Thread->CurrentFrame->State.cs_idx = 6 << 3;
auto &GDT = Thread->CurrentFrame->State.gdt[Thread->CurrentFrame->State.cs_idx >> 3];
Thread->CurrentFrame->State.SetGDTBase(&GDT, 0);
Thread->CurrentFrame->State.SetGDTLimit(&GDT, 0xF'FFFFU);
if (Config.Is64BitMode) {
GDT.L = 1; // L = Long Mode = 64-bit
GDT.D = 0; // D = Default Operand SIze = Reserved
}
else {
GDT.L = 0; // L = Long Mode = 32-bit
GDT.D = 1; // D = Default Operand Size = 32-bit
}
// Copy over the new thread state to the new object
if (NewThreadState) {
memcpy(&Thread->CurrentFrame->State, NewThreadState, sizeof(FEXCore::Core::CPUState));
@@ -520,12 +460,6 @@ void ContextImpl::OnCodeBufferAllocated(CPU::CodeBuffer& Buffer) {
void ContextImpl::ClearCodeCache(FEXCore::Core::InternalThreadState* Thread, bool NewCodeBuffer) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
if (CodeObjectCacheService) {
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
CodeSerialize::CodeObjectSerializeService::WaitForEmptyJobQueue(&Thread->ObjectCacheRefCounter);
}
if (NewCodeBuffer) {
// Allocate new CodeBuffer + L3 LookupCache and clear L1+L2 caches
Thread->CPUBackend->ClearCache();
@@ -579,7 +513,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
auto BlockInfo = Thread->FrontendDecoder->GetDecodedBlockInfo();
auto CodeBlocks = &BlockInfo->Blocks;
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode);
Thread->OpDispatcher->BeginFunction(GuestRIP, CodeBlocks, BlockInfo->TotalInstructionCount, BlockInfo->Is64BitMode,
AreMonoHacksActive() && MonoBackpatcherBlock.load(std::memory_order_relaxed) == GuestRIP);
const auto GPRSize = Thread->OpDispatcher->GetGPROpSize();
@@ -607,7 +542,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (InstsInBlock == 0) {
// Special case for an empty instruction block.
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry - GuestRIP));
}
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -637,11 +572,12 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->_GuestOpcode(InstAddress - GuestRIP);
}
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(ExistingCodePtr[0], ExistingCodePtr[1],
(uintptr_t)ExistingCodePtr - GuestRIP, DecodedInfo->InstSize);
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL || Block.ForceFullSMCDetection) {
auto ExistingCodePtr = reinterpret_cast<uint8_t*>(Block.Entry + BlockInstructionsLength);
auto InstAddressReg = Thread->OpDispatcher->_EntrypointOffset(GPRSize, InstAddress - GuestRIP);
std::array<uint8_t, 0x10> CodeOriginal;
memcpy(CodeOriginal.data(), ExistingCodePtr, DecodedInfo->InstSize);
auto CodeChanged = Thread->OpDispatcher->_ValidateCode(CodeOriginal, InstAddressReg, DecodedInfo->InstSize);
auto InvalidateCodeCond = Thread->OpDispatcher->CondJump(CodeChanged);
@@ -651,7 +587,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, InstAddress - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, InstAddress - GuestRIP));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -659,15 +595,21 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(NextOpBlock);
}
if (TableInfo && TableInfo->OpcodeDispatcher) {
auto Fn = TableInfo->OpcodeDispatcher;
if (TableInfo && TableInfo->OpcodeDispatcher.OpDispatch) {
auto Fn = TableInfo->OpcodeDispatcher.OpDispatch;
Thread->OpDispatcher->ResetHandledLock();
Thread->OpDispatcher->ResetDecodeFailure();
IR::ForceTSOMode ForceTSO =
BlockInForceTSOValidRange ?
(InstForceTSOIt != ForceTSOInstructions.end() && *InstForceTSOIt == InstAddress ? IR::ForceTSOMode::ForceEnabled :
IR::ForceTSOMode::ForceDisabled) :
IR::ForceTSOMode::NoOverride;
IR::ForceTSOMode ForceTSO = IR::ForceTSOMode::NoOverride;
if (BlockInForceTSOValidRange) {
if (InstForceTSOIt != ForceTSOInstructions.end() && *InstForceTSOIt == InstAddress) {
ForceTSO = IR::ForceTSOMode::ForceEnabled;
} else {
ForceTSO = IR::ForceTSOMode::ForceDisabled;
}
} else if (DecodedInfo->Flags & X86Tables::DecodeFlags::FLAG_FORCE_TSO) {
ForceTSO = IR::ForceTSOMode::ForceEnabled;
}
Thread->OpDispatcher->SetForceTSO(ForceTSO);
std::invoke(Fn, Thread->OpDispatcher, DecodedInfo);
if (Thread->OpDispatcher->HadDecodeFailure()) {
@@ -717,7 +659,8 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (NeedsBlockEnd) {
// We had some instructions. Early exit
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_InlineEntrypointOffset(GPRSize, Block.Entry + BlockInstructionsLength - GuestRIP));
break;
}
@@ -760,27 +703,10 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
// JIT Code object cache lookup
if (CodeObjectCacheService) {
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
if (CodeCacheEntry) {
auto CompiledCode = Thread->CPUBackend->RelocateJITObjectCode(GuestRIP, CodeCacheEntry);
if (CompiledCode) {
return {
.CompiledCode = {},
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.StartAddr = 0, // Unused
.Length = 0, // Unused
.NeedsAddGuestCodeRanges = false,
};
}
}
}
if (SourcecodeResolver && Config.GDBSymbols()) {
auto AOTIRCacheEntry = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
if (AOTIRCacheEntry.Entry && !AOTIRCacheEntry.Entry->ContainsCode) {
AOTIRCacheEntry.Entry->SourcecodeMap = SourcecodeResolver->GenerateMap(AOTIRCacheEntry.Entry->Filename, AOTIRCacheEntry.Entry->FileId);
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
MappedSection->FileInfo.SourcecodeMap = SourcecodeResolver->GenerateMap(MappedSection->FileInfo.Filename, MappedSection->FileInfo.FileId);
}
}
@@ -855,51 +781,44 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
if (Config.BlockJITNaming()) {
auto FragmentBasePtr = CompiledCode.BlockBegin;
if (DebugData) {
auto GuestRIPLookup = SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
auto GuestRIPLookup = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (DebugData->Subblocks.size()) {
for (auto& Subblock : DebugData->Subblocks) {
auto BlockBasePtr = FragmentBasePtr + Subblock.HostCodeOffset;
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, GuestRIP, Subblock.HostCodeSize);
}
}
} else {
if (GuestRIPLookup.Entry) {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, CompiledCode.Size, GuestRIPLookup.Entry->Filename,
GuestRIP - GuestRIPLookup.VAFileStart);
if (DebugData->Subblocks.size()) {
for (auto& Subblock : DebugData->Subblocks) {
auto BlockBasePtr = FragmentBasePtr + Subblock.HostCodeOffset;
if (GuestRIPLookup) {
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, CompiledCode.Size, GuestRIPLookup->FileInfo.Filename,
GuestRIP - GuestRIPLookup->FileStartVA);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, CompiledCode.Size);
Symbols.Register(Thread->SymbolBuffer.get(), BlockBasePtr, GuestRIP, Subblock.HostCodeSize);
}
}
} else {
if (GuestRIPLookup) {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, CompiledCode.Size, GuestRIPLookup->FileInfo.Filename,
GuestRIP - GuestRIPLookup->FileStartVA);
} else {
Symbols.Register(Thread->SymbolBuffer.get(), FragmentBasePtr, GuestRIP, CompiledCode.Size);
}
}
}
// Tell the object cache service to serialize the code if enabled
if (CodeObjectCacheService && Config.CacheObjectCodeCompilation == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_READWRITE && DebugData) {
CodeObjectCacheService->AsyncAddSerializationJob(
fextl::make_unique<CodeSerialize::AsyncJobHandler::SerializationJobData>(CodeSerialize::AsyncJobHandler::SerializationJobData {
.GuestRIP = GuestRIP,
.GuestCodeLength = Length,
.GuestCodeHash = 0,
.HostCodeBegin = CompiledCode.BlockBegin,
.HostCodeLength = CompiledCode.Size,
.HostCodeHash = 0,
.ThreadJobRefCount = &Thread->ObjectCacheRefCounter,
.Relocations = std::move(*DebugData->Relocations),
}));
if (Config.LibraryJITNaming() || Config.GDBSymbols()) {
auto MappedSection = SyscallHandler->LookupExecutableFileSection(*Thread, GuestRIP);
if (MappedSection) {
if (Config.LibraryJITNaming()) {
Symbols.RegisterNamedRegion(Thread->SymbolBuffer.get(), CodePtr, DebugData->HostCodeSize, MappedSection->FileInfo.Filename);
}
if (Config.GDBSymbols()) {
GDBJITRegister(MappedSection->FileInfo, MappedSection->FileStartVA, GuestRIP, (uintptr_t)CodePtr, *DebugData);
}
}
}
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(Thread, CompiledCode.BlockBegin, GuestRIP, StartAddr, Length, {}, DebugData.get(), false)) {
// Early exit
return (uintptr_t)CodePtr;
if (!CodeCache.IsGeneratingCache) {
Thread->CPUBackend->ClearRelocations();
}
if (NeedsAddGuestCodeRanges) {
@@ -978,23 +897,6 @@ void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* T
InvalidateGuestThreadCodeRange(Thread, Accumulator, Start, Length);
}
void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) {
if (!Thread) {
return;
}
if (!IsMemoryShared) {
IsMemoryShared = true;
UpdateAtomicTSOEmulationConfig();
if (Config.TSOAutoMigration) {
// Only the lookup cache is cleared here, so that old code can keep running until next compilation.
// This will leak previously compiled blocks until the CodeBuffer is cleared for some other reason.
Thread->LookupCache->ClearCache();
}
}
}
bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to "
"be unique_locked here");
@@ -1002,6 +904,10 @@ bool ContextImpl::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thre
return Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
void ContextImpl::ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
static_cast<ContextImpl*>(Frame->Thread->CTX)->SyscallHandler->InvalidateGuestCodeRange(Frame->Thread, GuestRIP, 1);
}
std::optional<CustomIRResult>
ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void* Creator, void* Data) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
@@ -1041,10 +947,10 @@ void ContextImpl::AddThunkTrampolineIRHandler(uintptr_t Entrypoint, uintptr_t Gu
if (GPRSize == IR::OpSize::i64Bit) {
IR::Ref R = emit->_StoreRegister(emit->Constant(Entrypoint), GPRSize);
R->Reg = IR::PhysicalRegister(IR::GPRFixedClass, X86State::REG_R11).Raw;
R->Reg = IR::PhysicalRegister(IR::RegClass::GPRFixed, X86State::REG_R11).Raw;
} else {
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
emit->_StoreContextFPR(GPRSize, emit->_VCastFromGPR(IR::OpSize::i64Bit, IR::OpSize::i64Bit, emit->Constant(Entrypoint)),
offsetof(Core::CPUState, mm[0][0]));
}
emit->_ExitFunction(IR::OpSize::i64Bit, emit->Constant(GuestThunkEntrypoint), IR::BranchHint::None, emit->Invalid(), emit->Invalid());
},
@@ -1075,25 +981,36 @@ void ContextImpl::RemoveForceTSOInformation(uint64_t Address, uint64_t Size) {
ForceTSOInstructions.erase(ForceTSOInstructions.lower_bound(Address), ForceTSOInstructions.upper_bound(Address + Size));
}
void ContextImpl::RemoveCustomIREntrypoint(uintptr_t Entrypoint) {
void ContextImpl::MarkMonoBackpatcherBlock(uint64_t BlockEntry) {
MonoBackpatcherBlock.store(BlockEntry, std::memory_order_relaxed);
}
void ContextImpl::RemoveCustomIREntrypoint(FEXCore::Core::InternalThreadState* Thread, uintptr_t Entrypoint) {
LOGMAN_THROW_A_FMT(Config.Is64BitMode || !(Entrypoint >> 32), "64-bit Entrypoint in 32-bit mode {:x}", Entrypoint);
std::scoped_lock lk(CustomIRMutex);
InvalidatedEntryAccumulator Accumulator;
InvalidateGuestCodeRange(nullptr, Accumulator, Entrypoint, 1);
CustomIRHandlers.erase(Entrypoint);
HasCustomIRHandlers = !CustomIRHandlers.empty();
SyscallHandler->InvalidateGuestCodeRange(Thread, Entrypoint, 1);
}
IR::AOTIRCacheEntry* ContextImpl::LoadAOTIRCacheEntry(const fextl::string& filename) {
auto rv = IRCaptureCache.LoadAOTIRCacheEntry(filename);
return rv;
}
void ContextImpl::MonoBackpatcherWrite(FEXCore::Core::CpuStateFrame* Frame, uint8_t Size, uint64_t Address, uint64_t Value) {
auto Thread = Frame->Thread;
auto CTX = static_cast<ContextImpl*>(Thread->CTX);
{
auto lk = GuardSignalDeferringSection(CTX->CodeInvalidationMutex, Thread);
void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry* Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
if (Size == 8) {
*reinterpret_cast<uint64_t*>(Address) = Value;
} else if (Size == 4) {
*reinterpret_cast<uint32_t*>(Address) = Value;
} else {
ERROR_AND_DIE_FMT("Unexpected write size for backpatcher: {}", Size);
}
}
CTX->SyscallHandler->InvalidateGuestCodeRange(Thread, Address, Size);
}
void ContextImpl::ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) {
@@ -2,6 +2,7 @@
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
@@ -16,14 +17,20 @@
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <CodeEmitter/Emitter.h>
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#endif
#include <array>
#include <atomic>
#include <bit>
#include <condition_variable>
#include <csignal>
#include <cstring>
#include <signal.h>
namespace FEXCore::CPU {
@@ -134,8 +141,6 @@ void Dispatcher::EmitDispatcher() {
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
ARMEmitter::BiDirectionalLabel FullLookup {};
ARMEmitter::BiDirectionalLabel CallBlock {};
Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
@@ -143,23 +148,33 @@ void Dispatcher::EmitDispatcher() {
// Load in our RIP
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
#ifdef _M_ARM_64EC
// Clobbers TMP1/2
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::ForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
#endif
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbnz(ARMEmitter::Size::i32Bit, TMP1, &CompileSingleStep);
// L1 Cache
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP1, TMP1, 0);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(TMP4);
// L1C check failed, do a full lookup
Bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
@@ -287,40 +302,10 @@ void Dispatcher::EmitDispatcher() {
br(TMP1);
}
#ifdef _M_ARM_64EC
// Clobbers TMP1/2
auto EmitECExitCheck = [&]() {
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::ForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
str(REG_CALLRET_SP, STATE_PTR(CpuStateFrame, State.callret_sp));
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
};
#endif
// Need to create the block
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
EmitECExitCheck();
#endif
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -355,10 +340,6 @@ void Dispatcher::EmitDispatcher() {
{
Bind(&CompileSingleStep);
#ifdef _M_ARM_64EC
EmitECExitCheck();
#endif
EmitSignalGuardedRegion([&]() {
SpillStaticRegs(TMP1);
@@ -569,8 +550,8 @@ void Dispatcher::EmitDispatcher() {
FABI_F80_I16_I32_PTR,
FABI_F32_I16_F80_PTR,
FABI_F64_I16_F80_PTR,
FABI_F64_I16_F64_PTR,
FABI_F64_I16_F64_F64_PTR,
FABI_F64_F64_PTR,
FABI_F64_F64_F64_PTR,
FABI_I16_I16_F80_PTR,
FABI_I32_I16_F80_PTR,
FABI_I64_I16_F80_PTR,
@@ -578,7 +559,7 @@ void Dispatcher::EmitDispatcher() {
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_F64x2_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
}};
@@ -777,7 +758,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
mov(VABI1.Q(), VTMP1.Q());
}
mov(ARMEmitter::XReg::x1, STATE);
@@ -810,7 +791,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
FillF64Result();
} break;
case FABI_F64_I16_F64_PTR: {
case FABI_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -820,18 +801,17 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, uint64_t>(FallbackPointerReg);
GenerateIndirectRuntimeCall<double, double, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
FillF64Result();
} break;
case FABI_F64_I16_F64_F64_PTR: {
case FABI_F64_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -844,10 +824,9 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
fmov(VABI2.D(), VTMP2.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, double, uint64_t>(FallbackPointerReg);
GenerateIndirectRuntimeCall<double, double, double, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
@@ -1007,7 +986,7 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
FillF80x2Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
case FABI_F64x2_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -1016,14 +995,13 @@ uint64_t Dispatcher::GenerateABICall(FallbackABI ABI) {
SpillForABICall(CTX->HostFeatures.SupportsPreserveAllABI, TMP3, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
mov(ARMEmitter::XReg::x1, STATE);
mov(ARMEmitter::XReg::x0, STATE);
if (!TMP_ABIARGS) {
fmov(VABI1.D(), VTMP1.D());
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
// GenerateIndirectRuntimeCall<FEXCore::VectorScalarF64Pair, uint16_t, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
// GenerateIndirectRuntimeCall<FEXCore::VectorScalarF64Pair, FEXCore::VectorRegType, uint64_t>(FallbackPointerReg);
} else {
blr(FallbackPointerReg);
}
@@ -1125,6 +1103,53 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
}
}
SignalDelegatorConfig Dispatcher::MakeSignalDelegatorConfig() const {
// PF/AF are the final two SRA registers. We only want GPRs
const auto GPRCount = uint16_t(StaticRegisters.size() - 2);
const auto FPRCount = uint16_t(StaticFPRegisters.size());
const auto GetSRAGPRMapping = [GPRCount, this] {
SignalDelegatorConfig::SRAIndexMapping Mapping {};
for (size_t i = 0; i < GPRCount; ++i) {
Mapping[i] = StaticRegisters[i].Idx();
}
return Mapping;
};
const auto GetSRAFPRMapping = [FPRCount, this] {
SignalDelegatorConfig::SRAIndexMapping Mapping {};
for (size_t i = 0; i < FPRCount; ++i) {
Mapping[i] = StaticFPRegisters[i].Idx();
}
return Mapping;
};
return FEXCore::SignalDelegatorConfig {
.DispatcherBegin = Start,
.DispatcherEnd = End,
.AbsoluteLoopTopAddress = AbsoluteLoopTopAddress,
.AbsoluteLoopTopAddressFillSRA = AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = SignalHandlerReturnAddressRT,
.PauseReturnInstruction = PauseReturnInstruction,
.ThreadPauseHandlerAddressSpillSRA = ThreadPauseHandlerAddressSpillSRA,
.ThreadPauseHandlerAddress = ThreadPauseHandlerAddress,
// Stop handlers.
.ThreadStopHandlerAddressSpillSRA = ThreadStopHandlerAddressSpillSRA,
.ThreadStopHandlerAddress = ThreadStopHandlerAddress,
// SRA information.
.SRAGPRCount = GPRCount,
.SRAFPRCount = FPRCount,
.SRAGPRMapping = GetSRAGPRMapping(),
.SRAFPRMapping = GetSRAFPRMapping(),
};
}
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl* CTX) {
return fextl::make_unique<Dispatcher>(CTX);
}
@@ -2,25 +2,18 @@
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/fextl/memory.h>
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#endif
#include <array>
#include <cstddef>
#include <cstdint>
#include <signal.h>
#include <stddef.h>
#include <stack>
#include <tuple>
namespace FEXCore {
struct GuestSigAction;
}
struct SignalDelegatorConfig;
} // namespace FEXCore
namespace FEXCore::Core {
struct CpuStateFrame;
@@ -42,6 +35,32 @@ public:
Dispatcher(FEXCore::Context::ContextImpl* ctx);
~Dispatcher();
void InitThreadPointers(FEXCore::Core::InternalThreadState* Thread);
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame);
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame, false);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
#endif
SignalDelegatorConfig MakeSignalDelegatorConfig() const;
protected:
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame, bool SingleInst);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
private:
/**
* @name Dispatch Helper functions
* @{ */
@@ -59,62 +78,14 @@ public:
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
uint64_t IntCallbackReturnAddress {};
uint64_t PauseReturnInstruction {};
std::array<uint64_t, FallbackABI::FABI_UNKNOWN> ABIPointers {};
/** @} */
uint64_t Start {};
uint64_t End {};
void InitThreadPointers(FEXCore::Core::InternalThreadState* Thread);
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame);
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame, false);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
#endif
uint16_t GetSRAGPRCount() const {
// PF/AF are the final two SRA registers.
// Only return the SRA for GPRs.
return StaticRegisters.size() - 2;
}
uint16_t GetSRAFPRCount() const {
return StaticFPRegisters.size();
}
void GetSRAGPRMapping(uint8_t Mapping[16]) const {
for (size_t i = 0; i < StaticRegisters.size() - 2; ++i) {
Mapping[i] = StaticRegisters[i].Idx();
}
}
void GetSRAFPRMapping(uint8_t Mapping[16]) const {
for (size_t i = 0; i < StaticFPRegisters.size(); ++i) {
Mapping[i] = StaticFPRegisters[i].Idx();
}
}
protected:
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame, bool SingleInst);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
private:
// Long division helpers
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
+169 -46
View File
@@ -71,7 +71,23 @@ Decoder::Decoder(FEXCore::Core::InternalThreadState* Thread)
: Thread {Thread}
, CTX {static_cast<FEXCore::Context::ContextImpl*>(Thread->CTX)}
, OSABI {CTX->SyscallHandler ? CTX->SyscallHandler->GetOSABI() : FEXCore::HLE::SyscallOSABI::OS_UNKNOWN}
, PoolObject {CTX->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {}
, PoolObject {CTX->FrontendAllocator, sizeof(FEXCore::X86Tables::DecodedInst) * DefaultDecodedBufferSize} {
FEX_CONFIG_OPT(ReducedPrecision, X87REDUCEDPRECISION);
if (ReducedPrecision) {
X87Table = &FEXCore::X86Tables::X87F64Ops;
} else {
X87Table = &FEXCore::X86Tables::X87F80Ops;
}
if (CTX->HostFeatures.SupportsAVX && CTX->HostFeatures.SupportsSVE256) {
VEXTable = &FEXCore::X86Tables::VEXTableOps;
VEXTableGroup = &FEXCore::X86Tables::VEXTableGroupOps;
} else if (CTX->HostFeatures.SupportsAVX) {
VEXTable = &FEXCore::X86Tables::VEXTableOps_AVX128;
VEXTableGroup = &FEXCore::X86Tables::VEXTableGroupOps_AVX128;
}
}
bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
// Treat FEX-internal X86 callbacks as always executable
@@ -83,6 +99,7 @@ bool Decoder::CheckRangeExecutable(uint64_t Address, uint64_t Size) {
auto RangeInfo = CTX->SyscallHandler->QueryGuestExecutableRange(Thread, Address);
ExecutableRangeBase = RangeInfo.Base;
ExecutableRangeEnd = RangeInfo.Base + RangeInfo.Size;
ExecutableRangeWritable = RangeInfo.Writable;
if (RangeInfo.Size == 0) {
return false;
@@ -242,13 +259,13 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
if (HasSIB) {
FEXCore::X86Tables::SIBDecoded SIB;
if (DecodeInst->DecodedSIB) {
if (DecodeInst->Flags & DecodeFlags::FLAG_DECODED_SIB) {
SIB.Hex = DecodeInst->SIB;
} else {
// Haven't yet grabbed SIB, pull it now
DecodeInst->SIB = ReadByte();
SIB.Hex = DecodeInst->SIB;
DecodeInst->DecodedSIB = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_SIB;
}
// If the SIB base is 0b101, aka BP or R13 then we have a 32bit displacement
@@ -314,6 +331,13 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
}
bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options) {
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
// TODO: Move this in to `NormalOpHeader`, Dispatch tables have a bug currently where some subtables don't inherit flags correctly.
// Can be seen by running FEX asm tests if this is removed.
return NormalOp(&Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0], Op);
}
DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
@@ -377,9 +401,9 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// If we require ModRM and haven't decoded it yet, do it now
// Some instructions have to read modrm upfront, others do it later
if (HasMODRM && !DecodeInst->DecodedModRM) {
if (HasMODRM && !(DecodeInst->Flags & DecodeFlags::FLAG_DECODED_MODRM)) {
DecodeInst->ModRM = ReadByte();
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
}
// New instruction size decoding
@@ -412,9 +436,8 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_16BIT);
DestSize = 2;
} else if ((HasXMMDst || HasMMDst || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
} else if ((HasXMMDst || HasMMDst || BlockInfo.Is64BitMode) && (HasWideningDisplacement || DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_64BIT);
DestSize = 8;
} else {
@@ -441,9 +464,8 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
// See table 1-2. Operand-Size Overrides for this decoding
// If the default operating mode is 32bit and we have the operand size flag then the operating size drops to 16bit
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
} else if ((HasXMMSrc || HasMMSrc || BlockInfo.Is64BitMode) &&
(HasWideningDisplacement || SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
} else if ((HasXMMSrc || HasMMSrc || BlockInfo.Is64BitMode) && (HasWideningDisplacement || SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BIT ||
SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_64BITDEF)) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_64BIT);
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_32BIT);
@@ -612,11 +634,20 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
Literal = static_cast<int32_t>(Literal);
}
DecodeInst->Src[CurrentSrc].Data.Literal.Size = DestSize;
DecodeInst->Src[CurrentSrc].Data.Literal.SignExtend = true;
}
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
++CurrentSrc;
if (Bytes == 8) [[unlikely]] {
DecodeInst->Src[CurrentSrc].Data.Literal.Size = 4;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal >> 32;
}
Bytes = 0;
DecodeInst->Src[CurrentSrc].Type = DecodedOperand::OpType::Literal;
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining", DecodeInst->PC,
@@ -626,7 +657,7 @@ bool Decoder::NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op,
}
bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op) {
DecodeInst->OP = Op;
DecodeInst->OPRaw = DecodeInst->OP = Op;
DecodeInst->TableInfo = Info;
if (Info->Type == FEXCore::X86Tables::TYPE_UNKNOWN) {
@@ -642,10 +673,13 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
// A normal instruction is the most likely.
if (Info->Type == FEXCore::X86Tables::TYPE_INST) [[likely]] {
return NormalOp(Info, Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) [[unlikely]] {
// Dispatcher Op.
return NormalOp(&Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0], Op);
} else if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_11) {
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
@@ -662,24 +696,24 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
constexpr uint16_t PF_F2 = 3;
uint16_t PrefixType = PF_NONE;
if (DecodeInst->LastEscapePrefix == 0xF3) {
if (LastEscapePrefix == 0xF3) {
PrefixType = PF_F3;
} else if (DecodeInst->LastEscapePrefix == 0xF2) {
} else if (LastEscapePrefix == 0xF2) {
PrefixType = PF_F2;
} else if (DecodeInst->LastEscapePrefix == 0x66) {
} else if (LastEscapePrefix == 0x66) {
PrefixType = PF_66;
}
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
uint16_t LocalOp = OPD(Info->Type, PrefixType, ModRM.reg);
FEXCore::X86Tables::X86InstInfo* LocalInfo = &SecondInstGroupOps[LocalOp];
const FEXCore::X86Tables::X86InstInfo* LocalInfo = &SecondInstGroupOps[LocalOp];
#undef OPD
if (LocalInfo->Type == FEXCore::X86Tables::TYPE_SECOND_GROUP_MODRM && ModRM.mod == 0b11) {
// Everything in this group is privileged instructions aside from XGETBV
@@ -700,11 +734,16 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
uint16_t X87Op = ((Op - 0xD8) << 8) | ModRMByte;
return NormalOp(&X87Ops[X87Op], X87Op);
return NormalOp(&(*X87Table)[X87Op], X87Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
if (!VEXTable) {
// AVX not enabled.
return false;
}
uint16_t map_select = 1;
uint16_t pp = 0;
const uint8_t Byte1 = ReadByte();
@@ -742,7 +781,6 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
return false;
}
}
@@ -752,13 +790,13 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
Op = OPD(map_select, pp, VEXOp);
#undef OPD
FEXCore::X86Tables::X86InstInfo* LocalInfo = &VEXTableOps[Op];
const FEXCore::X86Tables::X86InstInfo* LocalInfo = &(*VEXTable)[Op];
if (LocalInfo->Type >= FEXCore::X86Tables::TYPE_VEX_GROUP_12 && LocalInfo->Type <= FEXCore::X86Tables::TYPE_VEX_GROUP_17) {
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
@@ -766,7 +804,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
Op = OPD(LocalInfo->Type, pp, ModRM.reg);
#undef OPD
return NormalOp(&VEXTableGroupOps[Op], Op, options);
return NormalOp(&(*VEXTableGroup)[Op], Op, options);
} else {
return NormalOp(LocalInfo, Op, options);
}
@@ -782,6 +820,7 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
bool Decoder::DecodeInstructionImpl(uint64_t PC) {
InstructionSize = 0;
LastEscapePrefix = 0;
Instruction.fill(0);
DecodeInst = &DecodedBuffer[DecodedSize];
@@ -803,7 +842,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
// Decode ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
DecodeInst->DecodedModRM = true;
DecodeInst->Flags |= DecodeFlags::FLAG_DECODED_MODRM;
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
@@ -842,7 +881,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
uint16_t LocalOp = (Prefix << 8) | ReadByte();
bool NoOverlay66 = (FEXCore::X86Tables::H0F38TableOps[LocalOp].Flags & InstFlags::FLAGS_NO_OVERLAY66) != 0;
if (DecodeInst->LastEscapePrefix == 0x66 && NoOverlay66) { // Operand Size
if (LastEscapePrefix == 0x66 && NoOverlay66) { // Operand Size
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather than modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
@@ -858,7 +897,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
constexpr uint16_t PF_3A_REX = (1 << 1);
uint16_t Prefix = PF_3A_NONE;
if (DecodeInst->LastEscapePrefix == 0x66) { // Operand Size
if (LastEscapePrefix == 0x66) { // Operand Size
Prefix = PF_3A_66;
}
@@ -884,17 +923,17 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
if (NoOverlay) { // This section of the table ignores prefix extention
return NormalOpHeader(&FEXCore::X86Tables::SecondBaseOps[EscapeOp], EscapeOp);
} else if (DecodeInst->LastEscapePrefix == 0xF3) { // REP
} else if (LastEscapePrefix == 0xF3) { // REP
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REP_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepModOps[EscapeOp], EscapeOp);
} else if (DecodeInst->LastEscapePrefix == 0xF2) { // REPNE
} else if (LastEscapePrefix == 0xF2) { // REPNE
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_REPNE_PREFIX;
return NormalOpHeader(&FEXCore::X86Tables::RepNEModOps[EscapeOp], EscapeOp);
} else if (DecodeInst->LastEscapePrefix == 0x66 && !NoOverlay66) { // Operand Size
} else if (LastEscapePrefix == 0x66 && !NoOverlay66) { // Operand Size
// Remove prefix so it doesn't effect calculations.
// This is only an escape prefix rather tan modifier now
DecodeInst->Flags &= ~DecodeFlags::FLAG_OPERAND_SIZE;
@@ -910,7 +949,7 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
}
case 0x66: // Operand Size prefix
DecodeInst->Flags |= DecodeFlags::FLAG_OPERAND_SIZE;
DecodeInst->LastEscapePrefix = Op;
LastEscapePrefix = Op;
DecodeFlags::PushOpAddr(&DecodeInst->Flags, DecodeFlags::FLAG_OPERAND_SIZE_LAST);
break;
case 0x67: // Address Size override prefix
@@ -941,11 +980,11 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
break;
case 0xF2: // REPNE prefix
DecodeInst->Flags |= DecodeFlags::FLAG_REPNE_PREFIX;
DecodeInst->LastEscapePrefix = Op;
LastEscapePrefix = Op;
break;
case 0xF3: // REP prefix
DecodeInst->Flags |= DecodeFlags::FLAG_REP_PREFIX;
DecodeInst->LastEscapePrefix = Op;
LastEscapePrefix = Op;
break;
case 0x64: // FS prefix
DecodeInst->Flags = (DecodeInst->Flags & ~FEXCore::X86Tables::DecodeFlags::FLAG_SEGMENTS) | DecodeFlags::FLAG_FS_PREFIX;
@@ -955,7 +994,10 @@ bool Decoder::DecodeInstructionImpl(uint64_t PC) {
break;
default:
[[likely]] { // Default base table
auto Info = &FEXCore::X86Tables::BaseOps[Op];
const X86InstInfo* Info = &FEXCore::X86Tables::BaseOps[Op];
if (Info->Type == FEXCore::X86Tables::TYPE_ARCH_DISPATCHER) {
Info = &Info->OpcodeDispatcher.Indirect[BlockInfo.Is64BitMode ? 1 : 0];
}
if (Info->Type == FEXCore::X86Tables::TYPE_REX_PREFIX) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_PREFIX;
@@ -1007,11 +1049,27 @@ Decoder::DecodedBlockStatus Decoder::DecodeInstruction(uint64_t PC) {
DecodeInst->TableInfo = nullptr;
DecodeInst->InstSize = 0;
return ErrorDuringDecoding ? DecodedBlockStatus::INVALID_INST : DecodedBlockStatus::NOEXEC_INST;
} else if (!DecodeInst->TableInfo || !DecodeInst->TableInfo->OpcodeDispatcher) {
} else if (!DecodeInst->TableInfo || (DecodeInst->TableInfo->Type == TYPE_INST && !DecodeInst->TableInfo->OpcodeDispatcher.OpDispatch)) {
// If there wasn't an error during decoding but we have no dispatcher for the instruction then claim invalid instruction.
return DecodedBlockStatus::INVALID_INST;
}
if (CTX->AreMonoHacksActive()) {
// Unity uses a standard SPSC ringbuffer with cached read/write pointers and thread waiting flags at the following
// offsets, which are consistent between 32-bit and 64-bit Unity versions from 2015 onwards.
auto IsKnownAtomicDisplacement = [](uint64_t Displacement) {
return Displacement == 0x80 || Displacement == 0x84 || Displacement == 0xC0 || Displacement == 0xC4;
};
if (DecodeInst->OP == 0x8b && DecodeInst->Src[0].IsGPRIndirect() &&
IsKnownAtomicDisplacement(DecodeInst->Src[0].Data.GPRIndirect.Displacement)) {
DecodeInst->Flags |= X86Tables::DecodeFlags::FLAG_FORCE_TSO;
}
if (DecodeInst->OP == 0x89 && DecodeInst->Dest.IsGPRIndirect() && IsKnownAtomicDisplacement(DecodeInst->Dest.Data.GPRIndirect.Displacement)) {
DecodeInst->Flags |= X86Tables::DecodeFlags::FLAG_FORCE_TSO;
}
}
return DecodedBlockStatus::SUCCESS;
}
@@ -1027,6 +1085,13 @@ void Decoder::BranchTargetInMultiblockRange() {
const auto InstEnd = DecodeInst->PC + DecodeInst->InstSize;
if (DecodeInst->TableInfo->Flags & FEXCore::X86Tables::InstFlags::FLAGS_CALL) {
if (ExecutableRangeWritable && CTX->AreMonoHacksActive()) {
// Mono generated code often contains noreturn calls with garbage following them, and calls are always backpatched
// after CIL compilation leading to n recompiles for a multiblock with n calls. Choose to minimize stutters over
// raw performance and disable tracking past calls for mono generated code.
return;
}
AddBranchTarget(InstEnd);
BlockInfo.EntryPoints.emplace(InstEnd);
return;
@@ -1058,10 +1123,16 @@ void Decoder::BranchTargetInMultiblockRange() {
TargetRIP &= 0xFFFFFFFFU;
}
if (Conditional) {
// If we are conditional then a target can be the instruction past the conditional instruction
AddBranchTarget(InstEnd);
}
// If the target RIP is x86 code within the symbol ranges then we are golden
// Forbid cross-page branches to both avoid massive (range-wise) code blocks in highly fragmented code and trying to decode unmapped branch targets
bool ValidMultiblockMember =
TargetRIP >= SymbolMinAddress && TargetRIP < std::min(FEXCore::AlignUp(InstEnd, FEXCore::Utils::FEX_PAGE_SIZE), SymbolMaxAddress);
// Forbid distant branches to have the cost code better match the guest code layout, avoiding massive (range-wise) code
// blocks in highly fragmented guest code. Such branches are often not-taken branches to garbage in obfuscated code.
constexpr uint64_t MAX_FORWARD_BRANCH_DIST = FEXCore::Utils::FEX_PAGE_SIZE * 4;
bool ValidMultiblockMember = TargetRIP >= SymbolMinAddress && TargetRIP < std::min(InstEnd + MAX_FORWARD_BRANCH_DIST, SymbolMaxAddress);
#ifdef _M_ARM_64EC
ValidMultiblockMember = ValidMultiblockMember && !RtlIsEcCode(TargetRIP);
@@ -1072,9 +1143,6 @@ void Decoder::BranchTargetInMultiblockRange() {
if (Conditional) {
MaxCondBranchForward = std::max(MaxCondBranchForward, TargetRIP);
MaxCondBranchBackwards = std::min(MaxCondBranchBackwards, TargetRIP);
// If we are conditional then a target can be the instruction past the conditional instruction
AddBranchTarget(InstEnd);
}
AddBranchTarget(TargetRIP);
@@ -1085,6 +1153,60 @@ void Decoder::BranchTargetInMultiblockRange() {
}
}
bool Decoder::IsBranchMonoTailcall(uint64_t NumInstructions) const {
// While the mono call backpatching block can easily be detected due it being the only one to contain SMC-faulting
// atomics, that can't be said for the tailcall jump backpatcher which has changed several times across versions and
// can be partially inlined. To work around this, instead detect the tailcall site itself and force full non-signal-based
// SMC detection for that single block.
if (!ExecutableRangeWritable) {
// We only care about jitted code
return false;
}
// See mini-{amd64,x86}.c in the mono codebase, specifically where METHOD_JUMP patches are emitted.
if (GetGPROpSize() == IR::OpSize::i32Bit) {
// Matches:
// LEAVE
// <none> / NOP / MOV EAX, EAX / LEA EBP, [EBP+0]
// JMP imm32
if (DecodeInst->OP != 0xE9 || NumInstructions < 2) {
return false;
}
auto PrevInst = std::prev(DecodeInst);
if (PrevInst->OP == 0xC9) {
return true;
}
if (NumInstructions < 3 || std::prev(PrevInst)->OP != 0xC9) {
return false;
}
return PrevInst->OP == 0x90 || (PrevInst->OP == 0x8B && PrevInst->ModRM == 0xC0) ||
(PrevInst->OP == 0x8D && PrevInst->ModRM == 0x6D && PrevInst->Src[1].IsLiteral() && PrevInst->Src[1].Literal() == 0);
} else {
FEXCore::X86Tables::ModRMDecoded ModRM;
ModRM.Hex = DecodeInst->ModRM;
if (DecodeInst->OPRaw == 0xFF && ModRM.reg == 4 && DecodeInst->Src[0].IsGPR()) {
if (DecodeInst->Src[0].Data.GPR.GPR == FEXCore::X86State::REG_RAX) {
// Found in versions of mono from 2024 onwards - matches:
// REX.W JMP rax
return (DecodeInst->Flags & (DecodeFlags::FLAG_REX_PREFIX | DecodeFlags::FLAG_REX_WIDENING | DecodeFlags::FLAG_REX_XGPR_B |
DecodeFlags::FLAG_REX_XGPR_X | DecodeFlags::FLAG_REX_XGPR_R)) ==
(DecodeFlags::FLAG_REX_PREFIX | DecodeFlags::FLAG_REX_WIDENING);
} else if (NumInstructions > 1 && DecodeInst->Src[0].Data.GPR.GPR == FEXCore::X86State::REG_R11) {
// Found in older versions of mono - match:
// MOV r11, imm64
// JMP r11
auto PrevInst = std::prev(DecodeInst);
return PrevInst->OP == 0xBB && PrevInst->Dest.IsGPR() && PrevInst->Dest.Data.GPR.GPR == FEXCore::X86State::REG_R11;
}
}
}
return false;
}
bool Decoder::InstCanContinue() const {
if (DecodeInst->PC + DecodeInst->InstSize == NextBlockStartAddress) {
return false;
@@ -1187,7 +1309,7 @@ const uint8_t* Decoder::AdjustAddrForSpecialRegion(const uint8_t* _InstStream, u
return _InstStream - EntryPoint + RIP;
}
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* _InstStream, uint64_t PC, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
BlockInfo.TotalInstructionCount = 0;
BlockInfo.Blocks.clear();
@@ -1199,8 +1321,8 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thre
DecodedBuffer = PoolObject.ReownOrClaimBuffer();
// Decode operating mode from thread's CS segment.
const auto CSSegment = Thread->CurrentFrame->State.gdt[Thread->CurrentFrame->State.cs_idx >> 3];
BlockInfo.Is64BitMode = CSSegment.L == 1;
const auto CSSegment = Core::CPUState::GetSegmentFromIndex(Thread->CurrentFrame->State, Thread->CurrentFrame->State.cs_idx);
BlockInfo.Is64BitMode = CSSegment->L == 1;
LOGMAN_THROW_A_FMT(BlockInfo.Is64BitMode == CTX->Config.Is64BitMode, "Expected operating mode to not change at runtime!");
// XXX: Load symbol data
@@ -1345,6 +1467,7 @@ void Decoder::DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thre
// If the branch target is within our multiblock range then we can keep going on
// We don't want to short circuit this since we want to calculate our ranges still
// NOTE: This will invalidate BlockIt, this is fine as we immediately break from the loop and EraseBlock cannot be true
BlockIt->ForceFullSMCDetection = CTX->AreMonoHacksActive() && IsBranchMonoTailcall(BlockIt->NumInstructions);
BranchTargetInMultiblockRange();
}
+16 -4
View File
@@ -4,18 +4,21 @@
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/IR/IR.h"
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <cstddef>
#include <cstdint>
#include <stddef.h>
#include <optional>
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::HLE {
enum class SyscallOSABI;
}
namespace FEXCore::Frontend {
class Decoder final {
@@ -34,6 +37,7 @@ public:
FEXCore::X86Tables::DecodedInst* DecodedInstructions;
DecodedBlockStatus BlockStatus;
bool IsEntryPoint {};
bool ForceFullSMCDetection {};
};
struct DecodedBlockInformation final {
@@ -45,7 +49,7 @@ public:
};
Decoder(FEXCore::Core::InternalThreadState* Thread);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState *Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
void DecodeInstructionsAtEntry(FEXCore::Core::InternalThreadState* Thread, const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst);
const DecodedBlockInformation* GetDecodedBlockInfo() const {
return &BlockInfo;
@@ -86,6 +90,7 @@ private:
DecodedBlockStatus DecodeInstruction(uint64_t PC);
void BranchTargetInMultiblockRange();
bool IsBranchMonoTailcall(uint64_t NumInstructions) const;
bool InstCanContinue() const;
void AddBranchTarget(uint64_t Target);
@@ -109,6 +114,7 @@ private:
uint64_t ExecutableRangeBase {};
uint64_t ExecutableRangeEnd {};
bool ExecutableRangeWritable {};
bool HitNonExecutableRange {};
const uint8_t* InstStream {};
@@ -119,6 +125,7 @@ private:
static constexpr size_t MAX_INST_SIZE = 15;
uint8_t InstructionSize {};
std::array<uint8_t, MAX_INST_SIZE> Instruction;
uint8_t LastEscapePrefix {};
FEXCore::X86Tables::DecodedInst* DecodeInst;
// This is for multiblock data tracking
@@ -147,6 +154,11 @@ private:
&FEXCore::Frontend::Decoder::DecodeModRM_16,
};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_X87_TABLE_SIZE>* X87Table;
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_TABLE_SIZE>* VEXTable {};
const std::array<X86Tables::X86InstInfo, X86Tables::MAX_VEX_GROUP_TABLE_SIZE>* VEXTableGroup {};
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
};
} // namespace FEXCore::Frontend
@@ -2,11 +2,13 @@
#pragma once
#include "Common/SoftFloat.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/SHMStats.h>
#include <FEXCore/Config/Config.h>
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW, bool Force80BitPrecision = false) {
@@ -77,6 +79,12 @@ struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static VectorRegType handle8(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame};
auto Context = static_cast<Context::ContextImpl*>(Frame->Thread->CTX);
auto ReducedPrecisionMode = Context->Config.x87ReducedPrecision;
auto StrictReducedPrecisionMode = Context->Config.x87StrictReducedPrecision;
if (!ReducedPrecisionMode || StrictReducedPrecisionMode) {
return X80SoftFloat::FromF64_PreserveNaN(&State.State, src);
}
return X80SoftFloat(&State.State, src);
}
};
@@ -115,6 +123,12 @@ struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, VectorRegType src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
ScopedSoftFloatState State {FCW, Frame};
auto Context = static_cast<Context::ContextImpl*>(Frame->Thread->CTX);
auto ReducedPrecisionMode = Context->Config.x87ReducedPrecision;
auto StrictReducedPrecisionMode = Context->Config.x87StrictReducedPrecision;
if (!ReducedPrecisionMode || StrictReducedPrecisionMode) {
return X80SoftFloat(src).ToF64_PreserveNan(&State.State);
}
return X80SoftFloat(src).ToF64(&State.State);
}
};
@@ -340,7 +354,7 @@ struct OpHandlers<IR::OP_F80SCALE> {
template<>
struct OpHandlers<IR::OP_F64SIN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return sin(src);
}
@@ -348,7 +362,7 @@ struct OpHandlers<IR::OP_F64SIN> {
template<>
struct OpHandlers<IR::OP_F64COS> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return cos(src);
}
@@ -356,7 +370,7 @@ struct OpHandlers<IR::OP_F64COS> {
template<>
struct OpHandlers<IR::OP_F64SINCOS> {
FEXCORE_PRESERVE_ALL_ATTR static VectorScalarF64Pair handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static VectorScalarF64Pair handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
double sin, cos;
#ifdef _WIN32
@@ -371,7 +385,7 @@ struct OpHandlers<IR::OP_F64SINCOS> {
template<>
struct OpHandlers<IR::OP_F64TAN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return tan(src);
}
@@ -379,7 +393,7 @@ struct OpHandlers<IR::OP_F64TAN> {
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return exp2(src) - 1.0;
}
@@ -387,7 +401,7 @@ struct OpHandlers<IR::OP_F64F2XM1> {
template<>
struct OpHandlers<IR::OP_F64ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return atan2(src1, src2);
}
@@ -395,7 +409,7 @@ struct OpHandlers<IR::OP_F64ATAN> {
template<>
struct OpHandlers<IR::OP_F64FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return fmod(src1, src2);
}
@@ -403,7 +417,7 @@ struct OpHandlers<IR::OP_F64FPREM> {
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return remainder(src1, src2);
}
@@ -411,7 +425,7 @@ struct OpHandlers<IR::OP_F64FPREM1> {
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
return src2 * log2(src1);
}
@@ -419,7 +433,7 @@ struct OpHandlers<IR::OP_F64FYL2X> {
template<>
struct OpHandlers<IR::OP_F64SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static double handle(uint16_t FCW, double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PRESERVE_ALL_ATTR static double handle(double src1, double src2, FEXCore::Core::CpuStateFrame* Frame) {
FEXCORE_PROFILE_INSTANT_INCREMENT(Frame->Thread, AccumulatedFloatFallbackCount, 1);
if (src1 == 0.0) { // src1 might be +/- zero
return src1; // this will return negative or positive zero if when appropriate
@@ -445,7 +459,6 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
uint64_t Tmp = Src1.ToI64(&State.State);
X80SoftFloat Rv;
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
for (size_t i = 0; i < 9; ++i) {
if (Tmp == 0) {
@@ -82,24 +82,22 @@ void InterpreterOps::FillFallbackIndexPointers(Core::FallbackABIInfo* Info, uint
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80SCALE>::handle)};
// Double Precision Unary
Info[Core::OPINDEX_F64SIN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle)};
Info[Core::OPINDEX_F64COS] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle)};
Info[Core::OPINDEX_F64SINCOS] = {ABIHandlers[FABI_F64x2_I16_F64_PTR],
Info[Core::OPINDEX_F64SIN] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SIN>::handle)};
Info[Core::OPINDEX_F64COS] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64COS>::handle)};
Info[Core::OPINDEX_F64SINCOS] = {ABIHandlers[FABI_F64x2_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SINCOS>::handle)};
Info[Core::OPINDEX_F64TAN] = {ABIHandlers[FABI_F64_I16_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle)};
Info[Core::OPINDEX_F64F2XM1] = {ABIHandlers[FABI_F64_I16_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle)};
Info[Core::OPINDEX_F64TAN] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64TAN>::handle)};
Info[Core::OPINDEX_F64F2XM1] = {ABIHandlers[FABI_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64F2XM1>::handle)};
// Double Precision Binary
Info[Core::OPINDEX_F64ATAN] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64ATAN>::handle)};
Info[Core::OPINDEX_F64FPREM] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64ATAN] = {ABIHandlers[FABI_F64_F64_F64_PTR], reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64ATAN>::handle)};
Info[Core::OPINDEX_F64FPREM] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM>::handle)};
Info[Core::OPINDEX_F64FPREM1] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64FPREM1] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FPREM1>::handle)};
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64FYL2X] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64FYL2X>::handle)};
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_I16_F64_F64_PTR],
Info[Core::OPINDEX_F64SCALE] = {ABIHandlers[FABI_F64_F64_F64_PTR],
reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F64SCALE>::handle)};
// SSE4.2 string instructions
@@ -220,21 +218,21 @@ bool InterpreterOps::GetFallbackHandler(const IR::IROp_Header* IROp, FallbackInf
return true; \
}
#define COMMON_UNARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
#define COMMON_UNARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_UNARYPAIR_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64x2_I16_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
#define COMMON_UNARYPAIR_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64x2_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
#define COMMON_BINARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_I16_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
#define COMMON_BINARY_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = {FABI_F64_F64_F64_PTR, Core::OPINDEX_F64##OP}; \
return true; \
}
// Unary
@@ -19,8 +19,8 @@ enum FallbackABI {
FABI_F80_I16_I32_PTR,
FABI_F32_I16_F80_PTR,
FABI_F64_I16_F80_PTR,
FABI_F64_I16_F64_PTR,
FABI_F64_I16_F64_F64_PTR,
FABI_F64_F64_PTR,
FABI_F64_F64_F64_PTR,
FABI_I16_I16_F80_PTR,
FABI_I32_I16_F80_PTR,
FABI_I64_I16_F80_PTR,
@@ -28,7 +28,7 @@ enum FallbackABI {
FABI_F80_I16_F80_PTR,
FABI_F80_I16_F80_F80_PTR,
FABI_F80x2_I16_F80_PTR,
FABI_F64x2_I16_F64_PTR,
FABI_F64x2_F64_PTR,
FABI_I32_I64_I64_V128_V128_I16,
FABI_I32_V128_V128_I16,
FABI_UNKNOWN,
+44 -20
View File
@@ -372,7 +372,7 @@ DEF_OP(CondSubNZCV) {
DEF_OP(Neg) {
auto Op = IROp->C<IR::IROp_Neg>();
if (Op->Cond == FEXCore::IR::COND_AL) {
if (Op->Cond == IR::CondClass::AL) {
neg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src));
} else {
cneg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src), MapCC(Op->Cond));
@@ -515,6 +515,12 @@ DEF_OP(AndWithFlags) {
}
}
DEF_OP(AndShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
and_(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1), GetReg(Op->Src2), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
DEF_OP(XorShift) {
auto Op = IROp->C<IR::IROp_XorShift>();
@@ -1040,24 +1046,19 @@ DEF_OP(Popcount) {
if (CTX->HostFeatures.SupportsCSSC) {
switch (OpSize) {
case IR::OpSize::i8Bit:
uxtb(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i16Bit:
uxth(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i32Bit:
cnt(ARMEmitter::Size::i32Bit, Dst, Src);
break;
case IR::OpSize::i64Bit:
cnt(ARMEmitter::Size::i64Bit, Dst, Src);
break;
default: LOGMAN_MSG_A_FMT("Unsupported Popcount size: {}", OpSize);
case IR::OpSize::i8Bit:
uxtb(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i16Bit:
uxth(ARMEmitter::Size::i32Bit, Dst, Src);
cnt(ARMEmitter::Size::i32Bit, Dst, Dst);
break;
case IR::OpSize::i32Bit: cnt(ARMEmitter::Size::i32Bit, Dst, Src); break;
case IR::OpSize::i64Bit: cnt(ARMEmitter::Size::i64Bit, Dst, Src); break;
default: LOGMAN_MSG_A_FMT("Unsupported Popcount size: {}", OpSize);
}
}
else {
} else {
switch (OpSize) {
case IR::OpSize::i8Bit:
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
@@ -1189,6 +1190,19 @@ DEF_OP(Rev) {
}
}
DEF_OP(Rbit) {
auto Op = IROp->C<IR::IROp_Rbit>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_A_FMT(OpSize == IR::OpSize::i32Bit || OpSize == IR::OpSize::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = ConvertSize48(IROp);
const auto Dst = GetReg(Node);
const auto Src = GetReg(Op->Src);
rbit(EmitSize, Dst, Src);
}
DEF_OP(Bfi) {
auto Op = IROp->C<IR::IROp_Bfi>();
const auto EmitSize = ConvertSize(IROp);
@@ -1272,6 +1286,16 @@ DEF_OP(Sbfe) {
sbfx(ConvertSize(IROp), Dst, Src, Op->lsb, Op->Width);
}
DEF_OP(MaskGenerateFromBitWidth) {
auto Op = IROp->C<IR::IROp_MaskGenerateFromBitWidth>();
auto BitWidth = GetReg(Op->BitWidth);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, -1);
cmp(ARMEmitter::Size::i64Bit, BitWidth, 0);
lslv(ARMEmitter::Size::i64Bit, TMP2, TMP1, BitWidth);
csinv(ARMEmitter::Size::i64Bit, GetReg(Node), TMP1, TMP2, ARMEmitter::Condition::CC_EQ);
}
DEF_OP(Select) {
auto Op = IROp->C<IR::IROp_Select>();
const auto OpSize = IROp->Size;
@@ -1368,12 +1392,12 @@ DEF_OP(VExtractToGPR) {
const auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto OpSize = IROp->Size;
[[maybe_unused]] constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto SSERegBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto ElementSizeBits = IR::OpSizeAsBits(Op->Header.ElementSize);
const auto Offset = ElementSizeBits * Op->Index;
[[maybe_unused]] const auto Is256Bit = Offset >= SSERegBitSize;
const auto Is256Bit = Offset >= SSERegBitSize;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
const auto Dst = GetReg(Node);
@@ -33,7 +33,7 @@ void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Sum));
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Pointer, false);
Relocations.emplace_back(MoveABI);
}
@@ -77,39 +77,36 @@ void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constan
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, EmitterCTX->Config.CacheObjectCodeCompilation());
LoadConstant(ARMEmitter::Size::i64Bit, Reg, Constant, false);
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations,
const char* EntryRelocations) {
size_t DataIndex {};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation* Reloc = reinterpret_cast<const FEXCore::CPU::Relocation*>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation> Relocations) {
const auto OrigBase = GetBufferBase();
const auto OrigSize = GetBufferSize();
const auto OrigOffset = GetCursorOffset();
switch (Reloc->Header.Type) {
SetBuffer(reinterpret_cast<std::uint8_t*>(Code.data()), Code.size_bytes());
for (auto& Reloc : Relocations) {
switch (Reloc.Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
uint64_t Pointer = GetNamedSymbolLiteral(Reloc.NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
SetCursorOffset(Reloc.NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc.NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
SetCursorOffset(Reloc.NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.NamedThunkMove.RegisterIndex), Pointer, true);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
@@ -117,18 +114,27 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
SetCursorOffset(Reloc.GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc.GuestRIPMove.RegisterIndex), Pointer, true);
break;
}
}
}
SetBuffer(OrigBase, OrigSize);
SetCursorOffset(OrigOffset);
return true;
}
fextl::vector<FEXCore::CPU::Relocation> Arm64JITCore::TakeRelocations() {
return std::move(Relocations);
}
} // namespace FEXCore::CPU
@@ -322,13 +322,26 @@ DEF_OP(AtomicFetchNeg) {
auto MemSrc = GetReg(Op->Addr);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
if (CTX->HostFeatures.SupportsAtomics) {
// Use a CAS loop to avoid needing to emulate unaligned LLSC atomics
ldr(SubEmitSize, TMP2, MemSrc);
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
mov(EmitSize, TMP4, TMP2);
neg(EmitSize, TMP3, TMP2);
casal(SubEmitSize, TMP2, TMP3, MemSrc);
sub(EmitSize, TMP3, TMP2, TMP4);
cbnz(EmitSize, TMP3, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
neg(EmitSize, TMP3, TMP2);
stlxr(SubEmitSize, TMP4, TMP3, MemSrc);
cbnz(EmitSize, TMP4, &LoopTop);
mov(EmitSize, GetReg(Node), TMP2.R());
}
}
DEF_OP(TelemetrySetValue) {
+58 -48
View File
@@ -146,6 +146,15 @@ DEF_OP(ExitFunction) {
} else {
stp<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::zr, ARMEmitter::XReg::zr, REG_CALLRET_SP, -0x10);
}
} else if (Op->Hint == IR::BranchHint::CheckTF) {
ARMEmitter::ForwardLabel TFUnset;
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &TFUnset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, NewRIP);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, State.rip));
ldr(TMP2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
blr(TMP2);
Bind(&TFUnset);
}
EmitLinkedBranch(NewRIP, Op->Hint == IR::BranchHint::Call);
@@ -205,21 +214,20 @@ DEF_OP(ExitFunction) {
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto Target = Op->TargetBlock;
PendingTargetLabel = &JumpTargets.try_emplace(Target.ID()).first->second;
PendingTargetLabel = JumpTarget(Op->TargetBlock);
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
auto TrueTargetLabel = JumpTarget(Op->TrueBlock);
if (Op->FromNZCV) {
b(MapCC(Op->Cond), TrueTargetLabel);
} else {
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
uint64_t Const;
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
auto Reg = GetReg(Op->Cmp1);
const auto Size = Op->CompareSize == IR::OpSize::i32Bit ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
@@ -227,16 +235,16 @@ DEF_OP(CondJump) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(isConst, "CondJump: Expected constant source");
if (Op->Cond.Val == FEXCore::IR::COND_EQ) {
if (Op->Cond == IR::CondClass::EQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbz(Size, Reg, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_NEQ) {
} else if (Op->Cond == IR::CondClass::NEQ) {
LOGMAN_THROW_A_FMT(Const == 0, "CondJump: Expected 0 source");
cbnz(Size, Reg, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_TSTZ) {
} else if (Op->Cond == IR::CondClass::TSTZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbz(Reg, Const, TrueTargetLabel);
} else if (Op->Cond.Val == FEXCore::IR::COND_TSTNZ) {
} else if (Op->Cond == IR::CondClass::TSTNZ) {
LOGMAN_THROW_A_FMT(Const < 64, "CondJump: Expected valid bit source");
tbnz(Reg, Const, TrueTargetLabel);
} else {
@@ -244,7 +252,7 @@ DEF_OP(CondJump) {
}
}
PendingTargetLabel = &JumpTargets.try_emplace(Op->FalseBlock.ID()).first->second;
PendingTargetLabel = JumpTarget(Op->FalseBlock);
}
DEF_OP(Syscall) {
@@ -438,48 +446,50 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
const auto* OldCode = (const uint8_t*)&Op->CodeOriginalLow;
auto OldCode = Op->CodeOriginal.data();
auto Base = GetReg(Op->Header.Args[0]).X();
int len = Op->CodeLength;
int idx = 0;
LoadConstant(ARMEmitter::Size::i64Bit, GetReg(Node), 0);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Entry + Op->Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, 1);
int Offset = 0;
ARMEmitter::ForwardLabel Fail;
const auto Dst = GetReg(Node);
while (len >= 8) {
ldr(ARMEmitter::XReg::x2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint64_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 8;
idx += 8;
}
while (len >= 4) {
ldr(ARMEmitter::WReg::w2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 4;
idx += 4;
}
while (len >= 2) {
ldrh(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint16_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 2;
idx += 2;
}
while (len >= 1) {
ldrb(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint8_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 1;
idx += 1;
}
auto EmitCheck = [&](size_t Size, auto&& LoadData) {
while (len >= Size) {
LoadData();
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &Fail);
len -= Size;
Offset += Size;
}
};
EmitCheck(8, [&]() {
ldr(TMP1, Base, Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, *(const uint64_t*)(OldCode + Offset));
});
EmitCheck(4, [&]() {
ldr(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint32_t*)(OldCode + Offset));
});
EmitCheck(2, [&]() {
ldrh(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint16_t*)(OldCode + Offset));
});
EmitCheck(1, [&]() {
ldrb(TMP1.W(), Base, Offset);
LoadConstant(ARMEmitter::Size::i32Bit, TMP2, *(const uint8_t*)(OldCode + Offset));
});
ARMEmitter::ForwardLabel End;
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 0);
b(&End);
Bind(&Fail);
LoadConstant(ARMEmitter::Size::i32Bit, Dst, 1);
Bind(&End);
}
DEF_OP(ThreadRemoveCodeEntry) {
@@ -423,11 +423,11 @@ DEF_OP(Vector_FToI) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Nearest: frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::NegInfinity: frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::PosInfinity: frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::TowardsZero: frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Host: frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
}
} else {
const auto IsScalar = ElementSize == OpSize;
@@ -449,21 +449,21 @@ DEF_OP(Vector_FToI) {
}
switch (Op->Round) {
case IR::Round_Nearest.Val: ROUNDING_FN(frintn); break;
case IR::Round_Negative_Infinity.Val: ROUNDING_FN(frintm); break;
case IR::Round_Positive_Infinity.Val: ROUNDING_FN(frintp); break;
case IR::Round_Towards_Zero.Val: ROUNDING_FN(frintz); break;
case IR::Round_Host.Val: ROUNDING_FN(frinti); break;
case IR::RoundMode::Nearest: ROUNDING_FN(frintn); break;
case IR::RoundMode::NegInfinity: ROUNDING_FN(frintm); break;
case IR::RoundMode::PosInfinity: ROUNDING_FN(frintp); break;
case IR::RoundMode::TowardsZero: ROUNDING_FN(frintz); break;
case IR::RoundMode::Host: ROUNDING_FN(frinti); break;
}
#undef ROUNDING_FN
} else {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Nearest: frintn(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::NegInfinity: frintm(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::PosInfinity: frintp(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: frintz(SubEmitSize, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Host: frinti(SubEmitSize, Dst.Q(), Vector.Q()); break;
}
}
}
@@ -539,11 +539,11 @@ DEF_OP(Vector_F64ToI32) {
// Then convert to integers using fcvtzs.
auto CVTReg = Dst.Z();
switch (Round) {
case IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Towards_Zero.Val: CVTReg = Vector.Z(); break;
case IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::Nearest: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::NegInfinity: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::PosInfinity: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::RoundMode::TowardsZero: CVTReg = Vector.Z(); break;
case IR::RoundMode::Host: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
}
fcvtzs(Dst.Z(), ARMEmitter::SubRegSize::i32Bit, Mask, CVTReg, ARMEmitter::SubRegSize::i64Bit);
@@ -567,11 +567,11 @@ DEF_OP(Vector_F64ToI32) {
///< Round float to integral depending on rounding mode.
switch (Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Nearest: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::NegInfinity: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::PosInfinity: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::TowardsZero: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case IR::RoundMode::Host: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
}
// Now narrow from f64 to f32.
@@ -0,0 +1,35 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/AllocatorHooks.h>
#include <FEXCore/fextl/vector.h>
#include <cstdint>
namespace FEXCore::CPU {
union Relocation;
} // namespace FEXCore::CPU
namespace FEXCore::Core {
struct DebugDataSubblock {
uint32_t HostCodeOffset;
uint32_t HostCodeSize;
};
struct DebugDataGuestOpcode {
uint64_t GuestEntryOffset;
ptrdiff_t HostEntryOffset;
};
/**
* @brief Contains debug data for a block of code for later debugger analysis
*
* Needs to remain around for as long as the code could be executed at least
*/
struct DebugData : public FEXCore::Allocator::FEXAllocOperators {
uint64_t HostCodeSize; ///< The size of the code generated in the host JIT
fextl::vector<DebugDataSubblock> Subblocks;
fextl::vector<DebugDataGuestOpcode> GuestOpcodes;
fextl::vector<FEXCore::CPU::Relocation>* Relocations;
};
} // namespace FEXCore::Core
@@ -16,7 +16,7 @@ DEF_OP(VAESImc) {
DEF_OP(VAESEnc) {
const auto Op = IROp->C<IR::IROp_VAESEnc>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -41,7 +41,7 @@ DEF_OP(VAESEnc) {
DEF_OP(VAESEncLast) {
const auto Op = IROp->C<IR::IROp_VAESEncLast>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -64,7 +64,7 @@ DEF_OP(VAESEncLast) {
DEF_OP(VAESDec) {
const auto Op = IROp->C<IR::IROp_VAESDec>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -89,7 +89,7 @@ DEF_OP(VAESDec) {
DEF_OP(VAESDecLast) {
const auto Op = IROp->C<IR::IROp_VAESDecLast>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Key = GetVReg(Op->Key);
@@ -322,7 +322,7 @@ DEF_OP(VSha256U1) {
DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
[[maybe_unused]] const auto OpSize = IROp->Size;
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1);
+66 -55
View File
@@ -12,14 +12,13 @@ $end_info$
*/
#include "Common/SoftFloat.h"
#include "FEXCore/Utils/Telemetry.h"
#include "FEXCore/Utils/TypeDefines.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include "Interface/Core/JIT/DebugData.h"
#include "Interface/Core/JIT/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
@@ -30,15 +29,16 @@ $end_info$
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXCore/Utils/TypeDefines.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
#include <stdio.h>
#include <unistd.h>
#include <string.h>
#include <cstdio>
#include <cstring>
#include <limits>
#include <unistd.h>
namespace {
struct DivRem {
@@ -222,7 +222,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
FillF64Result();
} break;
case FABI_F64_I16_F64_PTR: {
case FABI_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -239,7 +239,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
FillF64Result();
} break;
case FABI_F64x2_I16_F64_PTR: {
case FABI_F64x2_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -264,7 +264,7 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::Ref Node) {
FillF64x2Result(DstLo, DstHi);
} break;
case FABI_F64_I16_F64_F64_PTR: {
case FABI_F64_F64_F64_PTR: {
// Linux Reg/Win32 Reg:
// tmp4 (x4/x13): FallbackHandler
// x30: return
@@ -538,7 +538,8 @@ uint64_t Arm64JITCore::ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEX
} else {
{
// Guard the LookupCache lock with the code invalidation mutex, to avoid issues with forking
auto lk_inval = GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
auto lk_inval =
GuardSignalDeferringSection<std::shared_lock>(static_cast<Context::ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
HostCode = Thread->LookupCache->FindBlock(GuestRip);
}
if (!HostCode) {
@@ -624,10 +625,10 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AddRegisters(FEXCore::IR::GPRClass, GeneralRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, StaticRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRClass, GeneralFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, StaticFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::GPR, GeneralRegisters.size());
RAPass->AddRegisters(IR::RegClass::GPRFixed, StaticRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPR, GeneralFPRegisters.size());
RAPass->AddRegisters(IR::RegClass::FPRFixed, StaticFPRegisters.size());
RAPass->PairRegs = PairRegisters;
{
@@ -639,6 +640,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::In
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadRemoveCodeEntryFromJit);
Common.MonoBackpatcherWrite = reinterpret_cast<uint64_t>(&Context::ContextImpl::MonoBackpatcherWrite);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -739,48 +741,48 @@ bool Arm64JITCore::IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode,
}
}
void Arm64JITCore::EmitInterruptChecks(bool CheckTF) {
if (CheckTF) {
ARMEmitter::ForwardLabel l_TFUnset;
ARMEmitter::ForwardLabel l_TFBlocked;
void Arm64JITCore::EmitTFCheck() {
ARMEmitter::ForwardLabel l_TFUnset;
ARMEmitter::ForwardLabel l_TFBlocked;
// Note that this needs to be before the below suspend checks, as X86 checks this flag immediately after executing an instruction.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
// Note that this needs to be before the below suspend checks, as X86 checks this flag immediately after executing an instruction.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
cbz(ARMEmitter::Size::i32Bit, TMP1, &l_TFUnset);
cbz(ARMEmitter::Size::i32Bit, TMP1, &l_TFUnset);
// X86 semantically checks TF after executing each instruction, so e.g. setting a context with TF set will execute a single instruction
// and then raise an exception. However on the FEX side this is simpler to implement by checking at the start of each instruction, handle this by having bit 1 being unset in the flag state indicate that TF is blocked for a single instruction.
tbz(TMP1, 1, &l_TFBlocked);
// X86 semantically checks TF after executing each instruction, so e.g. setting a context with TF set will execute a single instruction
// and then raise an exception. However on the FEX side this is simpler to implement by checking at the start of each instruction, handle this by having bit 1 being unset in the flag state indicate that TF is blocked for a single instruction.
tbz(TMP1, 1, &l_TFBlocked);
// Block TF for a single instruction when the frontend jumps to a new context by unsetting bit 1.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
and_(ARMEmitter::Size::i32Bit, TMP1, TMP1, ~(1 << 1));
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
// Block TF for a single instruction when the frontend jumps to a new context by unsetting bit 1.
ldrb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
and_(ARMEmitter::Size::i32Bit, TMP1, TMP1, ~(1 << 1));
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Core::FAULT_SIGTRAP,
.TrapNo = X86State::X86_TRAPNO_DB,
.si_code = 2,
.err_code = 0,
};
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Core::FAULT_SIGTRAP,
.TrapNo = X86State::X86_TRAPNO_DB,
.si_code = 2,
.err_code = 0,
};
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Constant);
str(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP));
br(TMP1);
Bind(&l_TFBlocked);
// If TF was blocked for this instruction, unblock it for the next.
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 0b11);
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Bind(&l_TFUnset);
}
Bind(&l_TFBlocked);
// If TF was blocked for this instruction, unblock it for the next.
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 0b11);
strb(TMP1, STATE_PTR(CpuStateFrame, State.flags[X86State::RFLAG_TF_RAW_LOC]));
Bind(&l_TFUnset);
}
void Arm64JITCore::EmitSuspendInterruptCheck() {
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
@@ -805,7 +807,9 @@ void Arm64JITCore::EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool C
adr(TMP1, &HeaderLabel);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader));
EmitInterruptChecks(CheckTF);
if (CheckTF) {
EmitTFCheck();
}
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
@@ -827,6 +831,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
CallReturnTargets.clear();
PendingJumpThunks.clear();
uint32_t SSACount = IR->GetSSACount();
JumpTargets.resize(IR->GetHeader()->BlockCount, {});
this->Entry = Entry;
this->DebugData = DebugData;
@@ -886,10 +891,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
auto BlockStartHostCode = GetCursorAddress<uint8_t*>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
const auto Target = &JumpTargets[BlockIROp->ID];
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
if (PendingTargetLabel && PendingTargetLabel != Target) {
if (PendingTargetLabel->Backward.Location) {
EmitSuspendInterruptCheck();
}
b(PendingTargetLabel);
PendingTargetLabel = nullptr;
}
@@ -900,7 +908,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
const auto IsReturnTarget = CallReturnTargets.try_emplace(Node).first;
if (PendingTargetLabel) {
// If there is a fallthrough branch to this block, skip over the entrypoint code.
b(&IsTarget->second);
b(Target);
} else if (PendingCallReturnTargetLabel && PendingCallReturnTargetLabel != &IsReturnTarget->second) {
// If we just emitted a call, but the block we're now emitting is not the return block so don't fallthrough.
b(PendingCallReturnTargetLabel);
@@ -921,7 +929,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
}
PendingTargetLabel = nullptr;
Bind(&IsTarget->second);
Bind(Target);
}
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
@@ -945,6 +953,9 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, uint64_t Size
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel) {
if (PendingTargetLabel->Backward.Location) {
EmitSuspendInterruptCheck();
}
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
+84 -57
View File
@@ -10,13 +10,17 @@ $end_info$
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Relocations.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IntrusiveIRList.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
@@ -24,16 +28,20 @@ $end_info$
#include <array>
#include <cstdint>
#include <functional>
#include <optional>
#include <utility>
#include <variant>
namespace FEXCore::Core {
struct InternalThreadState;
}
namespace FEXCore::Context {
struct ExitFunctionLinkData;
}
namespace FEXCore::IR {
class RegisterAllocationPass;
}
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
@@ -67,7 +75,13 @@ private:
uint64_t Entry {};
CPUBackend::CompiledCode CodeData {};
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
fextl::vector<ARMEmitter::BiDirectionalLabel> JumpTargets;
ARMEmitter::BiDirectionalLabel* JumpTarget(IR::OrderedNodeWrapper Node) {
auto Block = IR->GetOp<IR::IROp_CodeBlock>(Node);
return &JumpTargets[Block->ID];
}
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> CallReturnTargets;
struct PendingJumpThunk {
@@ -83,11 +97,13 @@ private:
[[nodiscard]]
ARMEmitter::Register GetReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::GPRFixedClass.Val) {
LOGMAN_THROW_A_FMT(RegClass == IR::RegClass::GPRFixed || RegClass == IR::RegClass::GPR, "Unexpected Class: {}", Reg.Class);
if (RegClass == IR::RegClass::GPRFixed) {
return StaticRegisters[Reg.Reg];
} else if (Reg.Class == IR::GPRClass.Val) {
} else if (RegClass == IR::RegClass::GPR) {
return GeneralRegisters[Reg.Reg];
}
@@ -106,11 +122,13 @@ private:
[[nodiscard]]
ARMEmitter::VRegister GetVReg(IR::PhysicalRegister Reg) const {
LOGMAN_THROW_A_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::FPRFixedClass.Val) {
LOGMAN_THROW_A_FMT(RegClass == IR::RegClass::FPRFixed || RegClass == IR::RegClass::FPR, "Unexpected Class: {}", Reg.Class);
if (RegClass == IR::RegClass::FPRFixed) {
return StaticFPRegisters[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
} else if (RegClass == IR::RegClass::FPR) {
return GeneralFPRegisters[Reg.Reg];
}
@@ -128,8 +146,8 @@ private:
}
[[nodiscard]]
FEXCore::IR::RegisterClassType GetRegClass(IR::Ref Node) const {
return FEXCore::IR::RegisterClassType {IR::PhysicalRegister(Node).Class};
static IR::RegClass GetRegClass(IR::Ref Node) {
return IR::PhysicalRegister(Node).AsRegClass();
}
[[nodiscard]]
@@ -146,7 +164,7 @@ private:
// Converts IR-base shift type to ARMEmitter shift type.
// Will be a no-op, only a type conversion since the two definitions match.
[[nodiscard]]
ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) const {
static ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) {
return Shift == IR::ShiftType::LSL ? ARMEmitter::ShiftType::LSL :
Shift == IR::ShiftType::LSR ? ARMEmitter::ShiftType::LSR :
Shift == IR::ShiftType::ASR ? ARMEmitter::ShiftType::ASR :
@@ -154,18 +172,23 @@ private:
}
[[nodiscard]]
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
static ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
return Op->Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
static ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->Size == IR::OpSize::i32Bit || Op->Size == IR::OpSize::i64Bit, "Invalid size");
return ConvertSize(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(IR::OpSize ElementSize) {
static ARMEmitter::Size ConvertSize(IR::OpSize Size) {
return Size == IR::OpSize::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
static ARMEmitter::SubRegSize ConvertSubRegSize16(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize == IR::OpSize::i8Bit || ElementSize == IR::OpSize::i16Bit || ElementSize == IR::OpSize::i32Bit ||
ElementSize == IR::OpSize::i64Bit || ElementSize == IR::OpSize::i128Bit,
"Invalid size");
@@ -177,105 +200,105 @@ private:
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
return ConvertSubRegSize16(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(IR::OpSize ElementSize) {
static ARMEmitter::SubRegSize ConvertSubRegSize8(IR::OpSize ElementSize) {
LOGMAN_THROW_A_FMT(ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSize16(ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
return ConvertSubRegSize8(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i64Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
static ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
return ARMEmitter::ToVectorSizePair(ConvertSubRegSize16(Op));
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i128Bit, "Invalid size");
return ConvertSubRegSizePair16(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
static ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
LOGMAN_THROW_A_FMT(Op->ElementSize != IR::OpSize::i8Bit, "Invalid size");
return ConvertSubRegSizePair8(Op);
}
[[nodiscard]]
ARMEmitter::Condition MapCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
static ARMEmitter::Condition MapCC(IR::CondClass Cond) {
switch (Cond) {
case IR::CondClass::EQ: return ARMEmitter::Condition::CC_EQ;
case IR::CondClass::NEQ: return ARMEmitter::Condition::CC_NE;
case IR::CondClass::SGE: return ARMEmitter::Condition::CC_GE;
case IR::CondClass::SLT: return ARMEmitter::Condition::CC_LT;
case IR::CondClass::SGT: return ARMEmitter::Condition::CC_GT;
case IR::CondClass::SLE: return ARMEmitter::Condition::CC_LE;
case IR::CondClass::UGE: return ARMEmitter::Condition::CC_CS;
case IR::CondClass::ULT: return ARMEmitter::Condition::CC_CC;
case IR::CondClass::UGT: return ARMEmitter::Condition::CC_HI;
case IR::CondClass::ULE: return ARMEmitter::Condition::CC_LS;
case IR::CondClass::FLU: return ARMEmitter::Condition::CC_LT;
case IR::CondClass::FGE: return ARMEmitter::Condition::CC_GE;
case IR::CondClass::FLEU: return ARMEmitter::Condition::CC_LE;
case IR::CondClass::FGT: return ARMEmitter::Condition::CC_GT;
case IR::CondClass::FU:
case IR::CondClass::VS: return ARMEmitter::Condition::CC_VS;
case IR::CondClass::FNU:
case IR::CondClass::VC: return ARMEmitter::Condition::CC_VC;
case IR::CondClass::MI: return ARMEmitter::Condition::CC_MI;
case IR::CondClass::PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
}
}
[[nodiscard]]
bool IsFPR(IR::RegisterClassType Class) const {
return Class == IR::FPRClass || Class == IR::FPRFixedClass;
static bool IsFPR(IR::RegClass Class) {
return Class == IR::RegClass::FPR || Class == IR::RegClass::FPRFixed;
}
[[nodiscard]]
bool IsGPR(IR::RegisterClassType Class) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
static bool IsGPR(IR::RegClass Class) {
return Class == IR::RegClass::GPR || Class == IR::RegClass::GPRFixed;
}
[[nodiscard]]
bool IsGPR(IR::Ref Node) {
static bool IsGPR(IR::Ref Node) {
return IsGPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsFPR(IR::Ref Node) {
static bool IsFPR(IR::Ref Node) {
return IsFPR(GetRegClass(Node));
}
[[nodiscard]]
bool IsGPR(IR::OrderedNodeWrapper Wrap) {
return IsGPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
static bool IsGPR(IR::OrderedNodeWrapper Wrap) {
return IsGPR(IR::PhysicalRegister(Wrap).AsRegClass());
}
[[nodiscard]]
bool IsFPR(IR::OrderedNodeWrapper Wrap) {
return IsFPR(IR::RegisterClassType {IR::PhysicalRegister(Wrap).Class});
static bool IsFPR(IR::OrderedNodeWrapper Wrap) {
return IsFPR(IR::PhysicalRegister(Wrap).AsRegClass());
}
[[nodiscard]]
@@ -374,7 +397,9 @@ private:
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
bool ApplyRelocations(uint64_t GuestEntry, std::span<std::byte> Code, std::span<const FEXCore::CPU::Relocation>);
fextl::vector<FEXCore::CPU::Relocation> TakeRelocations() override;
/** @} */
@@ -397,9 +422,11 @@ private:
void Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, IR::OpSize VectorIndexSize,
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale);
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale, IR::OpSize AddrSize);
void EmitInterruptChecks(bool CheckTF);
void EmitTFCheck();
void EmitSuspendInterruptCheck();
void EmitEntryPoint(ARMEmitter::BackwardLabel& HeaderLabel, bool CheckTF);
+84 -61
View File
@@ -21,7 +21,7 @@ DEF_OP(LoadContext) {
const auto Op = IROp->C<IR::IROp_LoadContext>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Dst = GetReg(Node);
switch (OpSize) {
@@ -52,7 +52,7 @@ DEF_OP(LoadContext) {
DEF_OP(LoadContextPair) {
const auto Op = IROp->C<IR::IROp_LoadContextPair>();
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
@@ -78,7 +78,7 @@ DEF_OP(StoreContext) {
const auto Op = IROp->C<IR::IROp_StoreContext>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Src = GetZeroableReg(Op->Value);
switch (OpSize) {
@@ -110,7 +110,7 @@ DEF_OP(StoreContextPair) {
const auto Op = IROp->C<IR::IROp_StoreContextPair>();
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
auto Src1 = GetZeroableReg(Op->Value1);
auto Src2 = GetZeroableReg(Op->Value2);
@@ -135,12 +135,12 @@ DEF_OP(StoreContextPair) {
DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
if (Op->Class == IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Reg < StaticRegisters.size(), "out of range reg");
mov(GetReg(Node).X(), StaticRegisters[Op->Reg].X());
} else if (Op->Class == IR::FPRClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
} else if (Op->Class == IR::RegClass::FPR) {
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "out of range reg");
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
@@ -175,13 +175,14 @@ DEF_OP(LoadAF) {
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
auto Reg = IR::PhysicalRegister(Node);
const auto Reg = IR::PhysicalRegister(Node);
const auto RegClass = Reg.AsRegClass();
if (Reg.Class == IR::GPRFixedClass) {
if (RegClass == IR::RegClass::GPRFixed) {
// Always use 64-bit, it's faster. Upper bits ignored for 32-bit mode.
mov(ARMEmitter::Size::i64Bit, GetReg(Reg), GetReg(Op->Value));
} else if (Reg.Class == IR::FPRFixedClass) {
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
} else if (RegClass == IR::RegClass::FPRFixed) {
const auto regSize = HostSupportsAVX256 ? IR::OpSize::i256Bit : IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
const auto guest = GetVReg(Reg);
@@ -193,7 +194,7 @@ DEF_OP(StoreRegister) {
mov(guest.Q(), host.Q());
}
} else {
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", Reg.Class);
LOGMAN_THROW_A_FMT(false, "Unhandled Op->Class {}", RegClass);
}
}
@@ -225,7 +226,7 @@ DEF_OP(LoadContextIndexed) {
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
switch (Op->Stride) {
case 1:
case 2:
@@ -288,7 +289,7 @@ DEF_OP(StoreContextIndexed) {
const auto Index = GetReg(Op->Index);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Value = GetReg(Op->Value);
switch (Op->Stride) {
@@ -348,12 +349,31 @@ DEF_OP(StoreContextIndexed) {
}
}
DEF_OP(FormContextAddress) {
const auto Op = IROp->C<IR::IROp_FormContextAddress>();
const auto Index = GetReg(Op->Index);
const auto Dst = GetReg(Node);
switch (Op->Stride) {
case 1:
case 2:
case 4:
case 8:
case 16:
case 32: {
add(ARMEmitter::Size::i64Bit, Dst, STATE, Index, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(Op->Stride));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled FormContextAddress stride: {}", Op->Stride); break;
}
}
DEF_OP(SpillRegister) {
const auto Op = IROp->C<IR::IROp_SpillRegister>();
const auto OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: {
@@ -394,7 +414,7 @@ DEF_OP(SpillRegister) {
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize); break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
} else if (Op->Class == FEXCore::IR::RegClass::FPR) {
const auto Src = GetVReg(Op->Value);
switch (OpSize) {
@@ -433,7 +453,7 @@ DEF_OP(SpillRegister) {
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize); break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled SpillRegister class: {}", Op->Class.Val);
LOGMAN_MSG_A_FMT("Unhandled SpillRegister class: {}", Op->Class);
}
}
@@ -442,7 +462,7 @@ DEF_OP(FillRegister) {
const auto OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
switch (OpSize) {
case IR::OpSize::i8Bit: {
@@ -483,7 +503,7 @@ DEF_OP(FillRegister) {
}
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize); break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
} else if (Op->Class == FEXCore::IR::RegClass::FPR) {
const auto Dst = GetVReg(Node);
switch (OpSize) {
@@ -522,7 +542,7 @@ DEF_OP(FillRegister) {
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize); break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled FillRegister class: {}", Op->Class.Val);
LOGMAN_MSG_A_FMT("Unhandled FillRegister class: {}", Op->Class);
}
}
@@ -559,14 +579,14 @@ ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(
return ARMEmitter::ExtendedMemOperand(Base.X(), ARMEmitter::IndexType::OFFSET, Const);
} else {
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
switch (OffsetType) {
case IR::MemOffsetType::SXTX:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
case IR::MEM_OFFSET_UXTW.Val:
case IR::MemOffsetType::UXTW:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::UXTW, FEXCore::ilog2(OffsetScale));
case IR::MEM_OFFSET_SXTW.Val:
case IR::MemOffsetType::SXTW:
return ARMEmitter::ExtendedMemOperand(Base.X(), RegOffset.X(), ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
default: LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetType: {}", OffsetType.Val); break;
default: LOGMAN_MSG_A_FMT("Unhandled GenerateMemOperand OffsetType: {}", OffsetType); break;
}
}
}
@@ -593,20 +613,20 @@ ARMEmitter::Register Arm64JITCore::ApplyMemOperand(IR::OpSize AccessSize, ARMEmi
add(ARMEmitter::Size::i64Bit, Tmp, Base, Tmp, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
} else {
auto RegOffset = GetReg(Offset);
switch (OffsetType.Val) {
case IR::MEM_OFFSET_SXTX.Val:
switch (OffsetType) {
case IR::MemOffsetType::SXTX:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::SXTX, FEXCore::ilog2(OffsetScale));
break;
case IR::MEM_OFFSET_UXTW.Val:
case IR::MemOffsetType::UXTW:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::UXTW, FEXCore::ilog2(OffsetScale));
break;
case IR::MEM_OFFSET_SXTW.Val:
case IR::MemOffsetType::SXTW:
add(ARMEmitter::Size::i64Bit, Tmp, Base, RegOffset, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
break;
default: LOGMAN_MSG_A_FMT("Unhandled OffsetType: {}", OffsetType.Val); break;
default: LOGMAN_MSG_A_FMT("Unhandled OffsetType: {}", OffsetType); break;
}
}
return Tmp;
@@ -657,7 +677,7 @@ ARMEmitter::SVEMemOperand Arm64JITCore::GenerateSVEMemOperand(IR::OpSize AccessS
// Note that we do nothing with the offset type and offset scale,
// since SVE loads and stores don't have the ability to perform an
// optional extension or shift as part of their behavior.
LOGMAN_THROW_A_FMT(OffsetType.Val == IR::MEM_OFFSET_SXTX.Val, "Currently only the default offset type (SXTX) is supported.");
LOGMAN_THROW_A_FMT(OffsetType == IR::MemOffsetType::SXTX, "Currently only the default offset type (SXTX) is supported.");
const auto RegOffset = GetReg(Offset);
return ARMEmitter::SVEMemOperand(Base.X(), RegOffset.X());
@@ -670,7 +690,7 @@ DEF_OP(LoadMem) {
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
switch (OpSize) {
@@ -704,7 +724,7 @@ DEF_OP(LoadMemPair) {
const auto Op = IROp->C<IR::IROp_LoadMemPair>();
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Dst1 = GetReg(Op->OutValue1);
const auto Dst2 = GetReg(Op->OutValue2);
@@ -732,17 +752,17 @@ DEF_OP(LoadMemTSO) {
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
LOGMAN_THROW_A_FMT(Op->OffsetScale == 1, "unexpected offset scale");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MEM_OFFSET_SXTX, "unexpected offset type");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MemOffsetType::SXTX, "unexpected offset type");
}
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
[[maybe_unused]] bool IsInline = IsInlineConstant(Op->Offset, &Offset);
bool IsInline = IsInlineConstant(Op->Offset, &Offset);
LOGMAN_THROW_A_FMT(IsInline, "expected immediate");
}
@@ -760,7 +780,7 @@ DEF_OP(LoadMemTSO) {
// Half-barrier once back-patched.
nop();
}
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
if (OpSize == IR::OpSize::i8Bit) {
// 8bit load is always aligned to natural alignment
@@ -775,7 +795,7 @@ DEF_OP(LoadMemTSO) {
// Half-barrier once back-patched.
nop();
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
if (OpSize == IR::OpSize::i8Bit) {
// 8bit load is always aligned to natural alignment
@@ -1020,7 +1040,7 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
ARMEmitter::VRegister IncomingDst, std::optional<ARMEmitter::Register> BaseAddr,
ARMEmitter::VRegister VectorIndexLow, std::optional<ARMEmitter::VRegister> VectorIndexHigh,
ARMEmitter::VRegister MaskReg, IR::OpSize VectorIndexSize, size_t DataElementOffsetStart,
size_t IndexElementOffsetStart, uint8_t OffsetScale) {
size_t IndexElementOffsetStart, uint8_t OffsetScale, IR::OpSize AddrSize) {
LOGMAN_THROW_A_FMT(ElementSize >= IR::OpSize::i8Bit && ElementSize <= IR::OpSize::i64Bit, "Invalid element size");
const auto PerformSMove = [this](IR::OpSize ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
@@ -1096,17 +1116,17 @@ void Arm64JITCore::Emulate128BitGather(IR::OpSize Size, IR::OpSize ElementSize,
// Calculate memory position for this gather load
if (BaseAddr.has_value()) {
if (VectorIndexSize == IR::OpSize::i32Bit) {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
add(ConvertSize(AddrSize), TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
} else {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
add(ConvertSize(AddrSize), TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
}
} else {
///< In this case we have no base address, All addresses come from the vector register itself
if (VectorIndexSize == IR::OpSize::i32Bit) {
// Sign extend and shift in to the 64-bit register
sbfiz(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
sbfiz(ConvertSize(AddrSize), TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
} else {
lsl(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
lsl(ConvertSize(AddrSize), TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
}
}
@@ -1164,7 +1184,8 @@ DEF_OP(VLoadVectorGatherMasked) {
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
const bool SupportsSVELoad = (HostSupportsSVE128 || HostSupportsSVE256) &&
(OffsetScale == 1 || OffsetScale == IR::OpSizeToSize(VectorIndexSize)) && VectorIndexSize == IROp->ElementSize;
(OffsetScale == 1 || OffsetScale == IR::OpSizeToSize(VectorIndexSize)) &&
VectorIndexSize == IROp->ElementSize && Op->AddrSize == IR::OpSize::i64Bit;
if (SupportsSVELoad) {
uint8_t SVEScale = FEXCore::ilog2(OffsetScale);
@@ -1222,7 +1243,7 @@ DEF_OP(VLoadVectorGatherMasked) {
} else {
LOGMAN_THROW_A_FMT(!Is256Bit, "Can't emulate this gather load in the backend! Programming error!");
Emulate128BitGather(IROp->Size, IROp->ElementSize, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg,
VectorIndexSize, DataElementOffsetStart, IndexElementOffsetStart, OffsetScale);
VectorIndexSize, DataElementOffsetStart, IndexElementOffsetStart, OffsetScale, Op->AddrSize);
}
}
@@ -1247,7 +1268,9 @@ DEF_OP(VLoadVectorGatherMaskedQPS) {
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh)) : std::nullopt;
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
if (HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4)) {
const bool SupportsSVELoad = HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4) && Op->AddrSize == IR::OpSize::i64Bit;
if (SupportsSVELoad) {
ARMEmitter::SVEModType ModType = ARMEmitter::SVEModType::MOD_NONE;
if (OffsetScale != 1) {
ModType = ARMEmitter::SVEModType::MOD_LSL;
@@ -1301,7 +1324,7 @@ DEF_OP(VLoadVectorGatherMaskedQPS) {
}
} else {
Emulate128BitGather(IR::OpSize::i128Bit, IR::OpSize::i32Bit, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg,
IR::OpSize::i64Bit, 0, 0, OffsetScale);
IR::OpSize::i64Bit, 0, 0, OffsetScale, Op->AddrSize);
}
}
@@ -1602,7 +1625,7 @@ DEF_OP(StoreMem) {
const auto MemReg = GetReg(Op->Addr);
const auto MemSrc = GenerateMemOperand(OpSize, MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
switch (OpSize) {
case IR::OpSize::i8Bit: strb(Src, MemSrc); break;
@@ -1713,7 +1736,7 @@ DEF_OP(StoreMemPair) {
const auto OpSize = IROp->Size;
const auto Addr = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
const auto Src1 = GetZeroableReg(Op->Value1);
const auto Src2 = GetZeroableReg(Op->Value2);
switch (OpSize) {
@@ -1740,17 +1763,17 @@ DEF_OP(StoreMemTSO) {
const auto MemReg = GetReg(Op->Addr);
if (Op->Class == FEXCore::IR::GPRClass) {
if (Op->Class == IR::RegClass::GPR) {
LOGMAN_THROW_A_FMT(Op->Offset.IsInvalid() || CTX->HostFeatures.SupportsTSOImm9, "unexpected offset");
LOGMAN_THROW_A_FMT(Op->OffsetScale == 1, "unexpected offset scale");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MEM_OFFSET_SXTX, "unexpected offset type");
LOGMAN_THROW_A_FMT(Op->OffsetType == IR::MemOffsetType::SXTX, "unexpected offset type");
}
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
[[maybe_unused]] bool IsInline = IsInlineConstant(Op->Offset, &Offset);
bool IsInline = IsInlineConstant(Op->Offset, &Offset);
LOGMAN_THROW_A_FMT(IsInline, "expected immediate");
}
@@ -1767,7 +1790,7 @@ DEF_OP(StoreMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled StoreMemTSO size: {}", OpSize); break;
}
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
if (OpSize == IR::OpSize::i8Bit) {
@@ -2280,7 +2303,7 @@ DEF_OP(ParanoidLoadMemTSO) {
auto MemReg = GetReg(Op->Addr);
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
@@ -2301,7 +2324,7 @@ DEF_OP(ParanoidLoadMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", OpSize); break;
}
}
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == FEXCore::IR::GPRClass) {
} else if (CTX->HostFeatures.SupportsRCPC && Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
MemReg = ApplyMemOperand(OpSize, MemReg, TMP4, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (OpSize == IR::OpSize::i8Bit) {
@@ -2315,7 +2338,7 @@ DEF_OP(ParanoidLoadMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidLoadMemTSO size: {}", OpSize); break;
}
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Dst = GetReg(Node);
MemReg = ApplyMemOperand(OpSize, MemReg, TMP4, Op->Offset, Op->OffsetType, Op->OffsetScale);
switch (OpSize) {
@@ -2368,7 +2391,7 @@ DEF_OP(ParanoidStoreMemTSO) {
auto MemReg = GetReg(Op->Addr);
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == FEXCore::IR::GPRClass) {
if (CTX->HostFeatures.SupportsTSOImm9 && Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
uint64_t Offset = 0;
if (!Op->Offset.IsInvalid()) {
@@ -2388,7 +2411,7 @@ DEF_OP(ParanoidStoreMemTSO) {
default: LOGMAN_MSG_A_FMT("Unhandled ParanoidStoreMemTSO size: {}", OpSize); break;
}
}
} else if (Op->Class == FEXCore::IR::GPRClass) {
} else if (Op->Class == IR::RegClass::GPR) {
const auto Src = GetZeroableReg(Op->Value);
MemReg = ApplyMemOperand(OpSize, MemReg, TMP1, Op->Offset, Op->OffsetType, Op->OffsetScale);
switch (OpSize) {
@@ -2586,7 +2609,7 @@ DEF_OP(VStoreNonTemporalPair) {
const auto Op = IROp->C<IR::IROp_VStoreNonTemporalPair>();
const auto OpSize = IROp->Size;
[[maybe_unused]] const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
const auto Is128Bit = OpSize == IR::OpSize::i128Bit;
LOGMAN_THROW_A_FMT(Is128Bit, "This IR operation only operates at 128-bit wide");
const auto ValueLow = GetVReg(Op->ValueLow);
+50 -9
View File
@@ -10,10 +10,12 @@ $end_info$
#endif
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/DebugData.h"
#include "Interface/Core/JIT/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
namespace FEXCore::CPU {
@@ -46,10 +48,10 @@ DEF_OP(GuestOpcode) {
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::Fence_LoadStore.Val: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::Fence_Store.Val: dmb(ARMEmitter::BarrierScope::ST); break;
case IR::Fence_Inst.Val: isb(); break;
case IR::FenceType::Load: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::FenceType::LoadStore: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::FenceType::Store: dmb(ARMEmitter::BarrierScope::ST); break;
case IR::FenceType::Inst: isb(); break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
}
@@ -106,10 +108,10 @@ DEF_OP(GetRoundingMode) {
// zero. Just swapping 01 and 10. That's a bitfield reverse. Round mode is in
// bottom two bits. After reversing as a 32-bit operation, it'll be in [31:30]
// and ripe for reinsertion back at 0.
static_assert(IR::ROUND_MODE_NEAREST == 0);
static_assert(IR::ROUND_MODE_NEGATIVE_INFINITY == 1);
static_assert(IR::ROUND_MODE_POSITIVE_INFINITY == 2);
static_assert(IR::ROUND_MODE_TOWARDS_ZERO == 3);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::Nearest) == 0);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::NegInfinity) == 1);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::PosInfinity) == 2);
static_assert(FEXCore::ToUnderlying(IR::RoundMode::TowardsZero) == 3);
rbit(ARMEmitter::Size::i32Bit, TMP1, Dst);
bfi(ARMEmitter::Size::i64Bit, Dst, TMP1, 30, 2);
@@ -286,4 +288,43 @@ DEF_OP(Yield) {
yield();
}
DEF_OP(MonoBackpatcherWrite) {
auto Op = IROp->C<IR::IROp_MonoBackpatcherWrite>();
mov(ARMEmitter::Size::i64Bit, TMP3, GetReg(Op->Addr));
mov(ARMEmitter::Size::i64Bit, TMP4, GetReg(Op->Value));
PushDynamicRegs(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, IR::OpSizeToSize(Op->Size));
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, TMP3);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, TMP4);
}
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, TMP1, 1);
strb(TMP1.W(), TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
ldr(ARMEmitter::XReg::x4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.MonoBackpatcherWrite));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, uint8_t, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4);
}
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
FillStaticRegs();
PopDynamicRegs();
}
} // namespace FEXCore::CPU
@@ -18,18 +18,4 @@ DEF_OP(RMWHandle) {
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(IROp->Args[0]));
}
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A), B = GetReg(Op->B);
LOGMAN_THROW_A_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
mov(ARMEmitter::Size::i64Bit, A, B);
mov(ARMEmitter::Size::i64Bit, B, TMP1);
}
DEF_OP(Swap2) {
// Implemented above
}
} // namespace FEXCore::CPU
@@ -41,6 +41,7 @@ namespace FEXCore::CPU {
const auto Op = IROp->C<IR::IROp_##FEXOp>(); \
const auto OpSize = IROp->Size; \
const auto Is256Bit = OpSize == IR::OpSize::i256Bit; \
const auto Is128Bit = OpSize == IR::OpSize::i128Bit; \
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__); \
\
const auto Dst = GetVReg(Node); \
@@ -49,8 +50,10 @@ namespace FEXCore::CPU {
\
if (HostSupportsSVE256 && Is256Bit) { \
ARMOp(Dst.Z(), Vector1.Z(), Vector2.Z()); \
} else { \
} else if (Is128Bit) { \
ARMOp(Dst.Q(), Vector1.Q(), Vector2.Q()); \
} else { \
ARMOp(Dst.D(), Vector1.D(), Vector2.D()); \
} \
}
@@ -744,11 +747,11 @@ DEF_OP(VFToIScalarInsert) {
auto Src = *std::get_if<ARMEmitter::VRegister>(&SrcVar);
switch (RoundMode) {
case IR::Round_Nearest: frintn(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Negative_Infinity: frintm(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Positive_Infinity: frintp(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Towards_Zero: frintz(SubRegSize.Scalar, Dst, Src); break;
case IR::Round_Host: frinti(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::Nearest: frintn(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::NegInfinity: frintm(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::PosInfinity: frintp(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::TowardsZero: frintz(SubRegSize.Scalar, Dst, Src); break;
case IR::RoundMode::Host: frinti(SubRegSize.Scalar, Dst, Src); break;
}
};
@@ -1352,7 +1355,7 @@ DEF_OP(VFMin) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubRegSize = ConvertSubRegSize248(IROp);
[[maybe_unused]] const auto IsScalar = ElementSize == OpSize;
const auto IsScalar = ElementSize == OpSize;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -1425,7 +1428,7 @@ DEF_OP(VFMax) {
const auto ElementSize = Op->Header.ElementSize;
const auto SubRegSize = ConvertSubRegSize248(IROp);
[[maybe_unused]] const auto IsScalar = ElementSize == OpSize;
const auto IsScalar = ElementSize == OpSize;
const auto Is256Bit = OpSize == IR::OpSize::i256Bit;
LOGMAN_THROW_A_FMT(!Is256Bit || HostSupportsSVE256, "Need SVE256 support in order to use {} with 256-bit operation", __func__);
@@ -1,86 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct FEX_PACKED CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie {};
// Instructions per block configuration
int32_t MaxInstPerBlock {};
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Multiblock enabled
unsigned MultiBlock : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// TSO enabled
unsigned TSOEnabled : 1;
// ABI local flag unsafe optimization
unsigned ABILocalFlags : 1;
// Paranoid TSO mode enabled
unsigned ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
unsigned Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
unsigned x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 19;
bool operator==(const CodeObjectSerializationConfig& other) const {
return Cookie == other.Cookie && MaxInstPerBlock == other.MaxInstPerBlock && Arch == other.Arch && MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled && TSOEnabled == other.TSOEnabled && ABILocalFlags == other.ABILocalFlags &&
ParanoidTSO == other.ParanoidTSO && Is64BitMode == other.Is64BitMode && SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(const CodeObjectSerializationConfig& other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash {};
Hash <<= 32;
Hash |= other.MaxInstPerBlock;
Hash <<= 1;
Hash |= other.Arch;
Hash <<= 1;
Hash |= other.MultiBlock;
Hash <<= 1;
Hash |= other.HardwareTSOEnabled;
Hash <<= 1;
Hash |= other.TSOEnabled;
Hash <<= 1;
Hash |= other.ABILocalFlags;
Hash <<= 1;
Hash |= other.ParanoidTSO;
Hash <<= 1;
Hash |= other.Is64BitMode;
Hash <<= 2;
Hash |= other.SMCChecks;
Hash <<= 1;
Hash |= other.x87ReducedPrecision;
return Hash;
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is "
"generated!");
} // namespace FEXCore::CodeSerialize
@@ -1,121 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/Filesystem.h>
#include <fcntl.h>
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
#ifndef _WIN32
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
const fextl::string BaseFilename = FHU::Filesystem::GetFilename(filename);
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = fextl::make_unique<CodeRegionEntry>(Base, Size, Offset, filename, NamedRegionHandler->DefaultCodeHeader(Base, Offset));
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
} else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
#ifndef _WIN32
// Removing a named region through the job system
// We need to find the entry that we are deleting first
fextl::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
// Take the pointer from the map
EntryPointer = std::move(it->second);
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Remove this from the entry map
EntryMap.erase(it);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
}
} else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
} // namespace FEXCore::CodeSerialize
@@ -1,71 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/string.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename,
const fextl::string& filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion*>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion*>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
}
}
}
} // namespace FEXCore::CodeSerialize
@@ -1,85 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/Utils/Threads.h>
namespace {
static void* ThreadHandler(void* Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService* This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
} // namespace
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, AsyncHandler {&NamedRegionHandler, this}
, NamedRegionHandler {ctx} {
Initialize();
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
WorkerThreadShuttingDown = true;
// Kick the working thread
WorkAvailable.NotifyAll();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, fextl::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
const CodeObjectFileSection* CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
FEXCore::Threads::SetThreadName("ObjectCodeSeri\0");
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto& it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
} // namespace FEXCore::CodeSerialize
@@ -1,457 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Context/Context.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include "Interface/Core/ObjectCache/CodeObjectSerializationConfig.h"
#include "Interface/IR/AOTIR.h"
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/queue.h>
#include <FEXCore/fextl/robin_map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <shared_mutex>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData* Data;
const char* HostCode;
uint64_t NumRelocations;
const char* Relocations;
};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase {};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset {};
// Total amount of code that should be in this file
uint64_t TotalCodeSize {};
// Used to reserve the TSL map
uint64_t NumCodeEntries {};
// The number of relocations that point to this section
uint64_t NumRelocationsTo {};
// Total relocations in this file
uint64_t TotalRelocationsCount {};
};
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base {};
// Size of this code entry
uint64_t Size {};
// The offset inside the file that is mapped to Base
uint64_t Offset {};
// Filename of the object
fextl::string Filename {};
CodeObjectSerializationHeader EntryHeader {};
/** @} */
// The filename of the object cache for this entry
fextl::string ObjectEntrySourceFilename {};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char* CodeData {};
size_t FileSize {};
fextl::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
fextl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap {};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base, uint64_t Size, uint64_t Offset, const fextl::string& Filename, const CodeObjectSerializationHeader& DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = fextl::map<uint64_t, fextl::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = fextl::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
void* HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex* ThreadJobRefCount;
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex* ObjectJobRefCountMutexPtr;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
AsyncJobHandler(NamedRegionObjectHandler* NamedRegionHandler, CodeObjectSerializeService* CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const {
return Type;
}
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry} {}
const fextl::string BaseFilename;
const fextl::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, fextl::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
fextl::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler* NamedRegionHandler;
CodeObjectSerializeService* CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx);
void HandleNamedRegionObjectJobs();
const CodeObjectSerializationConfig& GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemAddNamedRegion>(base, filename, executable, entry));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion>(Base, Size, std::move(Entry)));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs {};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex {};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
fextl::queue<fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue {};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename, const fextl::string& filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx);
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(fextl::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex* ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
const CodeObjectFileSection* FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it);
CodeSerializationMutex& GetEntryMapMutex() {
return EntryMapMutex;
}
CodeSerializationMutex& GetUnrelocatedEntryMapMutex() {
return EntryMapMutex;
}
CodeRegionMapType& GetEntryMap() {
return AddressToEntryMap;
}
CodeRegionPtrMapType& GetUnrelocatedEntryMap() {
return UnrelocatedAddressToEntryMap;
}
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() {
WorkAvailable.NotifyOne();
}
private:
FEXCore::Context::ContextImpl* CTX;
Event WorkAvailable {};
fextl::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
};
} // namespace FEXCore::CodeSerialize
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -53,10 +53,11 @@ constexpr inline DispatchTableEntry OpDispatch_BaseOpTable[] = {
{0xAA, 2, &OpDispatchBuilder::STOSOp},
{0xAC, 2, &OpDispatchBuilder::LODSOp},
{0xAE, 2, &OpDispatchBuilder::SCASOp},
{0xB0, 16, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVGPROp, 0>},
{0xB0, 16, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVGPRImmediate>},
{0xC2, 2, &OpDispatchBuilder::RETOp},
{0xC8, 1, &OpDispatchBuilder::EnterOp},
{0xC9, 1, &OpDispatchBuilder::LEAVEOp},
{0xCA, 2, &OpDispatchBuilder::RETFARIndirectOp},
{0xCC, 2, &OpDispatchBuilder::INTOp},
{0xCF, 1, &OpDispatchBuilder::IRETOp},
{0xD7, 2, &OpDispatchBuilder::XLATOp},
@@ -75,33 +76,4 @@ constexpr inline DispatchTableEntry OpDispatch_BaseOpTable[] = {
{0xFA, 2, &OpDispatchBuilder::PermissionRestrictedOp},
{0xFC, 2, &OpDispatchBuilder::FLAGControlOp},
};
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable_64[] = {
{0x63, 1, &OpDispatchBuilder::MOVSXDOp},
{0xA0, 4, &OpDispatchBuilder::MOVOffsetOp},
};
constexpr inline DispatchTableEntry OpDispatch_BaseOpTable_32[] = {
{0x06, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x07, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX>},
{0x0E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX>},
{0x16, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX>},
{0x17, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX>},
{0x1E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX>},
{0x1F, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX>},
{0x27, 1, &OpDispatchBuilder::DAAOp},
{0x2F, 1, &OpDispatchBuilder::DASOp},
{0x37, 1, &OpDispatchBuilder::AAAOp},
{0x3F, 1, &OpDispatchBuilder::AASOp},
{0x40, 8, &OpDispatchBuilder::INCOp},
{0x48, 8, &OpDispatchBuilder::DECOp},
{0x60, 1, &OpDispatchBuilder::PUSHAOp},
{0x61, 1, &OpDispatchBuilder::POPAOp},
{0xA0, 4, &OpDispatchBuilder::MOVOffsetOp},
{0xCE, 1, &OpDispatchBuilder::INTOp},
{0xD4, 1, &OpDispatchBuilder::AAMOp},
{0xD5, 1, &OpDispatchBuilder::AADOp},
{0xD6, 1, &OpDispatchBuilder::SALCOp},
};
} // namespace FEXCore::IR
@@ -19,8 +19,12 @@ class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
@@ -32,24 +36,32 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
auto Tmp = _VAdd(OpSize::i128Bit, OpSize::i32Bit, Src, RotatedNode);
auto Result = _VInsElement(OpSize::i128Bit, OpSize::i32Bit, 3, 3, Src, Tmp);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref NewVec = _VExtr(OpSize::i128Bit, OpSize::i64Bit, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
Ref Result = _VXor(OpSize::i128Bit, OpSize::i8Bit, Dest, NewVec);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// ARM SHA1 mostly matches x86 semantics, except the input and outputs are both flipped from elements 0,1,2,3 to 3,2,1,0.
auto Src1 = SHADataShuffle(Dest);
@@ -58,13 +70,17 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
// The result is swizzled differently than expected
auto Result = SHADataShuffle(_VSha1SU1(Src1, Src2));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result {};
Ref ConstantVector {};
@@ -96,21 +112,29 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
case 3: Result = SHADataShuffle(_VSha1P(Src1, ZeroRegister, Src2)); break;
}
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
auto Result = _VSha256U0(Dest, Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
auto Src1 = _VExtr(OpSize::i128Bit, OpSize::i32Bit, Dest, Dest, 3);
auto DupDst = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
@@ -118,22 +142,16 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
auto Result = _VSha256U1(Src1, Src2);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
}
Ref OpDispatchBuilder::BitwiseAtLeastTwo(Ref A, Ref B, Ref C) {
// Returns whether at least 2/3 of A/B/C is true.
// Expressed as (A & (B | C)) | (B & C)
//
// Equivalent to expression in SHA calculations: (A & B) ^ (A & C) ^ (B & C)
auto And = _And(OpSize::i32Bit, B, C);
auto Or = _Or(OpSize::i32Bit, B, C);
return _Or(OpSize::i32Bit, _And(OpSize::i32Bit, A, Or), And);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsSHA) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
@@ -159,101 +177,121 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
auto B = _VSha256H2(EFGH, ABCD, Key);
auto Result = shuffle_abcd(A, B);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEnc(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESEncLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDec(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Result = _VAESDecLast(OpSize::i128Bit, Dest, Src, LoadZeroVector(OpSize::i128Bit));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
const auto DstSize = OpSizeFromDst(Op);
[[maybe_unused]] const auto Is128Bit = DstSize == OpSize::i128Bit;
const auto Is128Bit = DstSize == OpSize::i128Bit;
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Key = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const uint64_t RCON = Op->Src[1].Literal();
auto KeyGenSwizzle = LoadAndCacheNamedVectorConstant(OpSize::i128Bit, NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE);
@@ -261,28 +299,41 @@ Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsAES) {
UnimplementedOp(Op);
return;
}
Ref Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, OpSize::iInvalid);
StoreResultFPR(Op, Result);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
if (!CTX->HostFeatures.SupportsPMULL_128Bit) {
UnimplementedOp(Op);
return;
}
Ref Dest = LoadSourceFPR(Op, Op->Dest, Op->Flags);
Ref Src = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
auto Res = _PCLMUL(OpSize::i128Bit, Dest, Src, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
StoreResultFPR(Op, Res);
}
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
if (!CTX->HostFeatures.SupportsPMULL_128Bit) {
UnimplementedOp(Op);
return;
}
const auto DstSize = OpSizeFromDst(Op);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Src1 = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSourceFPR(Op, Op->Src[1], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Literal());
Ref Res = _PCLMUL(DstSize, Src1, Src2, Selector & 0b1'0001);
StoreResult(FPRClass, Op, Res, OpSize::iInvalid);
StoreResultFPR(Op, Res);
}
} // namespace FEXCore::IR
@@ -201,7 +201,7 @@ void OpDispatchBuilder::FixupAF() {
auto PFRaw = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
auto AFRaw = GetRFLAG(FEXCore::X86State::RFLAG_AF_RAW_LOC);
// Again 64-bit as masking is more expensive given our ConstProp design.
// Again 64-bit as masking is more expensive.
Ref XorRes = _Xor(OpSize::i64Bit, AFRaw, PFRaw);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -263,12 +263,10 @@ void OpDispatchBuilder::CalculateDeferredFlags() {
Ref OpDispatchBuilder::IncrementByCarry(OpSize OpSize, Ref Src) {
// If CF not inverted, we use .cc since the increment happens when the
// condition is false. If CF inverted, invert to use .cs. A bit mindbendy.
return _NZCVSelectIncrement(OpSize, {CFInverted ? COND_UGE : COND_ULT}, Src, Src);
return _NZCVSelectIncrement(OpSize, CFInverted ? CondClass::UGE : CondClass::ULT, Src, Src);
}
Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
Ref Res;
@@ -288,11 +286,11 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Src2PlusCF = IncrementByCarry(OpSize, Src2);
// Need to zero-extend for the comparison.
Res = _Add(OpSize, Src1, Src2PlusCF);
Res = Add(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
// TODO: We can fold that second Bfe in (cmp uxth).
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Res, Src2PlusCF, One, Zero);
auto SelectCFInv = Select01(OpSize, CondClass::UGE, Res, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -304,8 +302,6 @@ Ref OpDispatchBuilder::CalculateFlags_ADC(IR::OpSize SrcSize, Ref Src1, Ref Src2
}
Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2) {
auto Zero = _InlineConstant(0);
auto One = _InlineConstant(1);
auto OpSize = SrcSize == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(Src1, Src2);
@@ -325,10 +321,10 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(IR::OpSize SrcSize, Ref Src1, Ref Src2
auto Src2PlusCF = IncrementByCarry(OpSize, Src2);
Res = _Sub(OpSize, Src1, Src2PlusCF);
Res = Sub(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, IR::OpSizeAsBits(SrcSize), 0, Res);
auto SelectCFInv = _Select(FEXCore::IR::COND_UGE, Src1, Src2PlusCF, One, Zero);
auto SelectCFInv = Select01(OpSize, CondClass::UGE, Src1, Src2PlusCF);
SetNZ_ZeroCV(SrcSize, Res);
SetCFInverted(SelectCFInv);
@@ -349,10 +345,10 @@ Ref OpDispatchBuilder::CalculateFlags_SUB(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Res;
if (SrcSize >= OpSize::i32Bit) {
Res = _SubWithFlags(SrcSize, Src1, Src2);
Res = SubWithFlags(SrcSize, Src1, Src2);
} else {
_SubNZCV(SrcSize, Src1, Src2);
Res = _Sub(OpSize::i32Bit, Src1, Src2);
Res = Sub(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
@@ -379,10 +375,10 @@ Ref OpDispatchBuilder::CalculateFlags_ADD(IR::OpSize SrcSize, Ref Src1, Ref Src2
Ref Res;
if (SrcSize >= OpSize::i32Bit) {
Res = _AddWithFlags(SrcSize, Src1, Src2);
Res = AddWithFlags(SrcSize, Src1, Src2);
} else {
_AddNZCV(SrcSize, Src1, Src2);
Res = _Add(OpSize::i32Bit, Src1, Src2);
Res = Add(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
@@ -410,7 +406,7 @@ void OpDispatchBuilder::CalculateFlags_MUL(IR::OpSize SrcSize, Ref Res, Ref High
// If High = SignBit, then sets to nZCv. Else sets to nzcV. Since SF/ZF
// undefined, this does what we need after inverting carry.
auto Zero = _InlineConstant(0);
_CondSubNZCV(OpSize::i64Bit, Zero, Zero, CondClassType {COND_EQ}, 0x1 /* nzcV */);
_CondSubNZCV(OpSize::i64Bit, Zero, Zero, CondClass::EQ, 0x1 /* nzcV */);
CFInverted = true;
}
@@ -427,7 +423,7 @@ void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
// If High = 0, then sets to nZCv. Else sets to nzcV. Since SF/ZF undefined,
// this does what we need.
_CondSubNZCV(Size, Zero, Zero, CondClassType {COND_EQ}, 0x1 /* nzcV */);
_CondSubNZCV(Size, Zero, Zero, CondClass::EQ, 0x1 /* nzcV */);
CFInverted = true;
}
@@ -6,6 +6,7 @@ namespace FEXCore::IR {
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
@@ -71,9 +72,28 @@ constexpr DispatchTableEntry OpDispatch_H0F38Table[] = {
{OPD(PF_38_66, 0x40), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VMUL, OpSize::i32Bit>},
{OPD(PF_38_66, 0x41), 1, &OpDispatchBuilder::PHMINPOSUWOp},
{OPD(PF_38_NONE, 0xC8), 1, &OpDispatchBuilder::SHA1NEXTEOp},
{OPD(PF_38_NONE, 0xC9), 1, &OpDispatchBuilder::SHA1MSG1Op},
{OPD(PF_38_NONE, 0xCA), 1, &OpDispatchBuilder::SHA1MSG2Op},
{OPD(PF_38_NONE, 0xCB), 1, &OpDispatchBuilder::SHA256RNDS2Op},
{OPD(PF_38_NONE, 0xCC), 1, &OpDispatchBuilder::SHA256MSG1Op},
{OPD(PF_38_NONE, 0xCD), 1, &OpDispatchBuilder::SHA256MSG2Op},
{OPD(PF_38_66, 0xDB), 1, &OpDispatchBuilder::AESImcOp},
{OPD(PF_38_66, 0xDC), 1, &OpDispatchBuilder::AESEncOp},
{OPD(PF_38_66, 0xDD), 1, &OpDispatchBuilder::AESEncLastOp},
{OPD(PF_38_66, 0xDE), 1, &OpDispatchBuilder::AESDecOp},
{OPD(PF_38_66, 0xDF), 1, &OpDispatchBuilder::AESDecLastOp},
{OPD(PF_38_NONE, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_66, 0xF0), 2, &OpDispatchBuilder::MOVBEOp},
{OPD(PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF0), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66 | PF_38_F2, 0xF1), 1, &OpDispatchBuilder::CRC32},
{OPD(PF_38_66, 0xF6), 1, &OpDispatchBuilder::ADXOp},
{OPD(PF_38_F3, 0xF6), 1, &OpDispatchBuilder::ADXOp},
};
@@ -29,6 +29,7 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x40), 1, &OpDispatchBuilder::DPPOp<OpSize::i32Bit>},
{OPD(REX, PF_3A_66, 0x41), 1, &OpDispatchBuilder::DPPOp<OpSize::i64Bit>},
{OPD(REX, PF_3A_66, 0x42), 1, &OpDispatchBuilder::MPSADBWOp},
{OPD(REX, PF_3A_66, 0x44), 1, &OpDispatchBuilder::PCLMULQDQOp},
{OPD(REX, PF_3A_66, 0x60), 1, &OpDispatchBuilder::VPCMPESTRMOp},
{OPD(REX, PF_3A_66, 0x61), 1, &OpDispatchBuilder::VPCMPESTRIOp},
@@ -36,6 +37,8 @@ constexpr auto OpDispatchTableGenH0F3A = []() consteval {
{OPD(REX, PF_3A_66, 0x63), 1, &OpDispatchBuilder::VPCMPISTRIOp},
{OPD(REX, PF_3A_NONE, 0xCC), 1, &OpDispatchBuilder::SHA1RNDS4Op},
{OPD(REX, PF_3A_66, 0xDF), 1, &OpDispatchBuilder::AESKeyGenAssist},
};
return std::to_array(Table);
};
@@ -65,11 +68,6 @@ constexpr DispatchTableEntry OpDispatch_H0F3ATableNeedsREX0[] = {
{OPD(0, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i32Bit>},
};
constexpr DispatchTableEntry OpDispatch_H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x16), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PExtrOp, OpSize::i64Bit>},
{OPD(1, PF_3A_66, 0x22), 1, &OpDispatchBuilder::PINSROp<OpSize::i64Bit>},
};
#undef PF_3A_NONE
#undef PF_3A_66
@@ -117,7 +117,9 @@ constexpr DispatchTableEntry OpDispatch_PrimaryGroupTables[] = {
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 0), 1, &OpDispatchBuilder::INCOp}, // INC
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 1), 1, &OpDispatchBuilder::DECOp}, // DEC
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 2), 1, &OpDispatchBuilder::CALLAbsoluteOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 3), 1, &OpDispatchBuilder::CALLFARIndirectOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 4), 1, &OpDispatchBuilder::JUMPAbsoluteOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 5), 1, &OpDispatchBuilder::JUMPFARIndirectOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_5, OpToIndex(0xFF), 6), 1, &OpDispatchBuilder::PUSHOp},
// GROUP 11
@@ -69,10 +69,16 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
// GROUP 9
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_NONE, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 6), 1, &OpDispatchBuilder::RDRANDOp<false>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_66, 7), 1, &OpDispatchBuilder::RDRANDOp<true>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F2, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 1), 1, &OpDispatchBuilder::CMPXCHGPairOp},
{OPD(FEXCore::X86Tables::TYPE_GROUP_9, PF_F3, 7), 1, &OpDispatchBuilder::RDPIDOp},
// GROUP 12
@@ -145,6 +151,9 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_F2, 3), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 3>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_16, PF_F2, 4), 4, &OpDispatchBuilder::NOPOp},
// GROUP 17
{OPD(FEXCore::X86Tables::TYPE_GROUP_17, PF_66, 0), 1, &OpDispatchBuilder::Extrq_imm},
// GROUP P
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_NONE, 0), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, false, false, 1>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_NONE, 1), 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::Prefetch, true, false, 1>},
@@ -156,18 +165,6 @@ constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables[] = {
{OPD(FEXCore::X86Tables::TYPE_GROUP_P, PF_F2, 0), 8, &OpDispatchBuilder::NOPOp},
};
constexpr DispatchTableEntry OpDispatch_SecondaryGroupTables_64[] = {
// GROUP 15
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 0), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::ReadSegmentReg, OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 1), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::ReadSegmentReg, OpDispatchBuilder::Segment::GS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 2), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::WriteSegmentReg, OpDispatchBuilder::Segment::FS>},
{OPD(FEXCore::X86Tables::TYPE_GROUP_15, PF_F3, 3), 1,
&OpDispatchBuilder::Bind<&OpDispatchBuilder::WriteSegmentReg, OpDispatchBuilder::Segment::GS>},
};
#undef OPD
} // namespace FEXCore::IR
@@ -17,6 +17,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryModRMTables[] = {
// REG /7
{((3 << 3) | 0), 1, &OpDispatchBuilder::PermissionRestrictedOp},
{((3 << 3) | 1), 1, &OpDispatchBuilder::RDTSCPOp},
{((3 << 3) | 4), 1, &OpDispatchBuilder::CLZeroOp},
};
} // namespace FEXCore::IR
@@ -198,6 +198,8 @@ constexpr DispatchTableEntry OpDispatch_SecondaryRepNEModTables[] = {
{0x5E, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFDIVSCALARINSERT, OpSize::i64Bit>},
{0x5F, 1, &OpDispatchBuilder::VectorScalarInsertALUOp<IR::OP_VFMAXSCALARINSERT, OpSize::i64Bit>},
{0x70, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PSHUFWOp, true>},
{0x78, 1, &OpDispatchBuilder::Insertq_imm},
{0x79, 1, &OpDispatchBuilder::Insertq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i32Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i32Bit>},
{0xD0, 1, &OpDispatchBuilder::ADDSUBPOp<OpSize::i32Bit>},
@@ -256,6 +258,7 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0x75, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i16Bit>},
{0x76, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VCMPEQ, OpSize::i32Bit>},
{0x78, 1, nullptr}, // GROUP 17
{0x79, 1, &OpDispatchBuilder::Extrq},
{0x7C, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VFADDP, OpSize::i64Bit>},
{0x7D, 1, &OpDispatchBuilder::HSUBP<OpSize::i64Bit>},
{0x7E, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::MOVBetweenGPR_FPR, OpDispatchBuilder::VectorOpType::SSE>},
@@ -313,20 +316,4 @@ constexpr DispatchTableEntry OpDispatch_SecondaryOpSizeModTables[] = {
{0xFD, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i16Bit>},
{0xFE, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::VectorALUOp, IR::OP_VADD, OpSize::i32Bit>},
};
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable_64[] = {
{0x05, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::SyscallOp, true>},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
{0xA9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
};
constexpr DispatchTableEntry OpDispatch_TwoByteOpTable_32[] = {
{0x05, 1, &OpDispatchBuilder::NOPOp},
{0xA0, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA1, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_FS_PREFIX>},
{0xA8, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
{0xA9, 1, &OpDispatchBuilder::Bind<&OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_GS_PREFIX>},
};
} // namespace FEXCore::IR
File diff suppressed because it is too large. Load diff
@@ -28,36 +28,42 @@ class OrderedNode;
Ref OpDispatchBuilder::GetX87Top() {
// Yes, we are storing 3 bits in a single flag register.
// Deal with it
return _LoadContext(OpSize::i8Bit, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
return _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
Ref X87Empty = Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
Ref NewAbridgedFTW {};
_StackForceSlow(); // Invalidate x87 FTW register cache
for (int i = 0; i < 8; i++) {
Ref RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
Ref RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, Constant(1), Constant(0));
// For the output, we want a 1-bit for each pair not equal to 11 (Empty).
static_assert(static_cast<uint8_t>(FPState::X87Tag::Empty) == 0b11);
if (i) {
NewAbridgedFTW = _Orlshl(OpSize::i32Bit, NewAbridgedFTW, RegValid, i);
} else {
NewAbridgedFTW = RegValid;
}
}
// Make even bits 1 if the pair is equal to 11, and 0 otherwise.
FTW = _AndShift(OpSize::i32Bit, FTW, FTW, ShiftType::LSR, 1);
StoreContext(AbridgedFTWIndex, NewAbridgedFTW);
// Invert FTW and clear the odd bits. Even bits are 1 if the pair
// is not equal to 11, and odd bits are 0.
FTW = _Andn(OpSize::i32Bit, Constant(0x55555555), FTW);
// All that's left is to compact away the odd bits. That is a Morton
// deinterleave operation, which has a standard solution. See
// https://stackoverflow.com/questions/3137266/how-to-de-interleave-bits-unmortonizing
FTW = _And(OpSize::i32Bit, _Orlshr(OpSize::i32Bit, FTW, FTW, 1), Constant(0x33333333));
FTW = _And(OpSize::i32Bit, _Orlshr(OpSize::i32Bit, FTW, FTW, 2), Constant(0x0f0f0f0f));
FTW = _Orlshr(OpSize::i32Bit, FTW, FTW, 4);
// ...and that's it. StoreContext implicitly does the final masking.
_StoreContextGPR(OpSize::i8Bit, FTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
}
void OpDispatchBuilder::SetX87Top(Ref Value) {
_StoreContext(OpSize::i8Bit, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
_StoreContextGPR(OpSize::i8Bit, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
}
// Float LoaD operation with memory operand
void OpDispatchBuilder::FLD(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], Width, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
Ref ConvertedData = Data;
// Convert to 80bit float
if (Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
@@ -73,14 +79,14 @@ void OpDispatchBuilder::FLDFromStack(OpcodeArgs) {
void OpDispatchBuilder::FBLD(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
_PushStack(ConvertedData, Data, OpSize::i128Bit, true);
}
void OpDispatchBuilder::FBSTP(OpcodeArgs) {
Ref converted = _F80BCDStore(_ReadStackValue(0));
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
@@ -93,7 +99,7 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs, NamedVectorConstant K) {
void OpDispatchBuilder::FILD(OpcodeArgs) {
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSourceGPR_WithOpSize(Op, Op->Src[0], ReadWidth, Op->Flags);
// Sign extend to 64bits
if (ReadWidth != OpSize::i64Bit) {
@@ -106,15 +112,15 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
// Extract sign and make integer absolute
auto zero = Constant(0);
_SubNZCV(OpSize::i64Bit, Data, zero);
auto sign = _NZCVSelect(OpSize::i64Bit, CondClassType {COND_SLT}, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClassType {COND_MI});
auto sign = _NZCVSelect(OpSize::i64Bit, CondClass::SLT, Constant(0x8000), zero);
auto absolute = _Neg(OpSize::i64Bit, Data, CondClass::MI);
// left justify the absolute integer
auto shift = _Sub(OpSize::i64Bit, Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shift = Sub(OpSize::i64Bit, Constant(63), _FindMSB(IR::OpSize::i64Bit, absolute));
auto shifted = _Lshl(OpSize::i64Bit, absolute, shift);
auto adjusted_exponent = _Sub(OpSize::i64Bit, Constant(0x3fff + 63), shift);
auto zeroed_exponent = _Select(COND_EQ, absolute, zero, zero, adjusted_exponent);
auto adjusted_exponent = Sub(OpSize::i64Bit, Constant(0x3fff + 63), shift);
auto zeroed_exponent = _Select(OpSize::i64Bit, OpSize::i64Bit, CondClass::EQ, absolute, zero, zero, adjusted_exponent);
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
Ref ConvertedData = _VLoadTwoGPRs(shifted, upper);
@@ -160,12 +166,12 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
// Check for NaN/Infinity: exponent = 0x7fff
SaveNZCV();
_TestNZ(OpSize::i64Bit, Exponent, Constant(0x7fff));
Ref IsSpecial = _NZCVSelect(OpSize::i64Bit, {COND_EQ}, Constant(1), Constant(0));
Ref IsSpecial = _NZCVSelect01(CondClass::EQ);
// For overflow detection, check if exponent indicates a value >= 2^15
// Biased exponent for 2^15 is 0x3fff + 15 = 0x400e
_SubWithFlags(OpSize::i64Bit, Exponent, Constant(0x400e));
Ref IsOverflow = _NZCVSelect(OpSize::i64Bit, {COND_UGE}, Constant(1), Constant(0));
SubWithFlags(OpSize::i64Bit, Exponent, 0x400e);
Ref IsOverflow = _NZCVSelect01(CondClass::UGE);
// Set Invalid Operation flag if overflow or special value
Ref InvalidFlag = _Or(OpSize::i64Bit, IsSpecial, IsOverflow);
@@ -174,7 +180,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs, bool Truncate) {
Data = _F80CVTInt(Size, Data, Truncate);
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, Data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, Data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -200,10 +206,10 @@ void OpDispatchBuilder::FADD(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispa
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width);
}
@@ -230,10 +236,10 @@ void OpDispatchBuilder::FMUL(OpcodeArgs, IR::OpSize Width, bool Integer, OpDispa
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width);
}
@@ -267,10 +273,10 @@ void OpDispatchBuilder::FDIV(OpcodeArgs, IR::OpSize Width, bool Integer, bool Re
// We have one memory argument
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _F80CVTTo(arg, Width);
}
@@ -308,10 +314,10 @@ void OpDispatchBuilder::FSUB(OpcodeArgs, IR::OpSize Width, bool Integer, bool Re
// We have one memory argument
Ref Arg {};
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTToInt(Arg, Width);
} else {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _F80CVTTo(Arg, Width);
}
@@ -334,7 +340,7 @@ Ref OpDispatchBuilder::GetX87FTW_Helper() {
// bytes, we use the well-known bit twiddling algorithm:
//
// https://graphics.stanford.edu/~seander/bithacks.html#InterleaveBMN
Ref X = LoadContext(AbridgedFTWIndex);
Ref X = _LoadContextGPR(OpSize::i8Bit, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
X = _Orlshl(OpSize::i32Bit, X, X, 4);
X = _And(OpSize::i32Bit, X, Constant(0x0f0f0f0f));
X = _Orlshl(OpSize::i32Bit, X, X, 2);
@@ -375,41 +381,41 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
_SyncStackToSlow();
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
Ref Mem = LoadSourceGPR(Op, Op->Dest, Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
{
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMemGPR(Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMemGPR(Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1); }
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MemOffsetType::SXTX, 1);
}
}
@@ -421,11 +427,13 @@ Ref OpDispatchBuilder::ReconstructX87StateFromFSW_Helper(Ref FSW) {
auto C1 = _Bfe(OpSize::i32Bit, 1, 9, FSW);
auto C2 = _Bfe(OpSize::i32Bit, 1, 10, FSW);
auto C3 = _Bfe(OpSize::i32Bit, 1, 14, FSW);
auto IE = _Bfe(OpSize::i32Bit, 1, 0, FSW);
SetRFLAG<FEXCore::X86State::X87FLAG_C0_LOC>(C0);
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(C1);
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(C2);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(C3);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(IE);
return Top;
}
@@ -433,20 +441,20 @@ void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
_StackForceSlow();
const auto Size = OpSizeFromSrc(Op);
Ref Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Ref Mem = LoadSourceGPR(Op, Op->Src[0], Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref MemLocation = _Add(OpSize::i64Bit, Mem, Constant(IR::OpSizeToSize(Size) * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 1);
auto NewFSW = _LoadMemGPR(Size, MemLocation, Size);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
Ref MemLocation = _Add(OpSize::i64Bit, Mem, Constant(IR::OpSizeToSize(Size) * 2));
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
Ref MemLocation = Add(OpSize::i64Bit, Mem, IR::OpSizeToSize(Size) * 2);
SetX87FTW(_LoadMemGPR(Size, MemLocation, Size));
}
}
@@ -475,62 +483,61 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
Ref Top = GetX87Top();
{
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMemGPR(Size, Mem, FCW, Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1); }
{ _StoreMemGPR(Size, ReconstructFSW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1); }
auto ZeroConst = Constant(0);
{
// FTW
_StoreMem(GPRClass, Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, GetX87FTW_Helper(), Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction Offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 3), Size, MemOffsetType::SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 4), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer offset
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 5), Size, MemOffsetType::SXTX, 1);
}
{
// Data pointer selector
_StoreMem(GPRClass, Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MEM_OFFSET_SXTX, 1);
_StoreMemGPR(Size, ZeroConst, Mem, Constant(IR::OpSizeToSize(Size) * 6), Size, MemOffsetType::SXTX, 1);
}
auto OneConst = Constant(1);
auto SevenConst = Constant(7);
const auto LoadSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Ref data = _LoadContextFPRIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
_StoreMem(FPRClass, OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
_StoreMemFPR(OpSize::i128Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
// The final st(7) needs a bit of special handling here
Ref data = _LoadContextIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
Ref data = _LoadContextFPRIndexed(Top, LoadSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
if (ReducedPrecisionMode) {
data = _F80CVTTo(data, OpSize::i64Bit);
}
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMemFPR(OpSize::i64Bit, data, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
auto topBytes = _VDupElement(OpSize::i128Bit, OpSize::i16Bit, data, 4);
_StoreMem(FPRClass, OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
_StoreMemFPR(OpSize::i16Bit, topBytes, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (7 * 10) + 8), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// reset to default
FNINIT(Op);
@@ -541,8 +548,8 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
if (ReducedPrecisionMode) {
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
@@ -554,51 +561,48 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
_SetRoundingMode(roundingMode, false, roundingMode);
}
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 1), Size, MemOffsetType::SXTX, 1);
Ref Top = ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1));
}
auto OneConst = Constant(1);
auto SevenConst = Constant(7);
auto low = Constant(~0ULL);
auto high = Constant(0xFFFF);
Ref Mask = _VLoadTwoGPRs(low, high);
const auto StoreSize = ReducedPrecisionMode ? OpSize::i64Bit : OpSize::i128Bit;
for (int i = 0; i < 7; ++i) {
Ref Reg = _LoadMem(FPRClass, OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMemFPR(OpSize::i128Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * i)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
// Mask off the top bits
Reg = _VAnd(OpSize::i128Bit, OpSize::i128Bit, Reg, Mask);
if (ReducedPrecisionMode) {
// Convert to double precision
Reg = _F80CVT(OpSize::i64Bit, Reg);
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
_StoreContextFPRIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
Top = _And(OpSize::i32Bit, Add(OpSize::i32Bit, Top, 1), SevenConst);
}
// The final st(7) needs a bit of special handling here
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
Ref Reg = _LoadMem(FPRClass, OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref RegHigh =
_LoadMem(FPRClass, OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MEM_OFFSET_SXTX, 1);
Ref Reg = _LoadMemFPR(OpSize::i64Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7)), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Ref RegHigh = _LoadMemFPR(OpSize::i16Bit, Mem, Constant((IR::OpSizeToSize(Size) * 7) + (10 * 7) + 8), OpSize::i8Bit, MemOffsetType::SXTX, 1);
Reg = _VInsElement(OpSize::i128Bit, OpSize::i16Bit, 4, 0, Reg, RegHigh);
if (ReducedPrecisionMode) {
Reg = _F80CVT(OpSize::i64Bit, Reg); // Convert to double precision
}
_StoreContextIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit), FPRClass);
_StoreContextFPRIndexed(Reg, Top, StoreSize, MMBaseOffset(), IR::OpSizeToSize(OpSize::i128Bit));
}
// Load / Store Control Word
void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
auto FCW = _LoadContext(OpSize::i16Bit, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
StoreResult(GPRClass, Op, FCW, OpSize::iInvalid);
auto FCW = _LoadContextGPR(OpSize::i16Bit, offsetof(FEXCore::Core::CPUState, FCW));
StoreResultGPR(Op, FCW);
}
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
@@ -606,8 +610,8 @@ void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
// to switch for now to slow mode whenever these are manually changed.
// Remove the next line and try DF_04.asm in fast path.
_StackForceSlow();
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
Ref NewFCW = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
void OpDispatchBuilder::FXCH(OpcodeArgs) {
@@ -643,10 +647,10 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
b = _F80CVTToInt(arg, Width);
} else {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
b = _F80CVTTo(arg, Width);
}
} else {
@@ -677,6 +681,9 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs, IR::OpSize Width, bool Integer, OpDisp
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(PF);
}
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
if (PopTwice) {
_PopStackDestroy();
_PopStackDestroy();
@@ -698,6 +705,9 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Constant(0));
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(HostFlag_Unordered);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(HostFlag_ZF);
// Set Invalid Operation flag when unordered (NaN comparison)
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(HostFlag_Unordered);
}
void OpDispatchBuilder::X87OpHelper(OpcodeArgs, FEXCore::IR::IROps IROp, bool ZeroC2) {
@@ -756,10 +766,17 @@ Ref OpDispatchBuilder::ReconstructFSW_Helper(Ref T) {
void OpDispatchBuilder::X87FNSTSW(OpcodeArgs) {
Ref TopValue = _SyncStackToSlow();
Ref StatusWord = ReconstructFSW_Helper(TopValue);
StoreResult(GPRClass, Op, StatusWord, OpSize::iInvalid);
StoreResultGPR(Op, StatusWord);
}
void OpDispatchBuilder::FNCLEX(OpcodeArgs) {
// Clear the exception flag bit
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(_Constant(0));
}
void OpDispatchBuilder::FNINIT(OpcodeArgs) {
_SyncStackToSlow(); // Invalidate x87 register caches
auto Zero = Constant(0);
if (ReducedPrecisionMode) {
@@ -768,12 +785,12 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
// Init FCW to 0x037F
auto NewFCW = Constant(0x037F);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
// Set top to zero
SetX87Top(Zero);
// Tags all get marked as invalid
StoreContext(AbridgedFTWIndex, Zero);
_StoreContextGPR(OpSize::i8Bit, Zero, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
// Reinits the simulated stack
_InitStack();
@@ -782,6 +799,7 @@ void OpDispatchBuilder::FNINIT(OpcodeArgs) {
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_C3_LOC>(Zero);
SetRFLAG<FEXCore::X86State::X87FLAG_IE_LOC>(Zero);
}
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
@@ -827,11 +845,7 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
default: LOGMAN_MSG_A_FMT("Unhandled FCMOV op: 0x{:x}", Opcode); break;
}
auto ZeroConst = Constant(0);
auto AllOneConst = Constant(0xffff'ffff'ffff'ffffull);
Ref SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SrcCond);
Ref VecCond = _VDupFromGPR(OpSize::i128Bit, OpSize::i64Bit, SelectCC0All1(CC));
_F80VBSLStack(OpSize::i128Bit, VecCond, Op->OP & 7, 0);
}
@@ -847,11 +861,9 @@ void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
// Claim this is a normal number
// We don't support anything else
auto TopValid = _StackValidTag(0);
auto ZeroConst = Constant(0);
auto OneConst = Constant(1);
// In the case of top being invalid then C3:C2:C0 is 0b101
auto C3 = _Select(FEXCore::IR::COND_NEQ, TopValid, OneConst, OneConst, ZeroConst);
auto C3 = Select01(OpSize::i32Bit, CondClass::NEQ, TopValid, Constant(1));
auto C2 = TopValid;
auto C0 = C3; // Mirror C3 until something other than zero is supported
@@ -29,38 +29,38 @@ void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
const auto Size = OpSizeFromSrc(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, OpSize::i16Bit, Mem, OpSize::i16Bit);
auto NewFCW = _LoadMemGPR(OpSize::i16Bit, Mem, OpSize::i16Bit);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
auto NewFSW = _LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MEM_OFFSET_SXTX, 1);
auto NewFSW = _LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size)), Size, MemOffsetType::SXTX, 1);
ReconstructX87StateFromFSW_Helper(NewFSW);
{
// FTW
SetX87FTW(_LoadMem(GPRClass, Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MEM_OFFSET_SXTX, 1));
SetX87FTW(_LoadMemGPR(Size, Mem, Constant(IR::OpSizeToSize(Size) * 2), Size, MemOffsetType::SXTX, 1));
}
}
void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
_StackForceSlow();
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Ref NewFCW = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode, false, roundingMode);
_StoreContext(OpSize::i16Bit, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContextGPR(OpSize::i16Bit, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
// F64 ops
// Float load op with memory operand
void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
const auto ReadWidth = (Width == OpSize::f80Bit) ? OpSize::i128Bit : Width;
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], Width, Op->Flags);
// Convert to 64bit float
Ref ConvertedData = Data;
if (Width == OpSize::i32Bit) {
@@ -73,7 +73,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs, IR::OpSize Width) {
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
// Read from memory
Ref Data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], OpSize::i128Bit, Op->Flags);
Ref Data = LoadSourceFPR_WithOpSize(Op, Op->Src[0], OpSize::f80Bit, Op->Flags);
Ref ConvertedData = _F80BCDLoad(Data);
ConvertedData = _F80CVT(OpSize::i64Bit, ConvertedData);
_PushStack(ConvertedData, Data, OpSize::i64Bit, true);
@@ -82,7 +82,7 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
Ref converted = _F80CVTTo(_ReadStackValue(0), OpSize::i64Bit);
converted = _F80BCDStore(converted);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
StoreResultFPR_WithOpSize(Op, Op->Dest, converted, OpSize::f80Bit, OpSize::i8Bit);
_PopStackDestroy();
}
@@ -95,7 +95,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
const auto ReadWidth = OpSizeFromSrc(Op);
// Read from memory
Ref Data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], ReadWidth, Op->Flags);
Ref Data = LoadSourceGPR_WithOpSize(Op, Op->Src[0], ReadWidth, Op->Flags);
if (ReadWidth == OpSize::i16Bit) {
Data = _Sbfe(OpSize::i64Bit, IR::OpSizeAsBits(ReadWidth), 0, Data);
}
@@ -112,7 +112,7 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs, bool Truncate) {
} else {
data = _Float_ToGPR_S(Size == OpSize::i32Bit ? OpSize::i32Bit : OpSize::i64Bit, OpSize::i64Bit, data);
}
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, OpSize::i8Bit);
StoreResultGPR_WithOpSize(Op, Op->Dest, data, Size, OpSize::i8Bit);
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
_PopStackDestroy();
@@ -138,16 +138,16 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDi
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
@@ -176,16 +176,16 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpDi
Ref arg {};
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
} else {
FEX_UNREACHABLE;
}
@@ -228,16 +228,16 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
Arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
Arg = _Sbfe(OpSize::i64Bit, 16, 0, Arg);
}
Arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, Arg);
} else if (Width == OpSize::i32Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
Arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, Arg);
} else if (Width == OpSize::i64Bit) {
Arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -285,16 +285,16 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs, IR::OpSize Width, bool Integer, bool
if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
arg = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
arg = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -332,16 +332,16 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs, IR::OpSize Width, bool Integer, OpD
} else if (Width == OpSize::i16Bit || Width == OpSize::i32Bit || Width == OpSize::i64Bit) {
// Memory arg
if (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceGPR(Op, Op->Src[0], Op->Flags);
if (Width == OpSize::i16Bit) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(OpSize::i64Bit, Width == OpSize::i64Bit ? OpSize::i64Bit : OpSize::i32Bit, arg);
} else if (Width == OpSize::i32Bit) {
arg = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
arg = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
b = _Float_FToF(OpSize::i64Bit, OpSize::i32Bit, arg);
} else if (Width == OpSize::i64Bit) {
b = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
b = LoadSourceFPR(Op, Op->Src[0], Op->Flags);
}
} else {
FEX_UNREACHABLE;
@@ -382,7 +382,7 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
// non zero case
Ref ExpNZ = _Bfe(OpSize::i64Bit, 11, 52, Gpr);
ExpNZ = _Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
ExpNZ = Sub(OpSize::i64Bit, ExpNZ, Constant(1023));
Ref ExpNZV = _Float_FromGPR_S(OpSize::i64Bit, OpSize::i64Bit, ExpNZ);
Ref SigNZ = _And(OpSize::i64Bit, Gpr, Constant(0x800f'ffff'ffff'ffffLL));
@@ -393,8 +393,8 @@ void OpDispatchBuilder::X87FXTRACTF64(OpcodeArgs) {
SaveNZCV();
_TestNZ(OpSize::i64Bit, Gpr, Constant(0x7fff'ffff'ffff'ffffUL));
Ref Sig = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, {COND_EQ}, ExpZV, ExpNZV);
Ref Sig = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, SigZV, SigNZV);
Ref Exp = _NZCVSelectV(OpSize::i64Bit, CondClass::EQ, ExpZV, ExpNZV);
_PopStackDestroy();
_PushStack(Exp, Exp, OpSize::i64Bit, true);
@@ -33,7 +33,7 @@ X86GeneratedCode::X86GeneratedCode() {
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
memcpy(reinterpret_cast<void*>(CallbackReturn), &SignalReturnCode.at(0), SignalReturnCode.size());
memcpy(reinterpret_cast<void*>(CallbackReturn), SignalReturnCode.data(), SignalReturnCode.size());
mprotect(CodePtr, CODE_SIZE, PROT_READ);
#endif
@@ -1,29 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
meta: frontend|x86-tables ~ Metadata that drives the frontend x86/64 decoding
tags: frontend|x86-tables
$end_info$
*/
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/Core/Context.h>
namespace FEXCore::X86Tables {
void InitializeBaseTables(Context::OperatingMode Mode);
void InitializeSecondaryTables(Context::OperatingMode Mode);
void InitializeSecondaryGroupTables(Context::OperatingMode Mode);
void InitializePrimaryGroupTables(Context::OperatingMode Mode);
void InitializeH0F3ATables(Context::OperatingMode Mode);
void InitializeInfoTables(Context::OperatingMode Mode) {
InitializeBaseTables(Mode);
InitializeSecondaryTables(Mode);
InitializeSecondaryGroupTables(Mode);
InitializePrimaryGroupTables(Mode);
InitializeH0F3ATables(Mode);
}
} // namespace FEXCore::X86Tables
@@ -15,300 +15,425 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
enum Primary_LUT {
ENTRY_06,
ENTRY_07,
ENTRY_0E,
ENTRY_16,
ENTRY_17,
ENTRY_1E,
ENTRY_1F,
ENTRY_27,
ENTRY_2F,
ENTRY_37,
ENTRY_3F,
ENTRY_40,
ENTRY_48,
ENTRY_60,
ENTRY_61,
ENTRY_63,
ENTRY_9A,
ENTRY_A0,
ENTRY_A1,
ENTRY_A2,
ENTRY_A3,
ENTRY_CE,
ENTRY_D4,
ENTRY_D5,
ENTRY_D6,
ENTRY_EA,
ENTRY_MAX,
};
constexpr std::array<X86InstInfo[2], ENTRY_MAX> Primary_ArchSelect_LUT = {{
// ENTRY_06
{
{"PUSH ES", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_07
{
{"POP ES", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_ES_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_0E
{
{"PUSH CS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_CS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_16
{
{"PUSH SS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_17
{
{"POP SS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_SS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_1E
{
{"PUSH DS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::PUSHSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_1F
{
{"POP DS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::Bind<&IR::OpDispatchBuilder::POPSegmentOp, FEXCore::X86Tables::DecodeFlags::FLAG_DS_PREFIX> } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_27
{
{"DAA", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::DAAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_2F
{
{"DAS", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::DASOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_37
{
{"AAA", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::AAAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_3F
{
{"AAS", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::AASOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_40
{
{"INC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, { .OpDispatch = &IR::OpDispatchBuilder::INCOp } },
// REX
{"", TYPE_REX_PREFIX, FLAGS_NONE, 0},
},
// ENTRY_48
{
{"DEC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, { .OpDispatch = &IR::OpDispatchBuilder::DECOp } },
{"", TYPE_REX_PREFIX, FLAGS_NONE, 0},
},
// ENTRY_60
{
{"PUSHA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::PUSHAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_61
{
{"POPA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, { .OpDispatch = &IR::OpDispatchBuilder::POPAOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_63
{
{"ARPL", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"MOVSXD", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM, 0, { .OpDispatch = &IR::OpDispatchBuilder::MOVSXDOp } },
},
// ENTRY_9A
{
{"CALLF", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_A0
{
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A1
{
{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A2
{
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_A3
{
{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, { .OpDispatch = &IR::OpDispatchBuilder::MOVOffsetOp } },
},
// ENTRY_CE
{
{"INTO", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = &IR::OpDispatchBuilder::INTOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D4
{
{"AAM", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, { .OpDispatch = &IR::OpDispatchBuilder::AAMOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D5
{
{"AAD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, { .OpDispatch = &IR::OpDispatchBuilder::AADOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_D6
{
{"SALC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, { .OpDispatch = &IR::OpDispatchBuilder::SALCOp } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
// ENTRY_EA
{
{"JMPF", TYPE_INST, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
{"", TYPE_INVALID, FLAGS_NONE, 0, { .OpDispatch = nullptr } },
},
}};
const std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct BaseOpTable[] = {
// Prefixes
// Operand size overide
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0}},
// Address size override
{0x67, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x26, 1, X86InstInfo{"ES", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x2E, 1, X86InstInfo{"CS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x36, 1, X86InstInfo{"SS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x3E, 1, X86InstInfo{"DS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x67, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0}},
{0x26, 1, X86InstInfo{"ES", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x2E, 1, X86InstInfo{"CS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x36, 1, X86InstInfo{"SS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
{0x3E, 1, X86InstInfo{"DS", TYPE_LEGACY_PREFIX, FLAGS_NONE, 0}},
// These are still invalid on 64bit
{0x64, 1, X86InstInfo{"FS", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x65, 1, X86InstInfo{"GS", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF0, 1, X86InstInfo{"LOCK", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF2, 1, X86InstInfo{"REPNE", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x64, 1, X86InstInfo{"FS", TYPE_PREFIX, FLAGS_NONE, 0}},
{0x65, 1, X86InstInfo{"GS", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF0, 1, X86InstInfo{"LOCK", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF2, 1, X86InstInfo{"REPNE", TYPE_PREFIX, FLAGS_NONE, 0}},
{0xF3, 1, X86InstInfo{"REP", TYPE_PREFIX, FLAGS_NONE, 0}},
// Instructions
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x05, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x00, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x01, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x02, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x03, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_MODRM, 0}},
{0x04, 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x05, 1, X86InstInfo{"ADD", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x06, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_06] }}},
{0x07, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_07] }}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x15, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x08, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x09, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x0A, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x0B, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_MODRM, 0}},
{0x0C, 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x0D, 1, X86InstInfo{"OR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x0E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_0E] }}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0, nullptr}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x1D, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x10, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x11, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x12, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x13, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_MODRM, 0}},
{0x14, 1, X86InstInfo{"ADC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x15, 1, X86InstInfo{"ADC", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x16, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_16] }}},
{0x17, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_17] }}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x18, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x19, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_DISPLACE_SIZE_DIV_2, 0}},
{0x1A, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x1B, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_MODRM, 0}},
{0x1C, 1, X86InstInfo{"SBB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x1D, 1, X86InstInfo{"SBB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x1E, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1E] }}},
{0x1F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_1F] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x20, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x21, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x22, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x23, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_MODRM, 0}},
{0x24, 1, X86InstInfo{"AND", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x25, 1, X86InstInfo{"AND", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x35, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x27, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_27] }}},
{0x28, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x29, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x2A, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x2B, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_MODRM, 0}},
{0x2C, 1, X86InstInfo{"SUB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x2D, 1, X86InstInfo{"SUB", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x2F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_2F] }}},
{0x38, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x39, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x3A, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x3B, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x3C, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0x3D, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x30, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x31, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x32, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x33, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_MODRM, 0}},
{0x34, 1, X86InstInfo{"XOR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x35, 1, X86InstInfo{"XOR", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x50, 8, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0x58, 8, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0x37, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_37] }}},
{0x38, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x39, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x3A, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x3B, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_MODRM, 0}},
{0x3C, 1, X86InstInfo{"CMP", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0x3D, 1, X86InstInfo{"CMP", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x3F, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_3F] }}},
{0x62, 1, X86InstInfo{"", TYPE_GROUP_EVEX, FLAGS_NONE, 0, nullptr}},
{0x40, 8, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_40] }}},
{0x48, 8, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_48] }}},
{0x68, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SRC_SEXT, 4, nullptr}},
{0x69, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0x6A, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SRC_SEXT , 1, nullptr}},
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1, nullptr}},
{0x50, 8, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0x58, 8, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SF_REX_IN_BYTE | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0x60, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_60] }}},
{0x61, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_61] }}},
{0x62, 1, X86InstInfo{"", TYPE_GROUP_EVEX, FLAGS_NONE, 0}},
{0x63, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_63] }}},
{0x68, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_SRC_SEXT, 4}},
{0x69, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0x6A, 1, X86InstInfo{"PUSH", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SRC_SEXT , 1}},
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1}},
// This should just throw a GP
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x72, 1, X86InstInfo{"JB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x73, 1, X86InstInfo{"JNB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x74, 1, X86InstInfo{"JZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x75, 1, X86InstInfo{"JNZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x76, 1, X86InstInfo{"JBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x77, 1, X86InstInfo{"JNBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x78, 1, X86InstInfo{"JS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x79, 1, X86InstInfo{"JNS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7A, 1, X86InstInfo{"JP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7B, 1, X86InstInfo{"JNP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7C, 1, X86InstInfo{"JL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7D, 1, X86InstInfo{"JNL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7E, 1, X86InstInfo{"JLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x7F, 1, X86InstInfo{"JNLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x72, 1, X86InstInfo{"JB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x73, 1, X86InstInfo{"JNB", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x74, 1, X86InstInfo{"JZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x75, 1, X86InstInfo{"JNZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x76, 1, X86InstInfo{"JBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x77, 1, X86InstInfo{"JNBE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x78, 1, X86InstInfo{"JS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x79, 1, X86InstInfo{"JNS", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7A, 1, X86InstInfo{"JP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7B, 1, X86InstInfo{"JNP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7C, 1, X86InstInfo{"JL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7D, 1, X86InstInfo{"JNL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7E, 1, X86InstInfo{"JLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x7F, 1, X86InstInfo{"JNLE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x84, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x85, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x86, 1, X86InstInfo{"XCHG", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x87, 1, X86InstInfo{"XCHG", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x8A, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0, nullptr}},
{0x8B, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{0x8C, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM, 0, nullptr}},
{0x8E, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0, nullptr}},
{0x8F, 1, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_ZERO_REG | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x90, 8, X86InstInfo{"XCHG", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x98, 1, X86InstInfo{"CDQE", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x99, 1, X86InstInfo{"CQO", TYPE_INST, FLAGS_SF_DST_RDX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0x88, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x89, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x8A, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM, 0}},
{0x8B, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM, 0}},
{0x8C, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 0}},
{0x8D, 1, X86InstInfo{"LEA", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM, 0}},
{0x8E, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_16BIT) | FLAGS_MODRM, 0}},
{0x8F, 1, X86InstInfo{"POP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_ZERO_REG | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0x90, 8, X86InstInfo{"XCHG", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_SF_SRC_RAX, 0}},
{0x98, 1, X86InstInfo{"CDQE", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0}},
{0x99, 1, X86InstInfo{"CQO", TYPE_INST, FLAGS_SF_DST_RDX | FLAGS_SF_SRC_RAX, 0}},
// These three are all X87 instructions
{0x9B, 1, X86InstInfo{"FWAIT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9C, 1, X86InstInfo{"PUSHF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0, nullptr}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_BLOCK_END, 0, nullptr}},
{0x9B, 1, X86InstInfo{"FWAIT", TYPE_INST, FLAGS_NONE, 0}},
{0x9C, 1, X86InstInfo{"PUSHF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF), 0}},
{0x9D, 1, X86InstInfo{"POPF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_BLOCK_END, 0}},
{0x9E, 1, X86InstInfo{"SAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9F, 1, X86InstInfo{"LAHF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x9E, 1, X86InstInfo{"SAHF", TYPE_INST, FLAGS_NONE, 0}},
{0x9F, 1, X86InstInfo{"LAHF", TYPE_INST, FLAGS_NONE, 0}},
{0xA4, 1, X86InstInfo{"MOVSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA5, 1, X86InstInfo{"MOVS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA6, 1, X86InstInfo{"CMPSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA7, 1, X86InstInfo{"CMPS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xA0, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A0] }}},
{0xA1, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A1] }}},
{0xA2, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A2] }}},
{0xA3, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_A3] }}},
{0xA8, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1, nullptr}},
{0xA9, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{0xAA, 1, X86InstInfo{"STOS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAB, 1, X86InstInfo{"STOS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAC, 1, X86InstInfo{"LODS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xAD, 1, X86InstInfo{"LODS", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xAE, 1, X86InstInfo{"SCAS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xAF, 1, X86InstInfo{"SCAS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xA4, 1, X86InstInfo{"MOVSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA5, 1, X86InstInfo{"MOVS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA6, 1, X86InstInfo{"CMPSB", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xA7, 1, X86InstInfo{"CMPS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xB0, 8, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_REX_IN_BYTE , 1, nullptr}},
{0xB8, 8, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_DISPLACE_SIZE_MUL_2, 4, nullptr}},
{0xA8, 1, X86InstInfo{"TEST", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX , 1}},
{0xA9, 1, X86InstInfo{"TEST", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4}},
{0xAA, 1, X86InstInfo{"STOS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAB, 1, X86InstInfo{"STOS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAC, 1, X86InstInfo{"LODS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xAD, 1, X86InstInfo{"LODS", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xAE, 1, X86InstInfo{"SCAS", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xAF, 1, X86InstInfo{"SCAS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_SF_SRC_RAX, 0}},
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2, nullptr}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0, nullptr}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3, nullptr}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0, nullptr}},
{0xCA, 2, X86InstInfo{"RETF", TYPE_PRIV, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1, nullptr}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0, nullptr}},
{0xB0, 8, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_REX_IN_BYTE , 1}},
{0xB8, 8, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_REX_IN_BYTE | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_DISPLACE_SIZE_MUL_2, 4}},
{0xD7, 1, X86InstInfo{"XLAT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0xC2, 1, X86InstInfo{"RET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2}},
{0xC3, 1, X86InstInfo{"RET", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END , 0}},
{0xC8, 1, X86InstInfo{"ENTER", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 3}},
{0xC9, 1, X86InstInfo{"LEAVE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_DEBUG_MEM_ACCESS , 0}},
{0xCA, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 2}},
{0xCB, 1, X86InstInfo{"RETF", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
{0xCC, 1, X86InstInfo{"INT3", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xCD, 1, X86InstInfo{"INT", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 1}},
{0xCE, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_CE] }}},
{0xCF, 1, X86InstInfo{"IRET", TYPE_INST, FLAGS_SETS_RIP | FLAGS_BLOCK_END, 0}},
{0xE0, 1, X86InstInfo{"LOOPNE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE1, 1, X86InstInfo{"LOOPE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE2, 1, X86InstInfo{"LOOP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1, nullptr}},
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0xD4, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D4] }}},
{0xD5, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D5] }}},
{0xD6, 1, X86InstInfo{"", TYPE_ARCH_DISPATCHER, FLAGS_NONE, 0, { .Indirect = Primary_ArchSelect_LUT[ENTRY_D6] }}},
{0xD7, 1, X86InstInfo{"XLAT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0}},
{0xE0, 1, X86InstInfo{"LOOPNE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE1, 1, X86InstInfo{"LOOPE", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE2, 1, X86InstInfo{"LOOP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_SF_SRC_RCX, 1}},
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1}},
// Should just throw GP
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 1, nullptr}},
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 1}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 1}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END | FLAGS_CALL , 4, nullptr}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1, nullptr}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END | FLAGS_CALL , 4}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1}},
// Should just throw GP
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0}},
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0}},
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0}},
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0}},
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0}},
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0}},
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0}},
// Two Byte table
{0x0F, 1, X86InstInfo{"", TYPE_SECONDARY_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x0F, 1, X86InstInfo{"", TYPE_SECONDARY_TABLE_PREFIX, FLAGS_NONE, 0}},
// x87 table
{0xD8, 8, X86InstInfo{"", TYPE_X87_TABLE_PREFIX, FLAGS_MODRM, 0, nullptr}},
{0xD8, 8, X86InstInfo{"", TYPE_X87_TABLE_PREFIX, FLAGS_MODRM, 0}},
// ModRM table
// MoreBytes field repurposed for valid bits mask
{0x80, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 0, nullptr}},
{0x81, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 1, nullptr}},
{0x82, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 2, nullptr}},
{0x83, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 3, nullptr}},
{0xC0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 0, nullptr}},
{0xC1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 1, nullptr}},
{0xD0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 2, nullptr}},
{0xD1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 3, nullptr}},
{0xD2, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 4, nullptr}},
{0xD3, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 5, nullptr}},
{0xF6, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 0, nullptr}},
{0xF7, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 1, nullptr}},
{0xFE, 1, X86InstInfo{"", TYPE_GROUP_4, FLAGS_MODRM, 0, nullptr}},
{0xFF, 1, X86InstInfo{"", TYPE_GROUP_5, FLAGS_MODRM, 0, nullptr}},
{0x80, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 0}},
{0x81, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 1}},
{0x82, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 2}},
{0x83, 1, X86InstInfo{"", TYPE_GROUP_1, FLAGS_MODRM, 3}},
{0xC0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 0}},
{0xC1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 1}},
{0xD0, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 2}},
{0xD1, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 3}},
{0xD2, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 4}},
{0xD3, 1, X86InstInfo{"", TYPE_GROUP_2, FLAGS_MODRM, 5}},
{0xF6, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 0}},
{0xF7, 1, X86InstInfo{"", TYPE_GROUP_3, FLAGS_MODRM, 1}},
{0xFE, 1, X86InstInfo{"", TYPE_GROUP_4, FLAGS_MODRM, 0}},
{0xFF, 1, X86InstInfo{"", TYPE_GROUP_5, FLAGS_MODRM, 0}},
// Group 11
{0xC6, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 0, nullptr}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 1, nullptr}},
{0xC6, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 0}},
{0xC7, 1, X86InstInfo{"", TYPE_GROUP_11, FLAGS_MODRM, 1}},
// VEX table
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0}},
};
GenerateTable(&Table.at(0), BaseOpTable, std::size(BaseOpTable));
GenerateTable(Table.data(), BaseOpTable, std::size(BaseOpTable));
IR::InstallToTable(Table, IR::OpDispatch_BaseOpTable);
return Table;
}();
void InitializeBaseTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct BaseOpTable_64[] = {
{0x06, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x0E, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x16, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x1E, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x27, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2F, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x37, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x3F, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// REX
{0x40, 16, X86InstInfo{"", TYPE_REX_PREFIX, FLAGS_NONE, 0, nullptr}},
{0x60, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x63, 1, X86InstInfo{"MOVSXD", TYPE_INST, GenFlagsDstSize(SIZE_64BIT) | FLAGS_MODRM, 0, nullptr}},
{0x9A, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xA0, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA2, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA1, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xA3, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 8, nullptr}},
{0xCE, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xD4, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// `L1OM` Larrabee instructions used this as an escape byte.
// FEX will never support this.
{0xD6, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xEA, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
static constexpr U8U8InfoStruct BaseOpTable_32[] = {
{0x06, 1, X86InstInfo{"PUSH ES", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x07, 1, X86InstInfo{"POP ES", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x0E, 1, X86InstInfo{"PUSH CS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x16, 1, X86InstInfo{"PUSH SS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x17, 1, X86InstInfo{"POP SS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1E, 1, X86InstInfo{"PUSH DS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1F, 1, X86InstInfo{"POP DS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x27, 1, X86InstInfo{"DAA", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x2F, 1, X86InstInfo{"DAS", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x37, 1, X86InstInfo{"AAA", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x3F, 1, X86InstInfo{"AAS", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x40, 8, X86InstInfo{"INC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
{0x48, 8, X86InstInfo{"DEC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
{0x60, 1, X86InstInfo{"PUSHA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x61, 1, X86InstInfo{"POPA", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x63, 1, X86InstInfo{"ARPL", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x9A, 1, X86InstInfo{"CALLF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xA0, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA2, 1, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA1, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA3, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xCE, 1, X86InstInfo{"INTO", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xD4, 1, X86InstInfo{"AAM", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xD5, 1, X86InstInfo{"AAD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xD6, 1, X86InstInfo{"SALC", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX | FLAGS_SF_SRC_RAX, 0, nullptr}},
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
if (Mode == Context::MODE_64BIT) {
GenerateTable(&BaseOps.at(0), BaseOpTable_64, std::size(BaseOpTable_64));
IR::InstallToTable(BaseOps, IR::OpDispatch_BaseOpTable_64);
}
else {
GenerateTable(&BaseOps.at(0), BaseOpTable_32, std::size(BaseOpTable_32));
IR::InstallToTable(BaseOps, IR::OpDispatch_BaseOpTable_32);
}
}
}
Loaded 100 of 623 files, more files were not shown because too many files have changed in this diff. Show more