Compare commits

..
802 Commits
Author SHA1 Message Date
Ryan Houdek 2f5ebf1dd1 Docs: Update for release FEX-2212 2022-12-05 14:02:57 -08:00
Ryan Houdek 8f157e45bb Merge pull request #2198 from lioncash/add
OpcodeDispatcher: Handle VADDPD/VADDPS/VPADDB/VPADDW/VPADDD/VPADDQ
2022-12-05 13:19:30 -08:00
lioncash 6b259e2731 OpcodeDispatcher: Handle VPADDQ 2022-12-05 17:53:33 +00:00
lioncash 200660aba5 OpcodeDispatcher: Handle VPADDD 2022-12-05 17:44:12 +00:00
lioncash 065c12cfbb OpcodeDispatcher: Handle VPADDW 2022-12-05 17:33:52 +00:00
lioncash 318972620f OpcodeDispatcher: Handle VPADDB 2022-12-05 17:23:39 +00:00
lioncash e7f54d1592 OpcodeDispatcher: Handle VADDPD 2022-12-05 16:59:04 +00:00
lioncash a8571282b2 OpcodeDispatcher: Handle VADDPS 2022-12-05 16:43:26 +00:00
Ryan Houdek fc28062052 Merge pull request #2193 from Sonicadvance1/support_radeon_ioctl_emu
IoctlEmu: Support radeon
2022-12-05 06:36:44 -08:00
Mai cc6306aa32 Merge pull request #2194 from Sonicadvance1/fix_confusing_error
FEXServerClient: Disable confusing connection log
2022-12-05 13:39:50 +00:00
Ryan Houdek f0caa81253 FEXServerClient: Disable confusing connection log
On first FEXInterpreter execution, it is expected that `ConnectToServer`
will fail with `ECONNREFUSED` because FEXServer won't be running.

Skip printing this first messaage to stderr if configured.
If it is some other error message then ensure it is still printed.
2022-12-05 05:18:41 -08:00
Ryan Houdek cd98871f8f IoctlEmu: Support radeon 2022-12-05 05:12:04 -08:00
Mai 66e0d46d89 Merge pull request #2196 from Sonicadvance1/optimize_symlink_following
Linux: Improve performance of hot paths in path searching
2022-12-05 13:05:22 +00:00
Mai 1dd5642e46 Merge pull request #2195 from Sonicadvance1/improve_interpreter_check
FEXLoader: Make `IsInterpreterInstalled` check less horrible.
2022-12-05 13:04:11 +00:00
Mai 9ca34ca306 Merge pull request #2178 from Sonicadvance1/const_jit
Arm64: Const on unmodified argument
2022-12-05 13:03:21 +00:00
Ryan Houdek 69f39a0bc3 Arm64: Const on unmodified argument
Just noticed this when tinkering around the JIT.
These arguments can safely be const.
2022-12-05 03:54:17 -08:00
Ryan Houdek 50cf74db0a Linux: Improve performance of hot paths in path searching
`GetEmulatedPath` and `OpenAt` are called a /lot/ in applications.
std::filesystem::path handling here is quite heavy and costly for what
we are trying to achieve.

Remove this usage and instead use lstat, access, and readlink directly
which is a heck of a lot faster.

In particular this helps out pressure-vessel, shaving off launch times
by 1-2 seconds.
Going from ~22 seconds down to ~20 seconds.
2022-12-05 03:51:50 -08:00
Ryan Houdek 6886d8ff65 FEXLoader: Make IsInterpreterInstalled check less horrible.
`std::filesystem::exists` is particularly gnarly in how it checks to see
if the file exists.
It allocates memory, it creates lists, it splits things, then eventually
checking the status

Remove all this overhead to help out minorly for applications that
execve a lot.
2022-12-05 02:58:15 -08:00
Mai c1d118c1d4 Merge pull request #2191 from Sonicadvance1/minor_aeskeygenassist_optimization
Arm64: Minor optimization in AESKEYGENASSIST
2022-12-04 06:59:31 +00:00
Ryan Houdek 5e46d63c42 Arm64: Minor optimization in AESKEYGENASSIST
The less number of FPR<->GPR movement instructions the better.
This removes one instance of `ins` and replaces the other with a 64-bit
`dup` instead.
The LoadConstant still turns in to a single `movz` instruction with the
shift.
2022-12-03 03:59:27 -08:00
xianwei zheng 863a59a8e2 Thunk: Crash on XSetErrorHandler(NULL) (#2190)
* BUGFIX:Adding the nullptr check to FinalizeHostTrampolineForGuestFunction. It will lead crash on interpreter code set function nullptr callback. eg. XSetErrorHandler(NULL)

* fix #2189
2022-11-30 21:57:09 -08:00
Ryan Houdek c37fcf136a Merge pull request #2187 from lioncash/and
OpcodeDispatcher: Handle VANDPD/VANDPS/VPAND/VANDNPD/VANDNPS/VPANDN
2022-11-30 16:22:42 -08:00
lioncash 02a2292115 OpcodeDispatcher: Handle VPANDN 2022-11-30 15:51:10 +00:00
lioncash a483bc9837 OpcodeDispatcher: Handle VANDNPD 2022-11-30 15:51:10 +00:00
lioncash 120a6b85f4 OpcodeDispatcher: Handle VANDNPS 2022-11-30 15:51:05 +00:00
Ryan Houdek 9fea774a93 Merge pull request #2188 from lioncash/pclmul
unittests: Expand VPCLMULQDQ unit test
2022-11-29 17:15:34 -08:00
lioncash bf1e619ead unittests: Expand vpclmulqdq unit test
Now that we have some AVX instructions in place, we can make the test
use them and also enforce correctness behavior in the upper lane.
2022-11-29 22:07:38 +00:00
lioncash 0f8fcfc43e OpcodeDispatcher: Handle VPAND 2022-11-29 19:08:58 +00:00
lioncash 698b7fda06 OpcodeDispatcher: Handle VANDPD 2022-11-29 19:06:25 +00:00
lioncash 23caa6e20f OpcodeDispatcher: Handle VANDPS 2022-11-29 19:04:29 +00:00
Ryan Houdek 34e39c996e Merge pull request #2186 from lioncash/or
OpcodeDispatcher: Handle VORPD/VORPS/VPOR
2022-11-29 10:57:17 -08:00
lioncash 16ed20cfae OpcodeDispatcher: Handle VPOR 2022-11-29 18:43:38 +00:00
lioncash ef368ceafa OpcodeDispatcher: Handle VORPD 2022-11-29 18:40:04 +00:00
lioncash 45480ef32c OpcodeDispatcher: Handle VORPS 2022-11-29 18:38:13 +00:00
Ryan Houdek 4de69029e5 Merge pull request #2185 from lioncash/xor
OpcodeDispatcher: Handle VPXOR/VXORPD/VXORPS
2022-11-29 10:28:02 -08:00
lioncash c065770f48 OpcodeDispatcher: Handle VPXOR 2022-11-29 18:12:04 +00:00
lioncash 27957ea051 OpcodeDispatcher: Handle VXORPD 2022-11-29 18:05:28 +00:00
lioncash 94e9d1ab3b OpcodeDispatcher: Handle VXORPS 2022-11-29 18:04:29 +00:00
Ryan Houdek a374a9af35 Merge pull request #2184 from lioncash/vzero
OpcodeDispatcher: Handle VZEROUPPER/VZEROALL
2022-11-29 08:33:07 -08:00
lioncash 3e80416eb6 OpcodeDispatcher: Handle VZEROUPPER/VZEROALL 2022-11-29 16:15:32 +00:00
Ryan Houdek b35c6c6d22 Merge pull request #2183 from lioncash/vmovq
OpcodeDispatcher: Handle VMOVQ
2022-11-28 18:33:26 -08:00
lioncash 69d26cfee6 OpcodeDispatcher: Handle combined VMOVQ/VMOVD 2022-11-29 01:49:58 +00:00
lioncash e3be1540f1 OpcodeDispatcher: Handle VMOVQ
Fairly trivial, we can reuse the existing implementation for MOVQ.
2022-11-29 01:49:15 +00:00
Ryan Houdek 8e2b0d10e5 Merge pull request #2181 from Sonicadvance1/defer_cpuinfo
EmulatedFiles: Defer cpuinfo file initialization to first access
2022-11-28 17:48:42 -08:00
Ryan Houdek 57c5761920 Merge pull request #2179 from Sonicadvance1/tsl_maps
Core: Replace a couple maps with tsl robin_map
2022-11-28 17:48:27 -08:00
Ryan Houdek 0841ff5feb Merge pull request #2182 from lioncash/ntdq
OpcodeDispatcher: Handle VMOVNTDQ/VMOVNTDQA/VMOVNTPD/VMOVNTPS
2022-11-28 17:47:27 -08:00
lioncash 45115384b5 OpcodeDispatcher: Handle VMOVNTPD 2022-11-28 16:58:37 +00:00
lioncash dae2563850 OpcodeDispatcher: Handle VMOVNTPS 2022-11-28 16:56:00 +00:00
lioncash fb4df5a0b7 OpcodeDispatcher: Handle VMOVNTDQA 2022-11-28 16:51:06 +00:00
lioncash b2f0303d1e OpcodeDispatcher: Handle VMOVNTDQ 2022-11-28 16:49:01 +00:00
Ryan Houdek f8b2a0b4d8 Merge pull request #2180 from Sonicadvance1/disable_aot_stores
FEXLoader: Disables some AOT shutdown overhead when not enabled
2022-11-28 01:08:39 -08:00
Ryan Houdek e1fcb78ce3 FEXLoader: Disables some AOT shutdown overhead when not enabled
When AOT wasn't enabled it was still doing some accesses to the
filesystem on shutdown.

Slightly improves shutdown time.
2022-11-27 16:02:58 -08:00
Ryan Houdek d6309088c0 EmulatedFiles: Defer cpuinfo file initialization to first access
This improves startup time by a couple of milliseconds.

Most applications don't query cpuinfo so deferring improves most
application's startup times.
2022-11-27 15:34:10 -08:00
Ryan Houdek 1fb3a2e28f Core: Replace a couple maps with tsl robin_map
Improves the shutdown time a small amount and performance of these maps.
2022-11-27 04:35:27 -08:00
Mai df25d4e03e Merge pull request #2177 from Sonicadvance1/disable_multiblock_default
Config: Disable multiblock by default
2022-11-26 06:23:24 +00:00
Ryan Houdek 9c2f0287e0 Config: Disable multiblock by default
This causes users pain currently since the JIT isn't doing any caching
and our RA being a hack mess means it is quite slow.

Disable by default to improve JIT time performance.
2022-11-25 18:08:25 -08:00
Mai 9912d41714 Merge pull request #2174 from Sonicadvance1/emulated_sgdt
OpcodeDispatcher: Implement SGDT
2022-11-25 05:23:56 +00:00
Ryan Houdek 8b2cd87d9e unittests: Disable SGDT tests on host
The Zen+ CI runner doesn't support the UMIP hardware feature, so it
doesn't hit the kernel emulated path.

Instead the instruction returns real data on this hardware. Still in
kernel space, so it is unmapped as expected.
2022-11-24 18:29:05 -08:00
Ryan Houdek 3e6d23ae7e unittests: SGDT tests 2022-11-24 17:47:31 -08:00
Ryan Houdek c96c39d5b1 OpcodeDispatcher: Implement SGDT
Ran in to this when running Team Sonic Racing.
This game uses Denuvo Anti-Tamper which some versions rely on SGDT.

The Linux kernel catches and emulates this instruction.
It will always return a limit of 0 and a base of `0xFFFFFFFFFFFE0000ULL`
which is guaranteed to be in kernel space.

This gets the game slightly farther but still not running entirely.
2022-11-24 17:43:52 -08:00
Mai a4556e90cd Merge pull request #2170 from Sonicadvance1/optimize_sra_step_one
OpcodeDispatcher: Moves all GPR and XMM accesses to direct register accesses
2022-11-24 06:17:55 +00:00
Mai 4d2c4b4423 Merge pull request #2172 from Sonicadvance1/minor_vector_initialization_improvement
Syscalls: Minor optimization with initialization of syscall definition vector
2022-11-23 22:16:36 +00:00
Ryan Houdek 530de3f031 Syscalls: Minor optimization with initialization of syscall definition vector
Shaves a couple milliseconds off initialization time.

Noticed this while poking around.
2022-11-23 12:59:08 -08:00
Ryan Houdek c9622f6fd4 unittests/IR: Update tests for new IR semantics 2022-11-22 23:06:18 -08:00
Ryan Houdek 854628d959 IREmitter relies on CoreState.h now 2022-11-22 23:06:18 -08:00
Ryan Houdek 35dbf6c44b OpcodeDispatcher: Moves all GPR and XMM accesses to direct register accesses
This is a bit of a large commit since it is an all or nothing sort of
change.

Instead of doing LoadContext and StoreContext for GPRs and FPRs, any
that are statically allocated (GPR and XMM) should use
{Load,Store}Register directly.

This is now enforced that {Load,Store}Context{,Indexed} will assert if
trying to access these ranges.

This is the first step towards accelerating our JIT compile times in
that it removes the need for the SRA pass to convert all the IR over.

This ensures that from the Dispatcher directly we are handling SRA
accesses as "registers", ensuring that future changes don't need to run
in to this problem.

Even though it wasn't a target of this change, in a simple test
application this removes ~12% of the total JIT compile time.
2022-11-22 23:06:18 -08:00
Ryan Houdek b9fec7436f X86Tables: Fixes some incorrectly defined instruction sizes.
These were working because of partial loadstore handling with context
loadstores
2022-11-22 21:06:58 -08:00
Ryan Houdek 808d19c374 IR: Disallow loading and storing registers through {Load,Store}Context{,Indexed}
All SRA allocated registers must be handled explicitly through {Load,Store}Register
2022-11-22 21:06:58 -08:00
Ryan Houdek 441d7205ed Passes: Remove StaticRegisterAllocation pass 2022-11-22 21:06:58 -08:00
Ryan Houdek 8b3d3b68c6 Jit64: Implement {Load,Store}Register 2022-11-22 21:06:58 -08:00
Ryan Houdek aa0e038ef7 Interpreter: Implement {Load,Store}Register 2022-11-22 21:06:58 -08:00
Ryan Houdek 5e5e5a35d9 Merge pull request #2157 from Sonicadvance1/more_systemd_stuff
FEXServer: More Systemd fixes
2022-11-22 03:17:17 -08:00
Ryan Houdek 10e35a55ea FEXServerClient: Cleanup AF_UNIX abstract socket string math
Makes it a bit easier to reason about.
2022-11-22 03:05:33 -08:00
Ryan Houdek 83cebea780 FEXServerClient: Clean up comments in server mount folder. 2022-11-22 02:56:31 -08:00
Ryan Houdek 91bbb92c50 FEXServer: More Systemd fixes
Two changes here.

- Make the mount path follow server temp folder requirements.
  - Will be mounted in `/tmp/` or `$XDG_RUNTIME_DIR/` now
- Switch the FEXServer socket to an "abstract" AF_UNIX socket.
  - If the socket is in `/tmp/` then systemd will put the service in a
    private `/tmp` folder that only exists for the service.
  - If the socket is in `$XDG_RUNTIME_DIR` then pressure-vessel can't
    chroot anymore since they make their own runtime directory.
  - If it is in `$HOME/.fex-emu/` then it breaks usage where the
    filesystem is a mount that doesn't support AF_UNIX like sshfs.

The only reasonable thing to do is to switch over to `abstract` sockets
which will work in all cases.
Tested with pressure-vessel and systemd and now it works in all
situations.
2022-11-22 02:55:08 -08:00
Ryan Houdek c7dd6ff28a Merge pull request #2167 from Sonicadvance1/optimize_break_codegen
Arm64: Optimize Break IR op codegen
2022-11-21 23:37:39 -08:00
Ryan Houdek df761a99ce Merge pull request #2171 from lioncash/dqa
OpcodeDispatcher: Handle VMOVDQA/VMOVDQU
2022-11-21 23:29:21 -08:00
Ryan Houdek 359416e2b6 Arm64: Optimize Break IR op codegen
This isn't really a performance issue, more just something that is gross
looking at while looking at unit tests.

Break /usually/ isn't abused heavily by games (Except Denuvo) so not
really a perf concern regardless.
2022-11-21 23:20:15 -08:00
lioncash e140c0d60c OpcodeDispatcher: Handle VMOVDQU 2022-11-22 06:47:55 +00:00
Ryan Houdek 0030971f6f Merge pull request #2168 from Sonicadvance1/cmake_typo
CMake: Fix typo in clang thunks option.
2022-11-21 22:46:52 -08:00
lioncash 3a90aaf1e6 OpcodeDispatcher: Implement VMOVDQA 2022-11-22 06:40:07 +00:00
Ryan Houdek f8a199af49 Merge pull request #2169 from lioncash/movddup
OpcodeDispatcher: Handle VMOVDDUP
2022-11-21 22:21:35 -08:00
lioncash 3a5de8e10c OpcodeDispatcher: Handle VMOVDDUP 2022-11-22 05:33:53 +00:00
Ryan Houdek 3c881809f4 Merge pull request #2166 from Sonicadvance1/wrapnode_helper
IntrusiveIRList: Add a utility helper for getting an OrderedNodeWrapper
2022-11-21 21:33:30 -08:00
Ryan Houdek 8b6e9e08c0 Merge pull request #2165 from Sonicadvance1/remove_migrate_log
Core: Removes log about migrating to shared memory mode
2022-11-21 21:33:11 -08:00
Ryan Houdek 3c8da3e3b4 Merge pull request #2164 from Sonicadvance1/debug_logs_on_bad_socket
FEXServerClient: Add some debug logs for when FEX can't connect to se…
2022-11-21 21:33:00 -08:00
Ryan Houdek 0a39d909b2 CMake: Fix typo in clang thunks option. 2022-11-21 21:11:01 -08:00
Ryan Houdek d35d1092a4 IntrusiveIRList: Add a utility helper for getting an OrderedNodeWrapper
This is a nice helper that was otherwise missing.
2022-11-21 21:02:39 -08:00
Ryan Houdek 7ae655c56a Core: Removes log about migrating to shared memory mode
This hasn't ever caused problems and instead just adds a message that
appears in logs for nearly every application invocation.

Remove it because it isn't necessary to track.
2022-11-21 21:00:10 -08:00
Ryan Houdek 58f35ba413 Merge pull request #2163 from lioncash/movshdup
OpcodeDispatcher: Handle VMOVSHDUP/VMOVSLDUP
2022-11-21 20:58:49 -08:00
Ryan Houdek e61eb24ec2 FEXServerClient: Add some debug logs for when FEX can't connect to server
Sometimes when the socket fails to connect we have no debug information
at all as to why.

This at least gives us a little bit more.
2022-11-21 20:58:36 -08:00
lioncash 2e93d2ce51 OpcodeDispatcher: Simplify SSE MOVSLDUP
Like with MOVSHDUP, we only need to duplicate two values rather than
four.
2022-11-22 04:38:39 +00:00
lioncash 9d21e1efd5 OpcodeDispatcher: Handle VMOVSLDUP 2022-11-22 04:37:48 +00:00
lioncash e0e6b3ad6b OpcodeDispatcher: Simplify SSE MOVSHDUP
We only need to insert two values, rather than four.
2022-11-22 04:16:55 +00:00
lioncash 815cdc5b3c OpcodeDispatcher: Handle VMOVSHDUP 2022-11-22 04:16:34 +00:00
Ryan Houdek bc20f1e684 Merge pull request #2162 from lioncash/vmovhps
OpcodeDispatcher: Handle VMOVHPD/VMOVHPS
2022-11-21 18:02:06 -08:00
lioncash 20e5f2bec6 OpcodeDispatcher: Handle VMOVHPD 2022-11-22 01:06:47 +00:00
lioncash c3b6fa55b6 OpcodeDispatcher: Handle VMOVHPS 2022-11-22 01:01:28 +00:00
Ryan Houdek 69045db3a9 Merge pull request #2161 from lioncash/vmovlps
OpcodeDispatcher: Handle VMOVLPD/VMOVLPS
2022-11-21 14:20:35 -08:00
lioncash 7b2240c80b OpcodeDispatcher: Handle VMOVLPD 2022-11-21 21:45:58 +00:00
lioncash dcfbd90dd7 OpcodeDispatcher: Handle VMOVLPS 2022-11-21 21:45:36 +00:00
Ryan Houdek 70e6ab5782 Merge pull request #2160 from lioncash/mov
Arm64/VectorOps: Simplify VMov IR op on SVE
2022-11-21 12:40:02 -08:00
lioncash 175879823f Arm64/VectorOps: Add future clarifying context comment
Just so this doesn't get overlooked in the distant future.
2022-11-21 20:19:03 +00:00
Ryan Houdek 2271a90adb Merge pull request #2159 from lioncash/vmovapd
OpcodeDispatcher: Handle VMOVAPD/VMOVUPD/VMOVUPS
2022-11-21 12:08:55 -08:00
lioncash 2bdde5845e Arm64/VectorOps: Simplify VMov IR op on SVE
Initially I put this in very conservatively to make sure we always clear
out the upper lanes, but since Adv. SIMD operations have zero-extending
behavior when storing results, we can just use a lot of operations as
is, without needing to unnecessarily do the same work twice.
2022-11-21 20:05:58 +00:00
lioncash 0c40497a01 OpcodeDispatcher: Handle VMOVUPD 2022-11-21 17:14:25 +00:00
lioncash 45dff0f550 OpcodeDecoder: Handle VMOVUPS 2022-11-21 17:06:31 +00:00
lioncash e9035ef6ee OpcodeDecoder: Handle VMOVAPD 2022-11-21 17:06:27 +00:00
Ryan Houdek 02ca94e6e6 Merge pull request #2158 from Sonicadvance1/steam_appid_configs
Config: Add support for steamid based configurations.
2022-11-18 14:40:27 -08:00
Ryan Houdek 432b7d2dc8 Config: Add support for steamid based configurations.
This will be useful for keying specific executables to steamids.
This is sadly required because a bunch of games end up naming themselves
"game.exe" so we can't safely enable thunks for all things shipping a
generic name.
2022-11-17 18:27:42 -08:00
Ryan Houdek 181d315d2c Merge pull request #2155 from lioncash/x86
x86_64/VectorOps: Separate 128-bit/256-bit paths
2022-11-16 19:23:53 -08:00
Mai d5f7e616eb Merge pull request #2156 from Sonicadvance1/fexserver_systemd_fixes
Systemd fixes
2022-11-17 02:20:48 +00:00
Ryan Houdek 9a8869e8a4 FEXServer: Send shutdown signal to image mount programs on exit. 2022-11-16 18:08:32 -08:00
Ryan Houdek 85d56ed76f FEXServerClient: Print a message when the server socket fails 2022-11-16 18:08:32 -08:00
Ryan Houdek 06c827b5c8 FEXServerClient: Support XDG_RUNTIME_DIR
In the case of a platform enabling PrivateTmp then the FEXServer and
FEXInterpreter won't have a tmp folder that shares the socket location.

The runtime directory is a perfect place to share these across
processes.
2022-11-16 18:08:02 -08:00
Ryan Houdek 41259ff361 FEXServer: Don't deparent when running as a systemd service.
Due to how systemd watches PIDs, it will terminate the entire process
tree if the first pid exits.
2022-11-16 18:08:00 -08:00
Ryan Houdek cccee1a668 FEXServer: Ensure fatal messages are printed 2022-11-16 18:08:00 -08:00
lioncash c46b35362b x86_64/VectorOps: Separate 128-bit VShlI path 2022-11-16 19:46:12 +00:00
lioncash 2f96a6d8bf x86_64/VectorOps: Separate 128-bit VSShrI path 2022-11-16 19:44:46 +00:00
lioncash c78a47a3e5 x86_64/VectorOps: Separate 128-bit VUShrI path 2022-11-16 19:40:50 +00:00
lioncash b049721683 x86_64/VectorOps: Separate 128-bit VUABDL path 2022-11-16 19:37:01 +00:00
lioncash 11c06fe9fe x86_64/VectorOps: Separate 128-bit VMul path 2022-11-16 19:32:04 +00:00
lioncash c5f6e53d0d x86_64/VectorOps: Separate 128-bit VSShrS path 2022-11-16 19:29:13 +00:00
lioncash 1184672bb1 x86_64/VectorOps: Separate 128-bit VUShrS path 2022-11-16 19:27:47 +00:00
lioncash 57d3a2ba35 x86_64/VectorOps: Separate 128-bit VUShlS path 2022-11-16 19:26:20 +00:00
lioncash 1d813b0183 x86_64/VectorOps: Separate 128-bit VFCMPUNO path 2022-11-16 19:22:48 +00:00
lioncash 9b24931518 x86_64/VectorOps: Separate 128-bit VFCMPORD path 2022-11-16 19:22:01 +00:00
lioncash 7e5b8b7bdf x86_64/VectorOps: Separate 128-bit VFCMPLE path 2022-11-16 19:20:56 +00:00
lioncash 7e36473aff x86_64/VectorOps: Separate 128-bit VFCMPGT path 2022-11-16 19:19:28 +00:00
lioncash adbb512306 x86_64/VectorOps: Separate 128-bit VFCMPLT path 2022-11-16 19:18:01 +00:00
lioncash b951f4ad4b x86_64/VectorOps: Separate 128-bit VFCMPNEQ path 2022-11-16 19:16:57 +00:00
lioncash 48900662ae x86_64/VectorOps: Separate 128-bit VFCMPEQ path 2022-11-16 19:15:48 +00:00
lioncash da31a66c07 x86_64/VectorOps: Separate 128-bit VCMPLTZ path 2022-11-16 19:13:00 +00:00
lioncash ca710b1cbb x86_64/VectorOps: Separate 128-bit VCMPGTZ path 2022-11-16 19:10:39 +00:00
lioncash a4d7eec145 x86_64/VectorOps: Separate 128-bit VCMPGT path 2022-11-16 19:08:30 +00:00
lioncash 232c2fe87f x86_64/VectorOps: Separate 128-bit VCMPEQZ path 2022-11-16 19:06:54 +00:00
lioncash 661112cfd4 x86_64/VectorOps: Separate 128-bit VCMPEQ path 2022-11-16 19:01:31 +00:00
lioncash d4a84eaa9a x86_64/VectorOps: Separate 128-bit VBSL path 2022-11-16 18:59:40 +00:00
lioncash 99f8af64d0 x86_64/VectorOps: Separate 128-bit VSMax path 2022-11-16 18:57:16 +00:00
lioncash 1541ea9ffc x86_64/VectorOps: Separate 128-bit VUMax path 2022-11-16 18:55:23 +00:00
lioncash ced86e693c x86_64/VectorOps: Separate 128-bit VSMin path 2022-11-16 18:53:46 +00:00
lioncash 53920d5bd3 x86_64/VectorOps: Separate 128-bit VUMin path 2022-11-16 18:51:56 +00:00
lioncash 15c5a9dac0 x86_64/VectorOps: Separate 128-bit VNot path 2022-11-16 18:49:13 +00:00
lioncash d63cfdbe7d x86_64/VectorOps: Separate 128-bit VFNeg path 2022-11-16 18:46:35 +00:00
lioncash 569461a01c x86_64/VectorOps: Separate 128-bit VNeg path 2022-11-16 18:41:16 +00:00
lioncash 37c7dee236 x86_64/VectorOps: Separate 128-bit VFRSqrt path 2022-11-16 18:35:43 +00:00
lioncash cc230091c8 x86_64/VectorOps: Separate 128-bit VFSqrt path 2022-11-16 18:33:14 +00:00
lioncash bac33cf246 x86_64/VectorOps: Separate 128-bit VFRecp path 2022-11-16 18:31:38 +00:00
lioncash 7a9c0506b4 x86_64/VectorOps: Separate 128-bit VFMax path 2022-11-16 18:28:31 +00:00
lioncash fbd7c15a4b x86_64/VectorOps: Separate 128-bit VFMin path 2022-11-16 18:27:00 +00:00
lioncash 64edf24bc7 x86_64/VectorOps: Separate 128-bit VFDiv path 2022-11-16 18:25:00 +00:00
lioncash 3f456d683f x86_64/VectorOps: Separate 128-bit VFMul path 2022-11-16 18:23:21 +00:00
lioncash aac0824fd3 x86_64/VectorOps: Separate 128-bit VFSub path 2022-11-16 18:21:29 +00:00
lioncash 1b32ca0b93 x86_64/VectorOps: Separate 128-bit VFAddP path 2022-11-16 18:19:13 +00:00
lioncash c29456aac9 x86_64/VectorOps: Separate 128-bit VFAdd path 2022-11-16 18:17:23 +00:00
lioncash 715f25d059 x86_64/VectorOps: Separate 128-bit VAbs path 2022-11-16 18:13:51 +00:00
lioncash c85e31ec7a x86_64/VectorOps: Separate 128-bit VURAvg path 2022-11-16 18:10:53 +00:00
lioncash 0d6bfc4fa4 x86_64/VectorOps: Separate 128-bit VSQSub path 2022-11-16 18:07:59 +00:00
lioncash 106917add2 x86_64/VectorOps: Separate 128-bit VSQAdd path 2022-11-16 18:06:17 +00:00
lioncash 1e10a2bac3 x86_64/VectorOps: Separate 128-bit VUQSub path 2022-11-16 18:02:09 +00:00
lioncash 63dd09cd6f x86_64/VectorOps: Separate 128-bit VUQAdd path 2022-11-16 17:56:37 +00:00
lioncash c7e6935f42 x86_64/VectorOps: Separate 128-bit VSub path 2022-11-16 17:54:57 +00:00
lioncash 1be5054e86 x86_64/VectorOps: Separate 128-bit VAdd path 2022-11-16 17:51:57 +00:00
lioncash f40755aca7 x86_64/VectorOps: Separate 128-bit VXor path 2022-11-16 17:49:25 +00:00
lioncash d49b78cf34 x86_64/VectorOps: Separate 128-bit VOr path 2022-11-16 17:46:49 +00:00
lioncash 10e80ae064 x86_64/VectorOps: Separate 128-bit VBic path 2022-11-16 17:44:10 +00:00
lioncash f3ebb214aa x86_64/VectorOps: Separate 128-bit VAnd path 2022-11-16 17:40:25 +00:00
lioncash dc41c3bd9f x86_64/VectorOps: Separate 128-bit VectorImm path 2022-11-16 17:40:14 +00:00
Ryan Houdek 56ff09f3ac Merge pull request #2154 from lioncash/decoder
OpcodeDispatcher: Handle VMOVAPS
2022-11-15 12:10:41 -08:00
lioncash 1707b27d14 OpcodeDecoder: Only install AVX ops if host supports it 2022-11-15 19:57:08 +00:00
lioncash 7c92963eca x86_64/MemoryOps: Use unaligned loads/stores for 256-bit cases
Alignment will always be a little finicky in this case, so let's just
use unaligned variants to make behavior always consistent.
2022-11-15 19:57:08 +00:00
lioncash ecf82c90ee OpcodeDispatcher: Handle VMOVAPS 2022-11-15 19:57:05 +00:00
lioncash 580f06fe00 Frontend: Handle 256-bit vectors 2022-11-15 03:11:34 +00:00
Ryan Houdek 5a403b7765 Merge pull request #2153 from lioncash/extr
IR: Handle 256-bit VExtr
2022-11-14 15:00:12 -08:00
lioncash f7367e56af IR: Handle 256-bit VExtr
Extends VExtr to handle 256-bit vectors.
2022-11-14 22:26:52 +00:00
Ryan Houdek f066abc151 Merge pull request #2152 from lioncash/vixl
Externals: Update vixl submodule
2022-11-14 09:49:36 -08:00
lioncash 2bee92f332 Externals: Update vixl submodule
Incorporates an upstreamed fix that allows movprfx to be used with
destructive SVE EXT.
2022-11-14 15:45:11 +00:00
Mai 7d9ed4e1bf Merge pull request #2151 from Sonicadvance1/remove_VSLI_VSRI
IR: Removes the only uses of VSLI and VSRI
2022-11-14 05:15:27 +00:00
Ryan Houdek 24696e6b98 IR: Removes the only uses of VSLI and VSRI
We use these two IR ops for PSRLQ and PSLLQ respectively, Which is
actually implemented in a quite inefficient way.

Instead switch the ops over to using VExtr which maps directly to one
instruction and can emulate both.
2022-11-13 19:51:31 -08:00
Mai 71f658b07d Merge pull request #2150 from Sonicadvance1/sort_named_rootfs
FEXConfig: Sort named rootfs vector
2022-11-13 23:14:38 +00:00
Ryan Houdek 920f56353f FEXConfig: Sort named rootfs vector
This was driving me nuts that this wasn't sorted by name.
2022-11-13 14:36:47 -08:00
Ryan Houdek e0fe9167ea Merge pull request #2149 from Sonicadvance1/calculate_minstack_size
ELFCodeLoader: Calculate AT_MINSIGSTKSZ
2022-11-13 14:26:59 -08:00
Ryan Houdek 45a2349a0d ELFCodeLoader: Calculate AT_MINSIGSTKSZ
This is the last remaining auxv value that we were missing.

This requires a little bit of setup to match what we are doing in
FEXCore's Dispatcher.

This only tracks how much space is required by the kernel to store its
required state.
2022-11-13 14:09:51 -08:00
Ryan Houdek d7b0e8469e Merge pull request #2148 from Sonicadvance1/fix_null_end
ELFCodeLoader: Fixes AT_PLATFORM null terminator
2022-11-13 13:58:47 -08:00
Mai 8afc3b8e23 Merge pull request #2147 from Sonicadvance1/auxv_secure
ELFCodeLoader: Pass through AT_SECURE
2022-11-13 21:49:27 +00:00
Ryan Houdek 9caa63d5b2 ELFCodeLoader: Fixes AT_PLATFORM null terminator
We were failing to null terminate the platform which is causing strcmp
to fail.
2022-11-13 13:46:36 -08:00
Mai 1d32df91ce Merge pull request #2146 from Sonicadvance1/fix_vsyscall_auxv
ELFCodeLoader: Ensure we set AT_SYSINFO for 32-bit
2022-11-13 20:52:46 +00:00
Ryan Houdek fa973f65bf ELFCodeLoader: Pass through AT_SECURE
When we are installed as a binfmt_misc handler the Linux kernel will
assign us AT_SECURE for setuid binaries.

Ensure we are passing this through for any application that will end up
needing it.
2022-11-13 12:50:58 -08:00
Ryan Houdek c027f02e5a ELFCodeLoader: Ensure we set AT_SYSINFO for 32-bit
AT_SYSINFO points to the vsyscall location that the kernel provides.
This AUXV value isn't used on x86-64.
If we have VDSO installed then we can use the one provided from there,
otherwise we need to provide a code page.

It seems like some behaviour changed with glibc provided in Ubuntu 22.10
that it now requires AT_SYSINFO.

Fixes wine 7.0 execution inside the Ubuntu 22.10 rootfs.
2022-11-13 12:39:07 -08:00
Ryan Houdek 9cee0126d7 Merge pull request #2145 from lioncash/memload
IR: Remove VLoadMemElement and VStoreMemElement
2022-11-11 11:44:54 -08:00
lioncash c713602d56 IR: Remove VLoadMemElement and VStoreMemElement
These are unused, so we can get rid of them.
2022-11-11 19:06:38 +00:00
Mai c8293cbbda Merge pull request #2137 from Sonicadvance1/pclmul_disable
OpcodeDispatcher: Disable PCLMUL if not supported on host
2022-11-10 22:23:48 +00:00
Ryan Houdek a9c51388cf Merge pull request #2144 from lioncash/reg
IR: Handle 256-bit LoadRegister/StoreRegister
2022-11-08 21:31:51 -08:00
lioncash ad1d65e91a IR: Handle 256-bit LoadRegister
Extends LoadRegister to handle 256-bit vectors.
2022-11-09 03:19:38 +00:00
lioncash 3a6c7803e8 IR: Handle 256-bit StoreRegister
Extends StoreRegister to handle 256-bit vectors.
2022-11-09 03:19:34 +00:00
Ryan Houdek aa837eddd5 Merge pull request #2143 from Sonicadvance1/kinetic_support
InstallFEX.py: Adds support for Kinetic
2022-11-08 14:38:52 -08:00
Ryan Houdek faef57838f InstallFEX.py: Adds support for Kinetic
Also removes Hirsute and Impish since Ubuntu's PPA system doesn't
support these anymore.

Fixes #2140
2022-11-08 14:26:59 -08:00
Ryan Houdek 04d4c5e017 Merge pull request #2141 from lioncash/addv
IR: Handle 256-bit VAddV
2022-11-08 12:14:47 -08:00
Ryan Houdek 9158877569 Merge pull request #2139 from neobrain/fix_thunks_guest_ide_integration
Thunks: Fix guest targets not being detected by IDEs
2022-11-08 12:09:44 -08:00
lioncash 1eae07f1b8 IR: Handle 256-bit VAddV
Extends VAddV to handle 256-bit vectors.
2022-11-08 18:21:30 +00:00
Tony Wasserka 14b22487f1 Thunks: Fix guest targets not being detected by IDEs
The IDE integration path didn't set up the BITNESS variable. This change
also unmarks that variable as a CMake option since it's not a boolean value.
2022-11-08 15:30:49 +01:00
Ryan Houdek 1ac7cd5835 OpcodeDispatcher: Disable PCLMUL if not supported on host
Might fix Steam on Pi4
2022-11-05 12:11:45 -07:00
Mai 5336f01725 Merge pull request #2135 from Sonicadvance1/release_process
Update release process to include AUR
2022-11-03 19:38:19 +00:00
Ryan Houdek b8b66b1829 Update release process to include AUR 2022-11-03 01:23:32 -07:00
Ryan Houdek fd3e988a20 Docs: Update for release FEX-2211 2022-11-02 23:25:10 -07:00
Ryan Houdek b1d98f4e58 Merge pull request #2134 from lioncash/temp
Arm64/ConversionOps: Eliminate use of temporary in Vector_FToF
2022-11-02 19:18:40 -07:00
Ryan Houdek 9e7daf61d0 Merge pull request #2133 from lioncash/inselem
IR: Handle 256-bit VInsElement
2022-11-02 18:59:18 -07:00
lioncash 6fbe25753b IR: Handle 256-bit VInsElement
Extends VInsElem to handle 256-bit vectors.
2022-11-03 01:43:42 +00:00
Ryan Houdek 03f0edc5b5 Merge pull request #2132 from lioncash/indexed
IR: Handle 256-bit LoadContextIndexed/StoreContextIndexed
2022-11-02 17:13:36 -07:00
lioncash 5536f1e835 Arm64/ConversionOps: Eliminate use of temporary in Vector_FToF
We can just use the destination register in this case.
2022-11-02 23:50:28 +00:00
lioncash 0de36706da IR.json: Expand allowed size in LoadContext and StoreContext IR ops
These can now handle 256-bit destinations
2022-11-02 16:19:05 +00:00
lioncash 17722dad6d IR: Handle 256-bit LoadContextIndexed
Extends LoadContextIndexed to handle 256-bit vectors.
2022-11-02 16:16:58 +00:00
lioncash 0371599996 IR: Handle 256-bit StoreContextIndexed
Extends StoreContextIndexed to handle 256-bit vectors.
2022-11-02 16:07:02 +00:00
Ryan Houdek 199649b30f Merge pull request #2131 from lioncash/simplify
Arm64/MemoryOps: Merge if statement into switch in ParanoidLoadMemTSO
2022-11-01 21:35:25 -07:00
lioncash 4ef35488db Arm64/MemoryOps: Merge if statement into switch in ParanoidLoadMemTSO
There's nothing preventing the OpSize == 1 case from being merged into
the switch, so we can do that to make things a little more consistent.
2022-11-02 03:45:17 +00:00
Ryan Houdek 70a91ee6ce Merge pull request #2130 from lioncash/memory
IR: Handle 256-bit StoreMem/StoreMemTSO/ParanoidStoreMemTSO
2022-11-01 20:41:34 -07:00
lioncash 418a27e47e IR: Handle 256-bit ParanoidStoreMemTSO
Extends ParanoidStoreMemTSO to handle 256-bit vectors.
2022-11-01 23:04:30 +00:00
lioncash 61c76d02cc IR: Handle 256-bit StoreMemTSO
Extends StoreMemTSO to handle 256-bit vectors.
2022-11-01 22:59:29 +00:00
lioncash d98641221d IR: Handle 256-bit StoreMem
Extends StoreMem to handle 256-bit vectors.
2022-11-01 21:39:52 +00:00
Ryan Houdek 8a14f87a44 Merge pull request #2129 from lioncash/memory
IR: Handle 256-bit LoadMem/LoadMemTSO/ParanoidLoadMemTSO
2022-11-01 14:23:19 -07:00
lioncash 02ce71734c IR: handle 256-bit ParanoidLoadTSO
Extends ParanoidLoadTSO to handle 256-bit vectors.
2022-11-01 21:02:42 +00:00
lioncash 96c2743280 IR: Handle 256-bit LoadMemTSO
Extends LoadMemTSO to handle 256-bit vectors.
2022-11-01 21:02:42 +00:00
lioncash 7bfc34b51c IR: Handle 256-bit LoadMem
Extends LoadMem to handle 256-bit vectors.
2022-11-01 21:02:38 +00:00
Ryan Houdek 40d820fd05 Merge pull request #2127 from lioncash/spill
Arm64/MemoryOps: Remove lingering unnecessary ptrue instances
2022-11-01 10:47:15 -07:00
lioncash d69287aaf7 Arm64/MemoryOps: Remove lingering unnecessary ptrue instances
Gets rid of some leftover bits from when we didn't have statically
allocated predicate registers.
2022-11-01 17:25:24 +00:00
Ryan Houdek d475b0ba9e Merge pull request #2126 from lioncash/ctx
IR: Handle 256-bit LoadContext/StoreContext
2022-11-01 10:22:26 -07:00
lioncash 1638b744b7 x86_64/MemoryOps: Ensure upper lane is cleared properly in FillRegister
Ensures that loaded values don't potentially have junk in the upper
lane. Will prevent potential wonky situations when implementing AVX
instructions.
2022-11-01 16:39:18 +00:00
lioncash 8b19894a06 IR: Handle 256-bit StoreContext
Extends StoreContext to handle 256-bit vectors.
2022-11-01 16:20:24 +00:00
Ryan Houdek d2e0dc99de Merge pull request #2125 from lioncash/unused
Interpreter/MiscOps: Remove unused StopThread() function
2022-11-01 09:13:31 -07:00
lioncash d04e40b5fd Interpreter/MiscOps: Remove unused StopThread() function
This has been unused since ff1d51c7bd

Silences a compiler warning.
2022-11-01 15:59:17 +00:00
lioncash 75d797b5cd IR: Handle 256-bit LoadContext
Extends LoadContext to handle 256-bit vectors.
2022-11-01 14:54:26 +00:00
Mai ecf4891087 Merge pull request #1668 from Sonicadvance1/wip_segment_register
Segment register index optimization
2022-11-01 02:54:44 +00:00
Ryan Houdek 0e1a418678 WIP: Segment register index optimization
Segment registers are indexed significantly more than they are changed.
Pay the cost of indexing during the set and store rather than the per
register index.

Should be a fairly significant performance improvement for 32-bit
applications. At least on hardware that doesn't have a data dependent
prefetcher.

Breaks Steam atm and isn't clean.
2022-10-31 19:42:30 -07:00
Mai 5bef13df94 Merge pull request #2124 from Sonicadvance1/gvisor_flakes
unittests/gvisor: Adds a bunch of tests to flakes
2022-10-31 21:03:38 +00:00
Ryan Houdek d8386121a8 Merge pull request #2115 from Sonicadvance1/fix_x11_thunk_recursion
Thunks/libX11: Fix recursive initialize
2022-10-31 13:41:07 -07:00
Ryan Houdek 000677abb6 Merge pull request #2078 from Sonicadvance1/fix_48bit_va_stack
Allocator: Expand stack space when stealing virtual address space
2022-10-31 13:11:20 -07:00
Ryan Houdek 64eb87e9b5 Merge pull request #2099 from Sonicadvance1/fix_infinite_loop
FEXServer: Be robust against invalid packets.
2022-10-31 13:11:13 -07:00
Ryan Houdek aa5e92bee2 Merge pull request #2083 from Sonicadvance1/fix_x87_flag_range
X87: Claim incoming float was in the range for trancendental ops
2022-10-31 13:10:42 -07:00
Ryan Houdek 0bf79dc5d6 unittests/gvisor: Adds a bunch of tests to flakes
These are getting annoying.
2022-10-31 12:53:04 -07:00
Ryan Houdek adb2171c0a Thunks/libX11: Fix recursive initialize
Fixes a crash that occurs due to `_XInitDisplayLock` due to the display
lock function being initialized to our own handler.

Once XInitThreads is called once then it becomes a no-op.

steamwebhelper was hitting this.
2022-10-31 12:38:12 -07:00
Ryan Houdek d6f8923f86 X87: Claim incoming float was in the range for trancendental ops
We don't detect the range of the long F80, so we need to set that the
source was in range to fix sin/cos/tan calculations.

If we don't set this flag to zero then glibc will do some additional
operations that causes the value to be incorrect.

Fixes the output of the test application in #2021, probably fixes some
camera orientation problems in games as well.
2022-10-31 12:36:49 -07:00
Ryan Houdek cf91ab9d5f Merge pull request #2123 from Sonicadvance1/fix_32bit_vdso
32bit: Fixes Debug build of VDSO
2022-10-31 12:17:10 -07:00
Ryan Houdek a0fb9531db FEXServer: Be robust against invalid packets.
Chrome seems to like sending us invalid packets of data sometimes. With
an invalid packet type just skip parsing the data entirely.

Fixes an infinite loop in Vampire Survivors.
2022-10-31 12:05:48 -07:00
Ryan Houdek eca9353b28 Merge pull request #2122 from Sonicadvance1/fix_rotate_right_of
OpcodeDispatcher: Fixes ROR imm OF calculation
2022-10-31 11:52:49 -07:00
Ryan Houdek a259730639 32bit: Fixes Debug build of VDSO
This was generating GOT prologues even on naked functions which was
breaking VDSO on 32-bit.

Fixes almost every 32-bit application when running with debug options.
2022-10-31 11:52:17 -07:00
Ryan Houdek 2e93d10eba OpcodeDispatcher: Fixes ROR imm OF calculation
Turns out this was calculating OF incorrectly, breaking Denuvo early in
its execution.

Changes the ROL imm OF calculation code as well to be more consistent
and not keep src1 alive longer than it needs to be.

Also adds two new unit tests to ensure this stays correct.
2022-10-31 10:28:47 -07:00
Mai 70a3ceb64e Merge pull request #2096 from Sonicadvance1/cleanup_64allocator
Utils/64BitAllocator: Minor cleanups and optimization for munmap
2022-10-31 16:47:53 +00:00
Mai b726f60afd Merge pull request #2098 from Sonicadvance1/fprem_tests
unittests/asm: Adds more extensive FPREM/FPREM1 tests
2022-10-31 16:38:42 +00:00
Mai 2fa1a64999 Merge pull request #2120 from Sonicadvance1/fix_proton_experimental_48bit
ELFCodeLoader: Fixes Proton Experimental on 48-bit VA systems
2022-10-31 16:38:07 +00:00
Ryan Houdek a42b659af9 ELFCodeLoader: Fixes Proton Experimental on 48-bit VA systems
This is a tricky situation that wine-preloader allocates the lower
32MB of stack space through fixed address mmap with MAP_FIXED.

They can't use mmap with an address hint nor MAP_FIXED_NOREPLACE because
it changes behaviour. mmap won't give you the allocation inside the
stack space even if you check `/proc/self/maps` that space isn't yet
allocated. The growable space of the stack blocks those allocations.

So the wine peeps might be SOL if they actually require this allocation
to exist.

To replicate this, allocate the application stack at the same location using an address hint.
This will give us the correct region on a 48-bit VA system, while also
letting it select a different region on a 36-bit VA system.
2022-10-28 02:30:18 -07:00
Ryan Houdek 004c3230a4 Merge pull request #2108 from Sonicadvance1/implement_thunk_disables
Thunks: Add support for disabling thunks in config
2022-10-26 23:38:55 -07:00
Ryan Houdek 2332c41510 Merge pull request #2119 from lioncash/tbl
IR: Handle 256-bit VTBL1
2022-10-26 20:44:23 -07:00
lioncash ec3039c5a2 IR: Handle 256-bit VTBL1
Extends VTBL1 to handle 256-bit vectors.
2022-10-27 00:35:45 +00:00
Ryan Houdek 639d6e6071 Merge pull request #2118 from lioncash/prfx
Arm64/VectorOps: Make use of MOVPRFX where applicable
2022-10-26 15:25:51 -07:00
lioncash cd518d4726 Arm64/VectorOps: Make use of MOVPRFX where applicable
Allows hardware to pack the move and following destructive operation
together into one constructive operation if possible.

e.g.

movprfx VTMP1.D, VectorLower.D
addp VTMP1.B, Pred, VTMP1.B, VectorUpper.B

is allowed to be merged as if it executed constructively like:

addp VTMP1.B, Pred, VectorLower.B, VectorUpper.B

if the hardware supports it. If it doesn't, then the instructions will
behave like a regular move and destructive addp operation separately.
2022-10-26 21:38:48 +00:00
Ryan Houdek b7d9c00dff Merge pull request #2117 from lioncash/ins
IR: Handle 256-bit VInsGPR
2022-10-26 13:22:35 -07:00
lioncash 4b17575f5a IR: Handle 256-bit VInsGPR
Extends VInsGPR to handle 256-bit vectors.
2022-10-26 19:45:27 +00:00
Ryan Houdek 5ba4bba138 Thunks: Add support for disabling thunks in config
Previously the config options could only have ever enabled thunks rather than
disable them.

Now sort the code so it can enable thunks, then following configs can
redisable them.  Allowing testing with global thunks enabled and
disabling problematic applications.

Also sorts the "ThunkConfigFile" config as lower priority than the
application configs. I wasn't thinking about ordering that hard for
these five configuration paths, but application configs should be higher
priority in this case.
2022-10-26 11:43:16 -07:00
Ryan Houdek b3ee5dba0f Merge pull request #2116 from lioncash/extract
IR: Handle 256-bit VExtractToGPR
2022-10-26 11:16:45 -07:00
lioncash d87ff5afa9 IR: Handle 256-bit VExtractToGPR
Extends VExtractToGPR to handle 256-bit vectors.
2022-10-26 17:43:31 +00:00
Ryan Houdek 62a24bd38f Merge pull request #2075 from Sonicadvance1/gpuvis_profiler
FEXCore: Adds support for a timeline profiler interface
2022-10-26 08:44:54 -07:00
Ryan Houdek 8d373c15b8 Merge pull request #2107 from Sonicadvance1/sort_and_upgrade_x11_thunk
Thunks/X11: Reorder and sort X11 interface by headers included.
2022-10-26 04:53:48 -07:00
Ryan Houdek 671f3e74a4 Merge pull request #2103 from Sonicadvance1/sse2_for_guest
Thunks/Guest: Enable SSE2 on thunks and set fpmath to sse
2022-10-26 04:51:35 -07:00
Ryan Houdek 7e810233d9 Merge pull request #2112 from lioncash/ftoi
IR: Handle 256-bit Vector_FToI
2022-10-26 00:07:04 -07:00
Ryan Houdek 4700dbd676 Thunks/Guest: Enable SSE2 on thunks and set fpmath to sse
Clang thunks already have these default enabled, but let's also enable
this on the GCC side.

sse2 will enable most things we care about, which matches ASIMD quite
closely.
fpmath=sse removes some x87 usage for 32-bit thunks specifically.

Should effectively be a non-functional-change
2022-10-26 00:05:16 -07:00
Ryan Houdek 74e18f4317 Merge pull request #2114 from lioncash/vec
Arm64/BranchOps: Remove unused std::vector
2022-10-25 22:02:09 -07:00
lioncash 1eea95cf18 Arm64/BranchOps: Remove unused std::vector
Removes a heap allocation for inline syscalls.
2022-10-26 04:34:57 +00:00
Ryan Houdek 7291b10727 Merge pull request #2113 from lioncash/scvtf
IR: Check for invalid conversion masks in Float_FromGPR_S
2022-10-25 21:32:43 -07:00
lioncash 6804916697 IR: Check for invalid conversion masks in Float_FromGPR_S
Previously this would silently ignore unhandled masks.
2022-10-26 03:59:25 +00:00
lioncash 819e61bf14 IR: Handle 256-bit Vector_FToI
Expands Vector_FToI to handle 256-bit vectors.
2022-10-26 03:42:23 +00:00
Ryan Houdek b8f7e4c8ec Merge pull request #2111 from lioncash/ftof
IR: Handle 256-bit Vector_FtoF
2022-10-25 20:17:26 -07:00
lioncash 17bcc0eed4 IR: Handle 256-bit Vector_FtoF
Extends Vector_FtoF to handle 256-bit vectors.
2022-10-26 02:58:06 +00:00
Ryan Houdek 13003da289 Merge pull request #2110 from lioncash/ftozs
IR: Handle 256-bit Vector_FToZS/Vector_FToS
2022-10-25 19:10:45 -07:00
lioncash 9273538955 IR: Handle 256-bit Vector_FToS
Extends Vector_FToS to handle 256-bit vectors.
2022-10-26 00:59:11 +00:00
lioncash 9750189def IR: Handle 256-bit Vector_FToZS
Extends Vector_FToZS to handle 256-bit vectors.
2022-10-26 00:53:53 +00:00
Ryan Houdek cb17ee9871 Merge pull request #2109 from lioncash/stof
IR: Handle 256-bit Vector_SToF
2022-10-25 17:14:16 -07:00
lioncash 4c3b78ba9a IR: Handle 256-bit Vector_SToF
Extends Vector_SToF to handle 256-bit vectors.
2022-10-25 23:54:41 +00:00
Ryan Houdek e00b6a401b Thunks/X11: Reorder and sort X11 interface by headers included.
Each one of these are sorted through the DefinitionExtracy.py script
running over a temporary header file for each set of includes.

eg:
```bash
$ cat test.h
 #include <X11/Xproto.h>
 #include <X11/XKBlib.h>
 #include <X11/Xlib.h>
 #include <X11/Xutil.h>
 #include <X11/Xresource.h>

 #include <X11/ImUtil.h>
$ ./Scripts/DefinitionExtract.h test.h > out.txt
```

Any custom defined types have been sorted appropriately.
A bunch of missing XKB definitions were missing and added in the
process.
I've had this stashed in my git stash for a while now, I just haven't
cleaned it up.

Fixes a bunch of thunks around X11 applications missing symbols.
2022-10-25 15:43:02 -07:00
Ryan Houdek ac0ab8a7b4 Thunks/X11: Ensure 11 headers are included with C linkage
Otherwise the compiler gets confused about some functions getting
declared with C++ linkage.
2022-10-25 15:32:40 -07:00
Ryan Houdek 0aff3941f4 Scripts/DefinitionExtract: Fixes some more function attributes
X11 has an attribute that was causing function declarations to be
missed.

These definitions exist in XLibint.h
eg:
```cpp
extern void _XEatData(
    Display*		/* dpy */,
    unsigned long	/* n */
) _X_COLD;
```

This `_X_COLD` attribute was causing these function definitions to get
missed.
2022-10-25 15:32:40 -07:00
Ryan Houdek 27b022d4d9 Merge pull request #2106 from lioncash/dup
IR: Handle 256-bit VDupElement
2022-10-25 13:49:29 -07:00
lioncash e188928742 IR: Handle 256-bit VDupElement
Extends VDupElement to handle 256-bit vectors.
2022-10-25 20:01:53 +00:00
Ryan Houdek 780e3c7fb7 Merge pull request #2105 from lioncash/unzip
IR: Handle 256-bit VUnZip/VUnZip2
2022-10-25 12:45:05 -07:00
lioncash 1b5146d3ac IR: Handle 256-bit VUnZip2
Extends VUnZip2 to handle 256-bit vectors.
2022-10-25 16:01:55 +00:00
lioncash 80cf3ca6b9 IR: Handle 256-bit VUnZip
Extends VUnZip to handle 256-bit vectors.
2022-10-25 15:49:02 +00:00
Ryan Houdek 2272b30a91 Merge pull request #2101 from Sonicadvance1/fix_thunk_versions
Thunks: Fixes missing thunk librarie so versions
2022-10-24 23:11:54 -07:00
Ryan Houdek b5fb1cb07c Merge pull request #2100 from Sonicadvance1/fix_thunk_loaded_check
ThunksDB: Fixes Thunks loaded boolean pointer check
2022-10-24 23:11:33 -07:00
Ryan Houdek 4ea34a9c22 Thunks: Fixes missing thunk librarie so versions
Some libraries were missing these version defines, which was causing
dlopen to fail.

This was causing thunks to break in pressure-vessel.
2022-10-24 20:59:54 -07:00
Ryan Houdek 48e7de9f9e ThunksDB: Fixes Thunks loaded boolean pointer check
Need to dereference the boolean to ensure we only load the thunksDB
files once.
2022-10-24 20:55:58 -07:00
Ryan Houdek 76dd2369a7 unittests/asm: Adds more extensive FPREM/FPREM1 tests
unit tests that show the difference of output between FPREM and FPREM1.
Setup as known failures on everything except for host since we don't
implement fprem correctly.

An incorrect fix to FPREM is as follows:
```diff
--- a/External/FEXCore/Source/Common/SoftFloat.h
+++ b/External/FEXCore/Source/Common/SoftFloat.h
@@ -158,6 +158,10 @@ struct X80SoftFloat {

     return Result;
 #else
+    BIGFLOAT lhs_ = lhs;
+    BIGFLOAT rhs_ = rhs;
+    BIGFLOAT Result = fmodl(lhs_, rhs_);
+    return Result;
     return extF80_rem(lhs, rhs);
 #endif
   }
```

But we shouldn't implement this fix. We should instead implement a new `extF80_mod`
function that handles the rounding differences between FPREM and FPREM1.

Fixes #2097.
Doesn't attempt to resolve #1538
2022-10-22 20:47:24 -07:00
Ryan Houdek 78e0cd6e77 Scripts: Updates testharness_runner to support runner specific known failures 2022-10-22 20:29:51 -07:00
Ryan Houdek 5514a04cb4 Utils/64BitAllocator: Minor cleanups and optimization for munmap
- Some minor cleanups in the VMARegion struct type.
- Switches over to using the FlexBitSet range scanning, based off this
implementation.
- Move memory region allocation to its own function instead of
  constructor
  - This will be used by a new constructor later for 48-bit host-side
    allocations
- Minor optimization to keep track of Munmap.
  - We were burning a bunch of time on backward scanning for free
    regions even though we never did a munmap to free anything.
  - Now only do backward scanning if a munmap occured.
  - Saves a bunch of CPU time
2022-10-21 21:07:09 -07:00
Ryan Houdek 99ca78b235 Utils/FlexBitSet: Adds range scanning functions
These were currently living in the 64BitAllocator class but can be moved
directly to the FlexBitSet.

Ideally in the future these routines can be optimized so our allocator
is faster but for now these are just moved.
2022-10-21 20:58:26 -07:00
Ryan Houdek ab45db1665 Merge pull request #2094 from lioncash/zip
IR: Handle 256-bit VZip/VZip2
2022-10-20 19:02:29 -07:00
Ryan Houdek bb38bcb67d Merge pull request #2093 from lioncash/shrn
IR: Handle 256-bit VUShrNI/VUShrNI2
2022-10-20 19:00:51 -07:00
Ryan Houdek 6cc2912542 Merge pull request #2092 from lioncash/vsqxtun
IR: Handle 256-bit VSQXTUN/VSQXTUN2
2022-10-20 18:57:24 -07:00
lioncash 340b2ca624 IR: Handle 256-bit VZip2
Extends VZip2 to handle 256-bit vectors.
2022-10-20 21:15:18 +00:00
lioncash 5baa15de03 IR: Handle 256-bit VZip
Extends VZip to handle 256-bit vectors.
2022-10-20 19:22:07 +00:00
lioncash 6ddca804d1 IR: Handle 256-bit VUShrNI2
Extends VUShrNI2 to handle 256-bit vectors.
2022-10-20 18:31:36 +00:00
lioncash f9831a85fb IR: Handle 256-bit VUShrNI
Extends VUShrNI to handle 256-bit vectors.
2022-10-20 17:51:47 +00:00
lioncash 7261033b7f IR: Handle 256-bit VSQXTUN2
Extends VSQXTUN2 to handle 256-bit vectors.
2022-10-20 17:08:00 +00:00
lioncash 3ad6866198 IR: Handle 256-bit VSQXTUN
Extends VSQXTUN to handle 256-bit vectors.
2022-10-20 16:51:53 +00:00
Ryan Houdek 1c7d4165ab Merge pull request #2091 from lioncash/vsqxtn
IR: Handle 256-bit VSQXTN/VSQXTN2
2022-10-19 20:46:53 -07:00
Ryan Houdek 3e48b1a8ac FEXCore: Adds support for a timeline profiler interface
This creates a generic interface that FEXCore can use for timeline
profiling. This allows us to create a generic interface which the
backend details are hidden so we can support multiple timeline profile
APIs.

The only API supported right now is ftrace/gpuvis. Which is extremely
lightweight of an interface with minimal overhead.

We must be careful here since in most cases will will have dozens of
FEX instances running at any given time. So a timeline profiler like
Microprofiler can have major issues since that only ever expects a
single process at a time.

Not enabled by default but just needs the `ENABLE_FEXCORE_PROFILER`
cmake option set to enable.
2022-10-19 19:56:35 -07:00
lioncash 07be100daf IR: Handle 256-bit VSQXTN2
Extends VSQXTN2 to handle 256-bit vectors.
2022-10-20 02:41:45 +00:00
lioncash de9351eefb IR: Handle 256-bit VSQXTN
Expands VSQXTN to handle 256-bit vectors.
2022-10-20 02:04:05 +00:00
Ryan Houdek 2c44b5b3a1 Allocator: Expand stack space when stealing virtual address space
If we take all of the stack space then the auto expanding stack doesn't
work and we get stuck with a small stack that breaks thunks.
2022-10-19 19:02:41 -07:00
Ryan Houdek 136f1e2fc7 Merge pull request #2090 from lioncash/sxtl
Arm64/VectorOps: Simplify SVE VSXTL/VSXTL2/VUXTL/VUXTL2 implementations
2022-10-19 17:02:17 -07:00
lioncash ad39add55f Arm64/VectorOps: Simplify VUXTL2 SVE implementation
Turns out there's an instruction that does what we need, but isn't named
similarly to UXTL2 at all.
2022-10-19 22:24:07 +00:00
lioncash f3c301e359 Arm64/VectorOps: Simplify VUXTL SVE implementation
Turns out there's an instruction that does what we need, but has a name
not similar to UXTL
2022-10-19 22:22:23 +00:00
lioncash 1c37a1b4d6 Arm64/VectorOps: Simplify VSXTL2 SVE implementation
Turns out there's a built-in instruction that does exactly what we want,
but just has a different name from SXTL2
2022-10-19 22:16:33 +00:00
lioncash 7222529904 Arm64/VectorOps: Simplify VSXTL SVE implementation
Was reading the ARM ARM and realized there's an instruction that does
exactly what we need right out of the box.
2022-10-19 22:12:22 +00:00
Ryan Houdek f26eccd00f Merge pull request #2089 from wannacu/main
Implements DAA, DAS, AAA, AAS, AAM and AAD instruction
2022-10-19 03:57:13 -07:00
wannacu 73375a76ac unittests: Adds DAA, DAS, AAA, AAS, AAM and AAD unit test 2022-10-19 13:56:21 +08:00
wannacu d4416d200e OpcodeDispatcher: Implements DAA, DAS, AAA, AAS, AAM and AAD instruction 2022-10-19 13:56:06 +08:00
Mai d1b235dd83 Merge pull request #2080 from Sonicadvance1/fix_64bit_syscall_mman
Syscalls: Fixes 64-bit mmap and munmap
2022-10-19 01:26:15 +00:00
Ryan Houdek 3ac5e0423a Merge pull request #2088 from lioncash/vsmull
IR: Handle 256-bit VSMull/VSMull2
2022-10-18 16:11:03 -07:00
Ryan Houdek fc6de5f3c0 Merge pull request #2087 from lioncash/vixl-narrow
External: Update vixl submodule
2022-10-18 16:08:31 -07:00
Mai b1e475d81d Merge pull request #2081 from Sonicadvance1/fix_rotate_flags
OpcodeDispatcher: Fixes flag calculation on ROR and ROL by immediate
2022-10-18 22:33:47 +00:00
Mai 4a09a4324f Merge pull request #2082 from Sonicadvance1/fix_c2_fprem1
OpcodeDispatcher: Fixes FPREM1 C2 flag calculation
2022-10-18 22:33:29 +00:00
lioncash 47f94327c5 IR: Amend x86_64 32->64 case for VUMull2
Realized I forgot to amend the registers used in the final multiply.
2022-10-18 16:23:16 +00:00
lioncash a009ed0b6b IR: Handle 256-bit VSMull2
Extends VSMull2 to handle 256-bit vectors.
2022-10-18 16:23:14 +00:00
lioncash 2476a686e7 IR: Handle 256-bit VSMull
Extends VSMull to handle 256-bit vectors.
2022-10-18 16:22:44 +00:00
lioncash f0db93773f unittests: Re-enable narrowing and widening tests
Now that the bug in vixl's simulator is fixed, we can enable these tests
again.
2022-10-18 15:18:14 +00:00
lioncash fabe824c8b Externals: Update vixl submodule
Includes fixes for the narrowing instructions.
2022-10-18 15:15:54 +00:00
Ryan Houdek 78a077397e Merge pull request #2085 from lioncash/vmull
IR: Handle 256-bit VUMull/VUMull2
2022-10-17 18:10:48 -07:00
lioncash 0c4b456aaa IR: Handle 256-bit VUMull2
Extends VUMull2 to handle 256-bit vectors.
2022-10-17 19:03:26 +00:00
lioncash 09185167bc OpcodeDispatcher/Vector: Amend and simplify PMULLOp
Allows PMULLOp to function correctly with the amended VPSHUFD entries.

Since this is only used to perform expanded multiplication from 32-bit
entries to 64-bit entries, we can simplify things a little bit.

All we need to do is yank the third 32-bit word down into the second
32-bit word's spot in the vector and let the VUMull/VSMull IR ops handle
it.
2022-10-17 19:03:26 +00:00
lioncash 5d78c3203c IR: Handle 256-bit VUMull
Extends VUMull to handle 256-bit vectors.

While we're at it, we can fix a typo in the VPSHUFD called for the
32->64-bit case.

To mirror UMULL, we need to replicate element 0 and 1, not 0 and 2

While we're at it, we can fix this with VSMull as well.
2022-10-17 19:03:23 +00:00
Ryan Houdek fa5322d3f9 Merge pull request #2084 from Sonicadvance1/more_auxv
ELFCodeLoader: Implement four more auxv values
2022-10-17 09:48:18 -07:00
Ryan Houdek d21aa5cac2 ELFCodeLoader: Implement four more auxv values
Implements AT_PLATFORM: Ends up being `i686` or `x86_64` depending on
ELF arch

Implements AT_HWCAP and AT_HWCAP2
AT_HWCAP is just CPUID function 01h EDX result
AT_HWCAP2 only has two defined bits in it, which we don't support
either.

Implements AT_RANDOM
Previously we were just sticking hardcoded values in to this.
Now we pass along the host's AT_RANDOM, or we generate our own if that
doesn't exist

Fixes #788
2022-10-16 19:45:19 -07:00
Ryan Houdek 102d5c57cb OpcodeDispatcher: Fixes FPREM1 C2 flag calculation
Accidentally didn't implement this for FPREM1 but it /was/ implemented
for FPREM. Fixes an infinite loop in cossin implementations.

Test code from the application returns an incorrect result, but it isn't
due to FPREM1.

```
$ `which wine` ./hello.exe
Sin 1.22460635382238E-16
Cos -1
$ FEXInterpreter `which wine` ./hello.exe
Sin 1.22460635382238E-16
Cos 0.54030230586814
```

Fixes #2021
2022-10-16 15:24:46 -07:00
Ryan Houdek ffb4de9fd9 Merge pull request #2079 from Sonicadvance1/ensure_armemitter_uses_allocator
Ensure Arm64Emitter uses FEX allocator
2022-10-15 15:37:23 -07:00
Ryan Houdek 6b3d8886e5 Merge pull request #2077 from Sonicadvance1/fix_thunks_with_lots_args
Thunks: Fixes indirect thunks with 8+ arguments
2022-10-15 15:37:09 -07:00
Ryan Houdek ddc10272a0 Merge pull request #2076 from Sonicadvance1/update_vulkan
Thunks: Update Vulkan thunk to v1.3.231
2022-10-15 15:14:30 -07:00
Ryan Houdek 0b5ef00165 Thunks: Fixes indirect thunks with 8+ arguments
Due to how we use a modified ABI for these indirect functions, we don't
have a clean way to say that the host_addr lives in a side-argument.

The previous inline asm that moved the value from r11 in to a variable
worked up until you hit functions with 8 or more arguments. At that
point the compiler was generating code before our inline assembly and
using r11 as a temporary, thus destroying our value.
Then a crash would occur and it was very hard to determine why. It would
end up calling some random function (0x1 in this case) from an indirect
call.

This made it /look/ like it was calling an invalid function returned
from the loader but in reality it was a corrupt register loading bad
data.

To work around this case, we can use an inline asm register variable and
a volatile asm block that "sets" the variable. In this case GCC and
Clang both seem to extend the live range of the register from the start
of the function to the use of the variable.

This resolves the issue for now, and I tested quite a large number of
function signatures to see if it would break in the future.

Theoretically our functional testing should catch this, but we don't
currently have something that abuses all the functions like this
currently.
2022-10-15 15:13:40 -07:00
Ryan Houdek 2a50416fc3 unittests: Ensures overloaded shifts don't result in JIT failure 2022-10-14 23:38:35 -07:00
Ryan Houdek ce514d9f83 unittests: Adds ROL and ROR CF flag calculation tests
This would have failed prior to the last commit
2022-10-14 23:37:46 -07:00
Ryan Houdek 9b77e7fd13 OpcodeDispatcher: Fixes flag calculation on ROR and ROL by immediate
These were being calculated incorrectly in the case of rotating with
values larger than 8-bit or 16-bit
2022-10-14 23:36:48 -07:00
Ryan Houdek abb44d3327 Merge pull request #2069 from wannacu/main
Flags: Refine _Bfe's shift
2022-10-14 22:53:05 -07:00
Ryan Houdek 11eaf3d48a Ensure Arm64Emitter uses FEX allocator
Otherwise we will end up allocating code buffers in the lower 32-bits,
consuming precious virtual address space.
2022-10-14 22:04:07 -07:00
Ryan Houdek 76c2cc2c3e Syscalls: Fixes 64-bit mmap and munmap
These should be using the real syscalls, not our provided allocators.

While not a problem currently since these redirect to host mmap and
munmap, it will become an issue once we have an allocator that lives
outside of x86-64 space.
2022-10-14 21:43:04 -07:00
Ryan Houdek 0e6c8bd12e Thunks: Update Vulkan thunk to v1.3.231
Only missing a few function definitions, resorted to match order of
definitions in the headers so future changes don't mix up as much
2022-10-14 01:51:35 -07:00
Ryan Houdek a9fb008317 External: Update Vulkan-Headers to v1.3.231 2022-10-14 01:50:49 -07:00
Ryan Houdek e9f3a5b3e4 Merge pull request #2074 from lioncash/vuxtl
IR: Handle 256-bit VUXTL/VUXTL2
2022-10-13 13:15:47 -07:00
Ryan Houdek ebc45dff45 Merge pull request #2070 from lioncash/vuabdl
IR: Handle 256-bit VUABDL
2022-10-13 12:17:47 -07:00
lioncash fc4a5ebfd3 IR: Handle 256-bit VUXTL2
Extends VUXTL2 to handle 256-bit vectors.
2022-10-13 19:15:05 +00:00
lioncash 1d7b688c55 IR: Handle 256-bit VUXTL
Extends VUXTL to handle 256-bit values.
2022-10-13 19:07:57 +00:00
Ryan Houdek f14a5ffbbf Merge pull request #2073 from lioncash/vsxtl
IR: Handle 256-bit VSXTL/VSXTL2
2022-10-13 11:53:12 -07:00
Ryan Houdek cada0d593c Merge pull request #2072 from lioncash/test
unittests: Amend mm register usage in H0F38/66_04.asm test
2022-10-13 11:28:27 -07:00
lioncash 0436540791 IR: Handle 256-bit VSXTL2
Extends VSXTL2 to handle 256-bit vectors.
2022-10-13 18:21:25 +00:00
lioncash a87ac86e18 IR: Handle 256-bit VSXTL
Extends VSXTL to handle 256-bit vectors.
2022-10-13 18:21:22 +00:00
lioncash 1278b23150 unittests: Amend mm register usage in H0F38/66_04.asm test
This should be using xmm2 rather than mm2.
2022-10-13 17:09:40 +00:00
lioncash 02f5ea4b9d IR: Handle 256-bit VUABDL
Extends VUABDL to handle 256-bit vectors.
2022-10-13 15:37:10 +00:00
wannacu 2e14e613d0 Flags: Refine _Bfe's shift 2022-10-13 16:41:04 +08:00
Ryan Houdek 85c2889652 Docs: Update for release FEX-2210 2022-10-13 00:46:59 -07:00
Ryan Houdek 23dd056b60 Merge pull request #2067 from lioncash/vmul
IR: Handle 256-bit VSMul/VUMul
2022-10-12 15:59:03 -07:00
lioncash 71043e372a IR: Handle 256-bit VSMul/VUMul
Extends VSMul and VUMul to handle 256-bit vectors.
2022-10-12 00:00:15 +00:00
Ryan Houdek c412d073b9 Merge pull request #2066 from lioncash/vrev64
IR: Handle 256-bit VRev64
2022-10-11 14:21:45 -07:00
lioncash 8fb03ff1b9 IR: Handle 256-bit VRev64
Extends VRev64 to handle 256-bit vectors.
2022-10-11 20:17:04 +00:00
Ryan Houdek c2b6aef6f4 Merge pull request #2065 from lioncash/shift-imm
IR: Handle 256-bit VShlI/VUShlI/VUShrI
2022-10-11 12:16:44 -07:00
lioncash 24547318c6 IR: Handle 256-bit VShlI
Extends VShlI to handle 256-bit vectors.
2022-10-11 18:19:23 +00:00
Ryan Houdek b693112c80 Merge pull request #2058 from Sonicadvance1/add_opencl_thunk_db
Add opencl thunk db
2022-10-11 11:06:20 -07:00
lioncash 48d1184066 IR: Handle 256-bit VSShrI
Extends VSShrI to handle 256-bit vectors.
2022-10-11 18:04:13 +00:00
Ryan Houdek 8da9ebc2e0 Merge pull request #2062 from wannacu/main
SMC: Fix possible deadlock
2022-10-11 10:13:02 -07:00
lioncash 5ba510474b IR: Handle 256-bit VUShrI
Extends VUShrI to handle 256-bit vectors.
2022-10-11 17:08:35 +00:00
Ryan Houdek 51214d1be1 Merge pull request #2064 from lioncash/vushls
IR: Handle 256-bit VSShrS/VUShlS/VUShrS
2022-10-11 09:59:03 -07:00
Ryan Houdek 4d6e15d7af Merge pull request #2063 from lioncash/interp-shift
Interpreter: Handle 256-bit VSShr/VUShl/VUShr
2022-10-11 09:22:11 -07:00
lioncash 4721894427 IR: Handle 256-bit VSShrS
Extends VSShrS to handle 256-bit vectors.
2022-10-11 16:21:50 +00:00
lioncash d429865b6e IR: Handle 256-bit VUShrS
Extends VUShrS to handle 256-bit vectors.
2022-10-11 16:08:09 +00:00
lioncash ca5881a72c IR: Handle 256-bit VUShlS
Extends VUShlS to handle 256-bit vectors.
2022-10-11 15:45:39 +00:00
lioncash 7151b9daff Interpreter/VectorOps: Remove lingering magic 32 constants
Makes these functions consistent with the rest that explicitly test for
256 bit width.
2022-10-11 15:02:41 +00:00
lioncash d9b5e28b22 Interpreter: Handle 256-bit VSShr
This is only implemented in the interpreter, so this is trivial.
2022-10-11 14:58:40 +00:00
lioncash aa7954a7d6 Interpreter: Handle 256-bit VUShr
This is only implemented in the interpreter, so this is trivial.
2022-10-11 14:56:19 +00:00
lioncash 802c70d1ab Interpreter: Handle 256-bit VUShl
This is only implemented in the interpreter, so this is trivial.
2022-10-11 14:55:17 +00:00
wannacu 8f905988e9 Use compatible syscall helpers 2022-10-11 16:35:17 +08:00
wannacu 478c5595ad SMC: Fix possible deadlock 2022-10-11 15:31:27 +08:00
Ryan Houdek 3977e1f29e Merge pull request #2055 from Sonicadvance1/update_description_ripping_script
Scripts: Updates DefinitionExtract
2022-10-10 10:11:16 -07:00
Ryan Houdek 2b1ef97354 Merge pull request #2060 from Sonicadvance1/clang_thunks
Thunks: Add support for building with clang
2022-10-10 09:44:58 -07:00
Mai eaddf7f1a5 Merge pull request #2061 from Sonicadvance1/fix_linker_script_depends
Thunks: Adds dependency on linker script
2022-10-10 12:17:20 -04:00
Ryan Houdek 3237de3085 Merge pull request #2056 from Sonicadvance1/guest_function_bool
Thunks/Host: Adds bool operator to fex_guest_function_ptr
2022-10-10 09:10:34 -07:00
Ryan Houdek df3d398d31 Thunks/Host: Adds bool operator to fex_guest_function_ptr
Lets us check if nullptr was passed in
2022-10-10 08:52:50 -07:00
Ryan Houdek b44b3401b7 Merge pull request #2015 from Sonicadvance1/map_regular_offset
ELFCodeLoader: Map primary ELF more like the kernel
2022-10-10 08:49:49 -07:00
Ryan Houdek c28ca0fac9 ELFCodeLoader: Map primary ELF more like the kernel
The kernel maps the primary ELF with a hint to some place *near* the
middle of the virtual address space. While the interpreter stays at the
top of the address space.

This also adds ASLR to the ELF loading, with a define for debugging and
testing purposes.

Requires both #2013 and #2014 merged first.
2022-10-10 08:38:15 -07:00
Ryan Houdek 235e2b6c2c FEXCore/Allocator: Store what the host VA is
Calling this function multiple times without this will change the
result.

Necessary so we can determine what the host VA is  from multiple
locations.
2022-10-10 08:35:40 -07:00
Mai d68b84bc27 Merge pull request #2013 from Sonicadvance1/fix_mapper
ELFCodeloader: Map once and then use MAP_FIXED to overwrite
2022-10-10 10:05:53 -04:00
Mai edca528608 Merge pull request #2039 from Sonicadvance1/fix_dynamic_non_interpreter_elfs
ELFCodeLoader: Fixes dynamic non-interpreter ELFs
2022-10-10 10:04:19 -04:00
Ryan Houdek b75e8f2abf Thunks: Add support for building with clang
Fairly straightforward, just requires enabling lld in this case since
cross-compiling doesn't work well with gnu linker.

Also lld doesn't understand the linker script program header symbolic
names for read/write/execute. So we need to use the raw number there.

Works around an issue where GCC 11 generates broken `init_array` section
and also plt sections that glibc doesn't understand.
2022-10-09 23:07:30 -07:00
Ryan Houdek ec3158e4cd Thunks: Adds dependency on linker script
Ensures the thunk is rebuilt if the linker scripts have changed.

Fixes #2054
2022-10-09 22:57:52 -07:00
Ryan Houdek 9c8c8041e0 ThunksDB: Adds OpenCL to the json
This will be used soon
2022-10-09 19:52:36 -07:00
Ryan Houdek e3adaacb51 Thunks: Adds another packed arguments template
This will be used soon
2022-10-09 19:47:37 -07:00
Ryan Houdek c49e11484f Scripts: Updates DefinitionExtract
Unused warning attribute wasn't getting ignored.
Also need to update output text to match new format
2022-10-09 19:44:49 -07:00
Ryan Houdek 7b4b9a80fa Merge pull request #2053 from lioncash/vfcmpord
IR: Handle 256-bit VFCMPORD/VFCMPUNO
2022-10-09 19:43:36 -07:00
lioncash c7ad066987 IR: Handle 256-bit VFCMPUNO
Extends VFCMPUNO to handle 256-bit vectors.
2022-10-06 15:43:01 +00:00
lioncash f4d229f1ba IR: Handle 256-bit VFCMPORD
Extends VFCMPORD to handle 256-bit vectors.
2022-10-06 15:33:30 +00:00
Ryan Houdek 25a8a00771 Merge pull request #2050 from lioncash/vfcmplt
IR: Handle 256-bit VFCMPLT/VFCMPGT/VFCMPLE
2022-10-04 14:44:54 -07:00
Ryan Houdek a67f7422b2 Merge pull request #2049 from lioncash/fcmeq
IR: Handle 256-bit VFCMPEQ/VFCMPNEQ
2022-10-04 14:43:52 -07:00
Ryan Houdek ed8150cfb6 Merge pull request #2048 from lioncash/vcmpgt
IR: Handle 256-bit VCMPGT/VCMPGTZ/VCMPLTZ
2022-10-04 14:42:53 -07:00
Ryan Houdek 462a163ba7 Merge pull request #2047 from lioncash/vcmpeq
IR: Handle 256-bit VCMPEQ/VCMPEQZ
2022-10-04 14:40:48 -07:00
Ryan Houdek 6374175a64 Merge pull request #2046 from lioncash/vbsl
IR: Handle 256-bit VBSL
2022-10-04 14:37:25 -07:00
lioncash 280b15ba2a IR: Handle 256-bit VFCMPLE
Extends VFCMPLE to handle 256-bit vectors.
2022-10-04 19:56:59 +00:00
lioncash 2a0b488e99 IR: Handle 256-bit VFCMPGT
Extends VFCMPGT to handle 256-bit vectors.
2022-10-04 19:43:32 +00:00
lioncash 3ef7c4ab51 IR: Handle 256-bit VFCMPLT
Extends VFCMPLT to handle 256-bit vectors.
2022-10-04 19:29:36 +00:00
lioncash e72d746036 IR: Handle 256-bit VFCMPNEQ
Extends VFCMPNEQ to handle 256-bit vectors.
2022-10-04 19:02:20 +00:00
lioncash 0bb4091e34 IR: Handle 256-bit VFCMPEQ
Extends VFCMPEQ to handle 256-bit vectors.
2022-10-04 18:44:12 +00:00
lioncash 6822fc595c IR: Handle 256-bit VCMPLTZ
Extends VCMPLTZ to handle 256-bit vectors.
2022-10-04 18:09:39 +00:00
lioncash a506a589dd IR: Handle 256-bit VCMPGTZ
Extends VCMPGTZ to handle 256-bit vectors.
2022-10-04 17:55:03 +00:00
lioncash 684a5977dd IR: Handle 256-bit VCMPGT
Extends VCMPGT to handle 256-bit registers.
2022-10-04 17:38:20 +00:00
lioncash ac3682e058 IR: Handle 256-bit VCMPEQZ
Extends VCMPEQZ to handle 256-bit vectors.
2022-10-04 16:57:19 +00:00
lioncash d5faf01f5a IR: Handle 256-bit VCMPEQ
Extends VCMPEQ to handle 256-bit vectors.
2022-10-04 16:33:25 +00:00
lioncash ecba1b6838 IR: Handle 256-bit VBSL
Extends VBSL to handle 256-bit vectors.
2022-10-03 18:06:27 +00:00
Ryan Houdek 8d8b029285 Merge pull request #2044 from lioncash/vumax
IR: Handle 256-bit VSMax/VUMax
2022-09-29 13:52:40 -07:00
lioncash bf6f855868 IR: Handle 256-bit VSMax
Extends VSMax to handle 256-bit vectors.
2022-09-29 20:31:19 +00:00
Ryan Houdek aa6a499329 Merge pull request #2043 from lioncash/vumin
IR: Handle 256-bit VSMin/VUMin
2022-09-29 13:19:07 -07:00
lioncash 0971650ef9 IR: Handle 256-bit VUMax
Extends VUMax to handle 256-bit vectors.
2022-09-29 20:18:42 +00:00
Ryan Houdek 64c4fdccf7 Merge pull request #2042 from lioncash/vnot
IR: Handle 256-bit VNot
2022-09-29 13:13:23 -07:00
Ryan Houdek d715ffbc8e Merge pull request #2041 from lioncash/vfneg
IR: Handle 256-bit VFNeg
2022-09-29 13:12:21 -07:00
lioncash aef801b5b5 IR: Handle 256-bit VSMin
Extends VSMin to handle 256-bit vectors.
2022-09-29 20:01:04 +00:00
lioncash 6c9e29796b IR: Handle 256-bit VUMin
Extends VUMin to handle 256-bit vectors.
2022-09-29 19:47:47 +00:00
lioncash 364bb3ac1e IR: Handle 256-bit VNot
Extends VNot to handle 256-bit vectors.
2022-09-29 14:40:46 +00:00
lioncash 428ea68507 IR: Handle 256-bit VFNeg
Extends VFNeg to handle 256-bit vectors.
2022-09-29 14:00:52 +00:00
Ryan Houdek 5cf59408a7 Merge pull request #2040 from Sonicadvance1/fix_vsyscall
VDSO: Fix vsyscall
2022-09-29 01:47:53 -07:00
Ryan Houdek 1596843015 VDSO: Fix vsyscall
The `mov ebp, ecx` was breaking vsyscall and was expected to be used
with the `syscall` instruction rather than `int 0x80`.
Remove that to fix it.

Also remove the pushes and pops around the syscall instruction, these
are unnecessary in an emulated environment, we won't clobber the
registers.

Fixes Steam execution with VDSO.
2022-09-28 17:34:24 -07:00
Ryan Houdek af6582ff5b ELFCodeLoader: Fixes dynamic non-interpreter ELFs
Specifically fixes /sbin/ldconfig.
Fixes 8df7c2d84f
Fixes Steam launching

I failed to test ELF files that are dynamic with no interpreter here, so
EntryPoint ended up being set to zero, which results in an instant
crash.

Ensure these values are set correctly after the primary ELF and
interpreter are loaded so starting RIP is correct.
2022-09-28 16:00:02 -07:00
Ryan Houdek 1799d4c675 Merge pull request #2038 from lioncash/vneg
IR: Handle 256-bit VNeg
2022-09-28 13:16:55 -07:00
Ryan Houdek 808e1c0330 Merge pull request #2029 from lioncash/interp
Interpreter: Use constant for AVX register size where applicable
2022-09-28 13:16:14 -07:00
Ryan Houdek dacd96cab5 Merge pull request #2037 from lioncash/vfrsqrt
IR: Handle 256-bit VFRSqrt
2022-09-28 13:15:50 -07:00
Ryan Houdek ca4d3bf64d Merge pull request #2036 from lioncash/vfsqrt
IR: Handle 256-bit VFSqrt
2022-09-28 13:14:25 -07:00
Ryan Houdek ea38b043c1 Merge pull request #2035 from lioncash/vfrecp
IR: Handle 256-bit VFRecp
2022-09-28 13:12:57 -07:00
Ryan Houdek a39746df2e Merge pull request #2034 from lioncash/vfmax
IR: Handle 256-bit VFMax
2022-09-28 13:11:04 -07:00
Ryan Houdek 2367a8e50b Merge pull request #2033 from lioncash/vfmin
IR: Handle 256-bit VFMin
2022-09-28 13:10:17 -07:00
Ryan Houdek cb121d7f17 Merge pull request #2032 from lioncash/vaddp
IR: Handle 256-bit VAddP
2022-09-28 13:08:40 -07:00
Ryan Houdek 412793c21d Merge pull request #2030 from lioncash/interp-mov
Interpreter: Handle 256-bit VMov
2022-09-28 13:05:16 -07:00
lioncash ce2286c48b IR: Handle 256-bit VNeg
Extends VNeg to handle 256-bit vectors.
2022-09-28 19:25:33 +00:00
lioncash d5694d6de0 IR: Handle 256-bit VFRSqrt
Extends VFRSqrt to handle 256-bit vectors.
2022-09-28 19:11:01 +00:00
lioncash 121218aa8a IR: Handle 256-bit VFSqrt
Extends VFSqrt to handle 256-bit vectors.
2022-09-28 18:46:43 +00:00
lioncash 6bb53fa758 IR: Handle 256-bit VFRecp
Extends VFRecp to handle 256-bit vectors.
2022-09-28 18:27:34 +00:00
lioncash 89aa0c5471 IR: Handle 256-bit VFMax
Extends VFMax to handle 256-bit vectors.
2022-09-28 17:40:16 +00:00
lioncash 53fcbf6afa IR: Handle 256-bit VFMin
Extends VFMin to handle 256-bit vectors.
2022-09-28 16:56:22 +00:00
lioncash 2f5643ae6b Arm64/VectorOps: Amend half-precision case in VFAddP
Noticed that I forgot to change the Zn register over to VTMP1.
This would have been caught by a vixl internal assert anyway.

Also make the behavior equal with VAddP, where we clear and only copy
over the exact amount of bytes instead of the whole register.
2022-09-28 15:10:49 +00:00
lioncash d162ac8b3d IR: Handle 256-bit VAddP
Extends VAddP to handle 256-bit vectors.
2022-09-28 15:01:31 +00:00
lioncash c9a704fbde Interpreter: Handle 256-bit VMov
Now we'll properly handle the move.

Also put an assert in place to catch any over-sized values.
2022-09-28 13:47:23 +00:00
lioncash 9d5a822a3a Interpreter: Use constant for AVX register size where applicable
Makes the previously implemented ops a little more self-documenting,
and, if we ever actually need to change this, there's a nice constant
that can be looked up instead of magic 32 values.
2022-09-28 13:38:03 +00:00
Ryan Houdek 50eba4066a Merge pull request #2028 from lioncash/vfdiv
IR: Handle 256-bit VFDiv
2022-09-27 13:39:02 -07:00
lioncash 6116ae5330 IR: Handle 256-bit VFDiv
Extends VFDiv to handle 256-bit vectors.
2022-09-27 20:20:31 +00:00
Ryan Houdek 447226576f Merge pull request #2027 from lioncash/vfmul
IR: Handle 256-bit VFMul
2022-09-27 13:18:38 -07:00
Ryan Houdek 3f8b872f17 Merge pull request #2026 from lioncash/vfsub
IR: Handle 256-bit VFSub
2022-09-27 13:17:00 -07:00
Ryan Houdek e573ddc2db Merge pull request #2025 from lioncash/vfaddp
IR: Handle 256-bit VFAddP
2022-09-27 12:56:58 -07:00
lioncash fe9aa681f0 IR: Handle 256-bit VFMul
Extends VFMul to handle 256-bit vectors.
2022-09-27 19:54:57 +00:00
lioncash 1b2f2c1559 IR: Handle 256-bit VFSub
Extends VFSub to be able to handle 256-bit vectors.
2022-09-27 19:33:46 +00:00
lioncash 84c75a86c3 IR: Handle 256-bit VFAddP
Extends VFAddP to be able to handle 256-bit vectors.
2022-09-27 19:03:44 +00:00
Ryan Houdek eedbde6f15 Merge pull request #2024 from lioncash/vfadd
IR: Handle 256-bit VFAdd
2022-09-27 09:53:25 -07:00
Ryan Houdek 4e441e5a08 Merge pull request #2023 from lioncash/vpopcnt
IR: Handle 256-bit VPopcount
2022-09-27 09:51:27 -07:00
Ryan Houdek 3e287a36c2 Merge pull request #2022 from lioncash/vabs
IR: Handle 256-bit VAbs
2022-09-27 09:48:26 -07:00
lioncash eadc477695 IR: Handle 256-bit VFAdd 2022-09-27 16:07:42 +00:00
lioncash 59aa324678 IR: Handle 256-bit VPopcount
Extends VPopcount to be able to handle 256-bit width vectors
2022-09-27 15:19:24 +00:00
lioncash 219bce1467 IR: Handle 256-bit VAbs
Extends VAbs to be able to handle 256-bit width vectors.
2022-09-27 13:56:38 +00:00
Ryan Houdek 46bde401bd Merge pull request #2019 from Sonicadvance1/remove_mov
IR: Removes Mov IR op
2022-09-26 19:49:52 -07:00
Ryan Houdek 01beac4956 Merge pull request #2018 from Sonicadvance1/remove_vextractelement
IR: Removes VExtractElement
2022-09-26 19:49:46 -07:00
Ryan Houdek fcd981e6b7 Merge pull request #2017 from Sonicadvance1/remove_vbitcast
IR: Removes unnecessary VBitcast IR op
2022-09-26 19:44:12 -07:00
Ryan Houdek 825833cfcc IR: Removes Mov IR op
This is unused and shouldn't ever be used.
2022-09-26 16:04:27 -07:00
Ryan Houdek 4a4c49bf68 IR: Removes VExtractElement
This is a duplicate of VDupElement since AArch64 doesn't support an
element extract plus zero of the rest of the register.

Removes and replaces its uses with VDupElement.
2022-09-26 15:55:41 -07:00
Ryan Houdek 763cea423a IR: Removes unnecessary VBitcast IR op
This instruction was purely a move that did format reinterpretation.
This was necessary with LLVM when we had implicit IR op sizes.

Now that all vector ops are explicitly sized, this is not only
redundant, but also completely unnecessary sicne we don't support LLVM
anymore.

Remove the op, which technically is a very minor optimization for the
two instructions that still used it.
2022-09-26 15:42:24 -07:00
Ryan Houdek 0fee355ff5 Merge pull request #2016 from lioncash/pred
Arm64/VectorOps: Make use of static predicate registers
2022-09-26 14:48:26 -07:00
Ryan Houdek 6f6f3c9dc5 Merge pull request #2012 from Sonicadvance1/32bit_vdso
32-bit VDSO support
2022-09-26 14:46:27 -07:00
Ryan Houdek 25e5d88ab2 VDSO Emulation: Wires up support for 32-bit VDSO 2022-09-26 14:35:38 -07:00
Ryan Houdek 87013340bb Thunks: Adds support for building 32-bit. Only VDSO for now. 2022-09-26 14:35:38 -07:00
Ryan Houdek 47c075ccc9 Thunks/VDSO: Add 32-bit linker script 2022-09-26 14:35:38 -07:00
Ryan Houdek b1a32d4ccf Thunks: Ensure fexthunks functions are hidden visible by default 2022-09-26 14:35:38 -07:00
Ryan Houdek 7af6a8dbdf Thunks/VDSO: Extend to support clock_gettime64 2022-09-26 14:35:38 -07:00
Ryan Houdek 383e99e4ef unittests: Extend VDSO test for gettime64 2022-09-26 14:35:38 -07:00
Ryan Houdek 8c7cfc4d11 Config: Adds support for unique 32-bit GuestThunk path 2022-09-26 14:01:48 -07:00
lioncash 496ee730c8 Arm64/VectorOps: Make use of static predicate registers
Since PR #2003, we now set up some predicate registers within
FillStaticRegs. We can now make use of those instead of manually setting
up predicate registers inside the IR opcodes.
2022-09-26 13:20:01 +00:00
Ryan Houdek 71f7ff5101 Merge pull request #2014 from Sonicadvance1/map_interp_first
ELFCodeLoader: Map interpreter first
2022-09-26 02:42:39 -07:00
Ryan Houdek 1ea00f68a2 Merge pull request #2010 from Sonicadvance1/add_support_for_32_bit_thunk_op
Thunks: Implement the Thunk IR op for 32-bit mode
2022-09-26 02:33:47 -07:00
Ryan Houdek 8df7c2d84f ELFCodeLoader: Map interpreter first
This more closely matches behaviour of the kernel.
Provides an example in source to ensure we don't break it in the future.
2022-09-25 18:07:31 -07:00
Ryan Houdek 46557a7a1f ELFCodeloader: Map once and then use MAP_FIXED to overwrite
Instead of mapping to find a range, unmapping, and then submapping
inside of it.

Map once, then mmap with MAP_FIXED to overwrite the mapping.
This fixes an issue where if you enabled ASAN then it would stick
additional mappings inbetween where we want to map. Thus breaking asan.
2022-09-25 17:50:16 -07:00
Ryan Houdek d8c2a8271f Merge pull request #2009 from Sonicadvance1/libvulkan_fix_print
Thunks/libvulkan: Fixes print for 32-bit
2022-09-25 13:33:00 -07:00
Ryan Houdek cc4c705fc0 Merge pull request #2008 from Sonicadvance1/disable_32_bit_x11
ThunkLibs: X11/Xext: Removes two functions that don't exist on 32-bit
2022-09-25 13:32:54 -07:00
Ryan Houdek 22f249fcf6 Thunks: Implement the Thunk IR op for 32-bit mode
Use the fastcall ABI for 32-bit x86 to make our lives easier.
Fastcall ABI puts the first two 32-bit arguments in ECX and EDX
respectively.

Compilers are nice today and allow us to do cross-abi function calls
like this.
2022-09-25 13:16:30 -07:00
Mai c262362a03 Merge pull request #2011 from Sonicadvance1/add_missing_flake
FEXLinuxTests: Adds missing pthread_cancel flake status
2022-09-24 20:53:22 -04:00
Ryan Houdek 212df9aa7b FEXLinuxTests: Adds missing pthread_cancel flake status
Missed the 32-bit version of this test
2022-09-24 10:05:36 -07:00
Ryan Houdek 691e39ec76 Thunks/libvulkan: Fixes print for 32-bit
Value passed in to this print will be 32-bit or 64-bit depending on
arch.

Noticed this while tinkering around and is easy enough to solve today.
2022-09-24 09:56:50 -07:00
Ryan Houdek 107cae2975 ThunkLibs: X11/Xext: Removes two functions that don't exist on 32-bit
_XData32 and _XRead32 don't exist as real functions in 32-bit versions
of these libraries, these end up just being defines that redirect to the
non-suffixed versions of the functions.

Noticed this while tinkering around and is easy enough to solve today.
2022-09-24 09:51:40 -07:00
Ryan Houdek 6742e0c376 Merge pull request #2003 from lioncash/svespill
JITs: Handle spilling/filling 256-bit vectors
2022-09-23 17:27:53 -07:00
Ryan Houdek 8f70137b1a Merge pull request #1981 from neobrain/feature_flt_catch2
FEXLinuxTests: Migrate to Catch2
2022-09-23 17:27:45 -07:00
lioncash 707db51b1b Arm64Emitter: Amend comment for GPR temporaries
Only x3 can be used across spill boundaries.
2022-09-24 00:12:01 +00:00
lioncash 5b5fa1aa29 x86_64/JIT: Handle pushing and popping 256-bit values 2022-09-24 00:11:56 +00:00
Ryan Houdek 2b9cc9666a Merge pull request #2006 from Sonicadvance1/remove_splat
IR: Removes SplatVector{2,4}
2022-09-23 14:42:44 -07:00
Ryan Houdek 82eba22292 Merge pull request #2007 from Sonicadvance1/remove_insscalar
IR: Removes VInsScalarElement
2022-09-23 14:42:36 -07:00
Ryan Houdek 5c84e8f23c IR: Removes VInsScalarElement
This IR op duplicates what VInsElement does.
2022-09-22 17:57:20 -07:00
Ryan Houdek 081b61677a IR: Removes SplatVector{2,4}
These IR ops are redundant and mostly unused.
VDupElement does exactly what these operations were already doing and
more closely matches what the hardware wants.
2022-09-22 17:49:28 -07:00
lioncash 9c54814b98 Arm64Emitter: Handle spilling 256-bit dynamic regs 2022-09-22 12:40:54 +00:00
lioncash 35d7b855ed Arm64Dispatcher: Increment code buffer size
vixl hits an assertion in CodeBuffer's Emit() function since there's no
space for any more instructions with the changes made to handle SVE.
2022-09-22 12:40:54 +00:00
lioncash 0b8799274c Arm64Emitter: Handle filling/spilling 256-bit static FPRs
Drops in handling for spilling/filling FPRs using SVE for supporting
AVX.

Also alters the dispatcher and JIT a little to avoid accidentally clobbering
TMP4 (x3 as of this commit)
2022-09-22 12:40:54 +00:00
lioncash ace2b737d8 Arm64Emitter: Initialize fixed predicate register values in FillStaticRegs
Allows us to have values set up in a way that we don't need to
constantly set up predicates in IR ops.
2022-09-22 12:40:54 +00:00
Mai ad85268524 Merge pull request #2005 from Sonicadvance1/fix_sve_vectorimm
Arm64: Fixes SVE VectorImm
2022-09-22 08:38:10 -04:00
Mai 832a320e22 Merge pull request #2004 from Sonicadvance1/update_vixl3
Update vixl external
2022-09-22 08:33:39 -04:00
Ryan Houdek 83763df6fd Arm64: Fixes SVE VectorImm
SVE DUP instruction does sign extension on the incoming immediate, while
ASIMD MOVI does zero extension.

If the immediate doesn't fit then move in to a GPR first and then DUP
from GPR.
2022-09-22 01:16:58 -07:00
Ryan Houdek bee868e9ba Update vixl external 2022-09-22 01:10:25 -07:00
Tony Wasserka c41de81694 FEXLinuxTests: Drop support for now unused "args:" annotations 2022-09-22 10:03:38 +02:00
Tony Wasserka 41aaeb1ff0 FEXLinuxTests: Migrate signal tests to Catch2 2022-09-22 10:03:37 +02:00
Tony Wasserka 16be2792ab FEXLinuxTests: Migrate FD test to Catch2 2022-09-22 10:03:36 +02:00
Tony Wasserka 6610bb355c FEXLinuxTests: Migrate VDSO test to Catch2 2022-09-22 10:03:35 +02:00
Tony Wasserka 4312fd7291 FEXLinuxTests: Migrate SMC tests to Catch2 2022-09-22 10:03:33 +02:00
Tony Wasserka 89a225a96d FEXLinuxTests: Enable use of Catch2 in tests 2022-09-22 10:03:31 +02:00
Ryan Houdek a590977639 Merge pull request #1984 from Sonicadvance1/functional_thunk_ci
Thunks: Adds functional thunk testing to CI
2022-09-20 11:14:03 -07:00
Ryan Houdek 0d0d116bde Merge pull request #2002 from lioncash/slots
JITs: Expand max spill slot size to 32 bytes
2022-09-19 15:22:25 -07:00
lioncash 341bdb5a54 JITs: Handle 32 byte spills and fills
Puts in the plumbing necessary to handle spilling and filling 256-bit
data.
2022-09-19 22:00:20 +00:00
lioncash 2f3dbfb289 JITs: Expand max spill slot size to 32 bytes
This will be necessary to handle spilling 256-bit vectors.
2022-09-19 19:51:32 +00:00
Ryan Houdek 169cfbbeed Merge pull request #2001 from lioncash/defmove
Arm64: Centralize location for register defines
2022-09-19 11:45:24 -07:00
Ryan Houdek f97a4afd8f Merge pull request #2000 from Sonicadvance1/fix_struct_verifier_ubuntu_20_04
CI: Fixes struct verifier on Ubuntu 20.04
2022-09-19 11:44:38 -07:00
lioncash 868e4a6d81 Arm64: Centralize location for register defines
Gets rid of a few repeated definitions and allows the emitter itself to
make use of these defines without causing a circular dependency on the
JIT.
2022-09-19 17:44:21 +00:00
Ryan Houdek 6f48f7d3ac CI: Fixes struct verifier on Ubuntu 20.04
Older clang fails to pull in these include paths when cross compiling.
Add them manually.
2022-09-19 01:27:54 -07:00
Tony Wasserka 1ed3ecb409 Merge pull request #1999 from neobrain/feature_toolchain_32bit
CMake: Add toolchain file for 32-bit cross-compiler
2022-09-19 09:30:57 +02:00
Tony Wasserka 2cb455b9d4 CMake: Add toolchain file for 32-bit cross-compiler 2022-09-19 09:19:22 +02:00
Ryan Houdek d4b5bf0f78 Merge pull request #1998 from Sonicadvance1/fix_struct_verifier
StructVerifier: Fixes CI failure
2022-09-18 19:26:21 -07:00
Ryan Houdek 69013772c1 StructVerifier: Fixes CI failure
The x86 runner had unattended-upgrades accidentally still enabled. It
upgraded a bunch of development packages which broke CI.

Fix the struct verifier so it works with the new packages.
Sadly python3-clang doesn't support all the new CursorKind types so we
need to self-define some of them for now.

Once this tool gets converted over to C++ it will be a non-issue.
2022-09-16 19:45:13 -07:00
Ryan Houdek 3448c83431 Merge pull request #1997 from neobrain/refactor_flt_unified_cmake
FEXLinuxTests: Build 32-bit and 64-bit test variants separately
2022-09-16 17:14:24 -07:00
Tony Wasserka e4c84542ea FEXLinuxTests: Build 32-bit and 64-bit test variants separately
This allows to use different toolchain files for each and it reduces
build system repetition in test target setup.

The "tests-32" directories has been integrated into the "tests" one. Tests
that should only run on 32-bit are detected by their filename ending with
".32.cpp" now.
2022-09-16 11:24:18 +02:00
Ryan Houdek acddc0323b Merge pull request #1995 from Sonicadvance1/disable_bad_test
unittests: Disable gvisor pselect test
2022-09-15 22:06:37 -07:00
Ryan Houdek c4285f0d30 unittests: Disable gvisor pselect test
This has a badly coded test that can hang forever. Our timeout kills it
at 5 minutes, which causes it to not even fall down the flake path.

Just disable it outright because of the bad test.
2022-09-15 16:29:02 -07:00
Ryan Houdek 977d6dd247 Merge pull request #1993 from lioncash/vuravg
VectorOps: Handle 256-bit VURAvg
2022-09-15 15:09:46 -07:00
lioncash 2bb27fffb7 VectorOps: Handle 256-bit VURAvg 2022-09-15 19:38:31 +00:00
Ryan Houdek 0261ed353d Merge pull request #1992 from lioncash/uminv
VectorOps: Handle 256-bit VUMinV
2022-09-15 12:05:27 -07:00
Ryan Houdek 96fecfd7c5 Merge pull request #1989 from Sonicadvance1/ci_flakes
CI: Adds support for flakes
2022-09-15 12:05:08 -07:00
Ryan Houdek 9fac1b8105 CI: Adds support for flakes
If a test is marked as a flake then it will be tried five times before
giving up.

Works around the problem of needing to babysit CI once a PR is pushed.
As long as we have all the flake tests marked.
2022-09-15 11:44:56 -07:00
lioncash 95fbcd7b9a VectorOps: Handle 256-bit VUMinV 2022-09-15 18:36:46 +00:00
Ryan Houdek 6adf227611 Merge pull request #1990 from Sonicadvance1/uninstall
cmake: Adds uninstall target
2022-09-15 11:33:27 -07:00
Ryan Houdek b5cb429243 Merge pull request #1991 from lioncash/bits
Interpreter: Handle 256-bit VAnd/VBic/VOr/VXor
2022-09-15 11:23:49 -07:00
Ryan Houdek 26ba8079a3 cmake: Adds uninstall target
Following guidance from cmake's FAQ:
https://gitlab.kitware.com/cmake/community/-/wikis/FAQ#can-i-do-make-uninstall-with-cmake

Due to some of the special handling that we do with installs, we need to
do additional uninstall handling that the install manifest doesn't cover.

Specifically we need to add additional uninstall targets for:
- FEXInterpreter
- binfmt_misc
- guest_thunks (Doing its own uninstall target, so passthrough)

While it isn't generally advised to install and uninstall through source
systems, this is something that users want to do all the time.
This has been asked for a couple of times now.

Fixes #1592
2022-09-15 11:22:24 -07:00
lioncash f999d30bc5 Interpreter: Handle 256-bit VAnd 2022-09-15 16:18:14 +00:00
lioncash ecb1cc4ed4 Interpreter: Handle 256-bit VBic 2022-09-15 16:18:14 +00:00
lioncash 5622bcae16 Interpreter: Handle 256-bit VOr 2022-09-15 16:18:14 +00:00
lioncash f4539ee289 Interpreter: Handle 256-bit VXor 2022-09-15 16:18:10 +00:00
Ryan Houdek d2138694b4 Merge pull request #1986 from neobrain/refactor_flt_cmake_cleanup
FEXLinuxTests: Use the build system instead of setting up compile flags via source-code annotations
2022-09-15 01:14:08 -07:00
Tony Wasserka 1767e21273 FEXLinuxTests: Use the build system instead of setting up compile flags via source-code annotations
The intent of these annotations was presumably to make it easier to adjust
build settings on a per-test basis, but doing this in the build system is
actually much cleaner.
2022-09-15 08:55:18 +02:00
Ryan Houdek 8f9d799342 Merge pull request #1988 from Sonicadvance1/fexserver_wait_old_kernel
FEXServer: Fix waiting on kernel version older than 5.3
2022-09-14 17:07:41 -07:00
Ryan Houdek a3b0b246f4 FEXServer: Fix waiting on kernel version older than 5.3
pidfd_open was added in kernel 5.3 so older kernel devices weren't able
to use `FEXServer -w`. If the syscall doesn't give us an FD to the
process, then use a pipe instead.

Since we are only polling for the FD to hangup this works for us.
2022-09-14 16:04:26 -07:00
Ryan Houdek dee85f14fe Merge pull request #1987 from Sonicadvance1/fix_fhu_syscalls
FHU: Convert to a interface target
2022-09-14 15:55:04 -07:00
Ryan Houdek 27309114be FHU: Convert to a interface target
Noticed recently that `FEXServer -w` was broken and couldn't understand
why. Turns out that FHU syscall handling was /always/ falling down the
`#else` path in the handlers since cmake `add_definitions` follows
folder scoping rules.

This means it was always returning -1, which was causing FEXServer's
pidfd_open usage to always receive -1, which meant the sendmsg with FD
was always failing, which meant the `FEXServer -w` would forever wait
for a message that was never sent.

Converting the utility over to a target not only fixes definition
scoping problems, but also makes the other paths actually work.

This found some compiling bugs and instead lets us define SYS_pidfd_open
if it doesn't exist. Letting the kernel return the ENOSYS if it doesn't
exist on that platform.

Main thing, fixes FEXServer -w hanging forever.
2022-09-14 14:58:23 -07:00
Ryan Houdek 704afed97b Merge pull request #1985 from neobrain/refactor_thunkgen_fmt
Thunks/gen: Use fmt for writing formatted output
2022-09-14 13:52:15 -07:00
Ryan Houdek 121f0a2c6c CI: FetchRootFS More robust rootfs fetching
Permissions mean we need to delete the folder before extracting.
On error make sure to delete the image file as well to ensure it reruns
everything.
2022-09-14 13:18:24 -07:00
Ryan Houdek 1c580ec92c Thunks: Adds functional thunk testing to CI
This is the bare minimum, it only tests glxinfo and vulkaninfo with and
without thunks. Nothing more special than that. Already found the .1 bug
with libvulkan host library loading.
2022-09-14 12:48:13 -07:00
Tony Wasserka ab8fc721a0 Thunks/gen: Use fmt for writing formatted output 2022-09-14 11:56:15 +02:00
Ryan Houdek 790447115c CI: Set CMAKE_INSTALL_PREFIX
This will be used in the next commit
2022-09-13 17:21:03 -07:00
Ryan Houdek 80abeac28a Thunks: Fixes a missing version number on libvulkan
Fixes an issue with loading libvulkan without development packages.
2022-09-13 17:20:06 -07:00
Ryan Houdek 54915f87ce Merge pull request #1978 from neobrain/refactor_astvisitor_to_frontendaction
Move thunk generator logic from ASTVisitor to ASTFrontendAction
2022-09-13 11:26:36 -07:00
Ryan Houdek f34f1309a7 Merge pull request #1983 from lioncash/vsqadd
VectorOps: Extend VSQAdd/VSQSub/VUQAdd/VUQSub
2022-09-13 11:26:11 -07:00
Ryan Houdek 0ad52b7d19 Merge pull request #1982 from lioncash/vadd
VectorOps: Extend VAdd/VSub
2022-09-13 11:24:39 -07:00
lioncash 809f60df06 VectorOps: Handle 256-bit VSQSub 2022-09-13 16:43:42 +00:00
lioncash cbdcd8253c VectorOps: Handle 256-bit VSQAdd 2022-09-13 16:20:37 +00:00
lioncash 7ac2cd7cc8 VectorOps: Handle 256-bit VUQSub 2022-09-13 16:03:30 +00:00
lioncash cf1bb1348c VectorOps: Handle 256-bit VUQAdd 2022-09-13 16:03:27 +00:00
lioncash b60a26ff9e VectorOps: Handle 256-bit VSub 2022-09-13 15:40:45 +00:00
lioncash b805c07342 VectorOps: Handle 256-bit VAdd 2022-09-13 15:40:06 +00:00
Mai 8d69f539ac Merge pull request #1979 from Sonicadvance1/fix_thunkconfig
FEXConfig: Ensure APP_CONFIG_NAME isn't stored in json
2022-09-13 11:38:24 -04:00
Ryan Houdek c8c0054f67 FEXConfig: Ensure APP_CONFIG_NAME isn't stored in json
Also in FEXLoader make sure to use `EraseSet` for these runtime options.

Fixes a bug where the config was being set to nothing, breaking the
ThunksDB configuration option.
2022-09-13 00:08:20 -07:00
Tony Wasserka 1085385bbe Thunks/gen: Move logic from ASTVisitor to ASTFrontendAction
ASTVisitor is great for iterating over AST nodes by type, but most of our
analysis is based on symbol names. For this task, a lookup in DeclContexts
after parsing is complete is better suited.
2022-09-12 18:52:33 +02:00
Tony Wasserka 56460b220c Thunks/gen: Move definition of GenerateThunkLibsAction into gen.cpp 2022-09-12 18:52:33 +02:00
Mai a583ebe590 Merge pull request #1977 from Sonicadvance1/extend_arch_check
CMake: Extend AArch64 check to include arm64
2022-09-09 22:36:30 -04:00
Ryan Houdek ce8175a800 CMake: Extend AArch64 check to include arm64
This has been seen in some build environments, forgot to commit this a
while ago.
2022-09-08 15:34:35 -07:00
Mai c987e1ef44 Merge pull request #1974 from Sonicadvance1/update_release_process
Docs: Update Release docs
2022-09-08 11:48:41 -04:00
Mai b36ec152d2 Merge pull request #1976 from Sonicadvance1/support_simulator
Add support for the vixl simulator
2022-09-08 11:47:46 -04:00
Ryan Houdek 44c62e703e github: Adds vixl simulator CI 2022-09-07 20:08:50 -07:00
Ryan Houdek 0f59c1d5e3 Add support for the vixl simulator
This will allow CI to test ARM features before we have any hardware that
supports it.
2022-09-07 19:54:07 -07:00
Ryan Houdek 5739f0b459 Update vixl 2022-09-07 19:10:13 -07:00
Ryan Houdek 4145fabfb6 Docs: Update Release docs
Reorder PPA building to be after the github tag. PPA takes a while to
run, so good to get it out of the way up front so it can be handled in
the background while doing the rest of the release.

Also update the link which was renamed.
2022-09-05 11:08:44 -07:00
Ryan Houdek ae34b1e521 Docs: Update for release FEX-2209 2022-09-05 10:32:07 -07:00
Ryan Houdek c17da25617 Merge pull request #1973 from neobrain/refactor_single_thunkgen_output
Thunks: Consolidate all generated code to one file per library per platform
2022-09-05 10:20:54 -07:00
Tony Wasserka cc8ef16240 Thunks/gen: Consolidate all generated code to one file per library per platform 2022-09-05 15:03:49 +02:00
Tony Wasserka 30fac81c41 Thunks/gen: Remove unused symtable generation code 2022-09-05 15:03:49 +02:00
Ryan Houdek bbcca80b60 Merge pull request #1971 from Sonicadvance1/termux_namespace_collision
Syscalls: Use underscored shm syscall names
2022-09-05 02:55:29 -07:00
Ryan Houdek 56ae3716bb Syscalls: Use underscored shm syscall names
Removes us needing to carry a patch downstream in the termux package
repo.
2022-09-05 02:39:42 -07:00
Ryan Houdek 2fba3a4a9a LinuxAllocator: Uppercase function names
Not only does this match our internal naming convention. This avoids the
shm symbol namespace collision as well.
2022-09-05 02:39:42 -07:00
Ryan Houdek 8ba0312d40 Syscalls: Prefix _ to shm syscall names to avoid namespace conflict
Termux uses defines for these, so our token pasting fails, but we also
still want to use their define so we can fall down their emulation
library whenever possible.

Prefix an underscore to be able to use both our number definitions and
their defines in the same file.
2022-09-05 02:39:42 -07:00
Ryan Houdek 53623ffa72 Merge pull request #1941 from Sonicadvance1/vdso_implementation
Thunks: Adds VDSO thunk library
2022-09-04 23:02:49 -07:00
Mai d6199687f4 Merge pull request #1970 from Sonicadvance1/fix_fresh_runner
Github: Fix fresh runner rootfs checkout
2022-09-02 21:49:22 -04:00
Ryan Houdek 7cb413dde6 Github: Fix fresh runner rootfs checkout
build folder doesn't exist on a freshly started runner.
Executing the rootfs fetch script doesn't care what the working
directory is. Doesn't need to be in `build/` which doesn't exist on a
fresh runner and will fail chdir.
2022-09-02 15:15:21 -07:00
Ryan Houdek c5a7fc1e2a Resolve most VDSO comments 2022-09-02 15:13:18 -07:00
Ryan Houdek 0869b0aa29 FEXLinuxTests: Adds a VDSO test
Ensures VDSO is working as correctly as it can.
2022-09-02 13:31:36 -07:00
Ryan Houdek 4e81f847f2 ELFCodeLoader: Load VDSO thunk at application startup
This ensures it will always be available to the application as long as
the library is installed.
2022-09-02 13:31:36 -07:00
Ryan Houdek 8fa27d2d9f Thunks: Adds VDSO thunk library
VDSO is heavily abused by Proton games to the point it is showing up as
CPU time.
Implement a guest-facing only thunk library using the hardcoded VDSO
interface in Thunks.

If available this will always be loaded on application load and set the
auxv value to support it.

This requires a bit of special treatment as our first user of linker
scripts since the format of the ELF must be careful crafted to not break
applications trying to parse it.

This library exposes a handful of symbols:
- clock_gettime
- clock_getres
- gettimeofday
- time
- getcpu
- All previous with `__vdso_` prefix
- LINUX_2.6

All of these symbols get routed directly to the host architecture VDSO
interface if they exist.
AArch64 doesn't have getcpu or time VDSO.

In a microbench, VDSO improved bench times substantially
x86-64 host: 3.612s -> 1.369s - 2.63x speed
AArch64 host: 3.821s -> 2.284s - 1.67x speed
  - AArch64 isn't as improved due to missing VDSO symbols

This is also our first /always/ enabled thunk as long as the file exists
2022-09-02 13:31:36 -07:00
Ryan Houdek e5a8a29efe Thunks: Add support for lds linker script on Guest libraries
This is going to be necessary in the next commit
2022-09-02 13:31:36 -07:00
Ryan Houdek f967f53176 Thunks: Adds VDSO specific thunks
x86-64 has five symbols within VDSO that we need to emulate.

Pass this through either glibc or host vdso if the symbol exists.
AArch64 doesn't have the time or getcpu vdso interface, so fall down
glibc instead.
2022-09-02 13:31:36 -07:00
Ryan Houdek 31fefaae0d Merge pull request #1969 from Sonicadvance1/fexrootfsfetcher_fix_crash
FEXRootFSFetcher: Fix crash if curl fails to download rootfs definition file
2022-09-02 11:07:03 -07:00
Ryan Houdek 7f9edbf39e Merge pull request #1968 from Sonicadvance1/new_domain
New domain.
2022-09-02 10:57:35 -07:00
Ryan Houdek 760b9c8e7f FEXRootFSFetcher: Fix crash if curl fails to download rootfs definition file
This can happen when the link goes down or internet blip.
2022-09-02 10:53:51 -07:00
Ryan Houdek cc7fb008fc New domain.
Needed to fix FEXRootFSFetcher from #1967
2022-09-02 10:43:07 -07:00
Ryan Houdek 98dbfbe654 Merge pull request #1949 from lioncash/interp-op
InterpreterOps: Extend SSAData size to accomodate 256-bit operations
2022-09-02 09:54:58 -07:00
lioncash 6444726614 InterpreterOps: Use designated initializer for IR op data
Same behavior, but keeps everything all initialized at the point of
declaration, rather than after the fact.
2022-09-02 12:40:25 -04:00
lioncash 0a562fbb10 InterpreterOps: Extend SSAData to handle 256-bit vectors 2022-09-02 12:40:14 -04:00
Ryan Houdek 7d8950de40 Merge pull request #1966 from lioncash/bool
Arm64/JIT: Rename CanUseSVE to HostSupportsSVE
2022-09-01 13:13:00 -07:00
lioncash 3fdde0c90b Arm64/JIT: Rename CanUseSVE to HostSupportsSVE
This is a much more descriptive name.

Spawned off of discussion in #1944
2022-09-01 13:22:41 -04:00
Ryan Houdek e776f4cd4e Merge pull request #1948 from lioncash/svebit
VectorOps: Extend VAnd/VBic/VOr/VXor
2022-08-30 17:57:51 -07:00
Ryan Houdek e7d7dd13d7 Merge pull request #1945 from lioncash/vectormov
VectorOps: Extend VMov
2022-08-30 17:57:00 -07:00
Ryan Houdek d5c83a2e45 Merge pull request #1950 from lioncash/x86run
HostRunner: Handle upper YMM lanes in sigsegv handler
2022-08-30 17:55:38 -07:00
Ryan Houdek 37ccb13917 Merge pull request #1946 from lioncash/x86dep
x86_64/JIT: Resolve lingering fmt deprecation warning
2022-08-30 17:54:57 -07:00
Ryan Houdek 8439cf410f Merge pull request #1944 from lioncash/sveimm
VectorOps: Extend VectorImm
2022-08-30 17:54:32 -07:00
lioncash 614d9448f7 HostRunner: Handle upper YMM lanes in sigsegv handler
Now we properly copy out the upper lanes instead of ignoring them.
2022-08-25 14:13:19 -04:00
lioncash 70efdbba5d VectorOps: Handle 256-bit VectorImm
Extends VectorImm to be capable of using SVE to handle 256-bit length
vectors.
2022-08-24 13:28:48 -04:00
Stefanos Kornilios Mitsis Poiitidis 12fee91ebb Merge pull request #1947 from neobrain/fix_tmpnam_warning
unittests/ThunkLibs: Fix warning about "dangerous" use of tmpnam
2022-08-24 20:13:04 +03:00
lioncash bcb7e20619 VectorOps: Handle 256-bit VXor 2022-08-24 12:52:05 -04:00
lioncash afc5e8a140 VectorOps: Handle 256-bit VOr 2022-08-24 12:52:05 -04:00
lioncash 4be6626c89 VectorOps: Handle 256-bit VBic 2022-08-24 12:52:02 -04:00
lioncash 5f2b6d629b VectorOps: Handle 256-bit VAnd 2022-08-24 12:43:51 -04:00
Tony Wasserka 1b8f5f08f0 unittests/ThunkLibs: Fix warning about "dangerous" use of tmpnam
tmpnam is considered insecure since it's vulnerable to TOCTOU issues.
This is not an issue for these tests, but replacing tmpnam is not any
more complicated than silencing the warning.
2022-08-24 18:15:51 +02:00
lioncash 79674c697a x86_64/JIT: Resolve lingering fmt deprecation warning
Just a log that was missed during the previous fmt deprecation cleanup.
2022-08-24 11:49:02 -04:00
lioncash 416d8c1df6 VectorOps: Handle 256-bit VMov
Kind of sucky that SVE doesn't have a convenient way to manipulate
predicate registers with immediates or anything to make this nicer (that
I know of).

Having to use a temp to clear the upper part of the vector reliably is
bleh.
2022-08-24 11:35:34 -04:00
Ryan Houdek d03b6a9382 Merge pull request #1942 from lioncash/zero
VectorOps: Extend VectorZero
2022-08-23 20:00:39 -07:00
lioncash df22e0c796 x86_64/JITClass: Add ToYMM helper
Will be used in subsequent changes to handle 256-bit operations in the
x86-64 backend
2022-08-23 12:58:25 -04:00
lioncash 79a3bd75cc VectorOps: Handle 256-bit VectorZero 2022-08-23 12:57:59 -04:00
Stefanos Kornilios Mitsis Poiitidis e6acdcc583 Merge pull request #1940 from neobrain/refactor_1868_cleanups
Thunks: Minor cleanups for signature-based function pointer thunking
2022-08-23 01:08:51 +03:00
Tony Wasserka cd78984228 Thunks: Define _M_X86_64/_M_ARM_64 when invoking thunkgen
This avoids the need to provide a fallback definition for platform-specific
macros. The definitions are only added host-side, since only Host.h is
included in any interface files.
2022-08-22 18:11:28 +02:00
Tony Wasserka 2f007c5f0a Thunks: Remove unused parameters of exports initializer 2022-08-22 18:11:28 +02:00
Tony Wasserka ae64a1e30c Thunks: Replace compiler-specific attributes with FEX_DEFAULT_VISIBILITY 2022-08-22 18:11:28 +02:00
Stefanos Kornilios Mitsis Poiitidis 097184c3e0 Merge pull request #1926 from Sonicadvance1/no_irloader_on_no_tests
IRLoader/TestHarnessLoader: Don't build if not building tests
2022-08-21 09:50:44 +03:00
Stefanos Kornilios Mitsis Poiitidis 84a95adae1 Merge pull request #1931 from Sonicadvance1/support_thunksdb_in_config
Thunks: Support direct thunk config in configuration files
2022-08-21 09:50:08 +03:00
Stefanos Kornilios Mitsis Poiitidis 123b6728e4 Merge pull request #1932 from Sonicadvance1/fix_allocator_perf_hit
64BitAllocator: Fixes a significant state tracking perf problem
2022-08-21 09:47:54 +03:00
Ryan Houdek 69b4fc98ed 64BitAllocator: Fixes a significant state tracking perf problem
Due to how the 64-bit allocator previously worked, it was never subjected to memory
regions larger than 64GB to be tracked. With the change in
PR #1885, this has changed to have regions that will hit sizes larger
than 170TB on some platforms.

Better yet, even with smaller regions it still had a performance issue,
it just wasn't as visible.

First problem: We used MemSet instead of MemClear for the live page
clearing. This caused pages to be claimed as "always in use".

This would cause us to always scan the entire region on allocation, find
that it didn't work and allocate a fresh region on every slab
allocation.
jemalloc saving us here since it allocates slabs from the OS fairly
aggressively.

Second problem: We used MemSet (now changed to MemClear) to "clear" the
state tracking for pages.

This causes ~600MB of memory to be used purely for state tracking.
This was physically backed since we were writing to every bit of
tracking for handling 256TB of VA.
This had a fault dance with the kernel for every new page being hit
here.
Instead of clearing the the bits with a memset, clear it with madvise so
it doesn't consume physical pages at all.

This means we use significantly less physical memory for 32-bit
applications.

With this change, pressure-vessel startup time goes from 24 seconds down
to 17 seconds. 70% of the original startup time.
But really the main savings here comes from the memory reduction that
PR #1885 ballooned, but has been an unseen problem before.

Before that PR we were burning 2MB of physical memory per region for no
reason.
After that PR we were burning up to 600MB of physical memory per region
for no reason. Changing a bit depending on how large the region ended up
being.

This now ends up being 2 pages starting out and grows as more pages are
are used. A significant improvement.
2022-08-20 15:32:40 -07:00
Ryan Houdek 35cf7703b1 Thunks: Support direct thunk config in configuration files
Previously in order to enable thunks, we needed an independent
description file of which thunks to be enabled. This is nice for quickly
testing out new games by setting `FEX_THUNKCONFIG` environment variable.

For users that just want to enable thunks this is an unwieldy
indirection that doesn't make much sense at a glance.

Previously this meant you needed two files as an example:
```
  ryanh@ubuntu-linux-20-04-desktop:~/.fex-emu$ cat thunks.json
  {
    "ThunksDB": {
      "GL": 1,
      "Vulkan": 1
    }
  }
  ryanh@ubuntu-linux-20-04-desktop:~/.fex-emu$ cat AppConfig/EnderLiliesSteam-Linux-Shipping.json
  {
    "Config": {
      "ThunkConfig":"~\/.fex-emu\/thunks.json"
    }
  }
```

Instead of this unwieldy redirection just support `ThunksDB` json
directly in the AppConfig.

```
  ryanh@ubuntu-linux-20-04-desktop:~/.fex-emu$ cat AppConfig/EnderLiliesSteam-Linux-Shipping.json
  {
    "Config": {
      <...>
    },
    "ThunksDB": {
      "GL": 1,
      "Vulkan": 1
    }
  }
```

As can be seen this makes this significantly easier for new users
getting in to thunks. Depending on which path to enable thunks the user
is more comfortable with, they can still enable them using the
`ThunkConfig` option or embedding directly in the application
configuration.

Additionally this removes the older non-ThunksDB path to loading thunks.
All users of it have moved on to using ThunksDB.
2022-08-20 10:02:13 -07:00
Ryan Houdek 8c1137543b Config: Adds APP_CONFIG_NAME meta config option
We have separate configurations for the Application path versus the
application name we are using as a configuration choice.

Example 1: FEXBash "wine Crysis64.exe"

Previous APP_FILENAME will contain `/usr/bin/wine`, which is still used
elsewhere.

This new APP_CONFIG_NAME will contain `Crysis64.exe`

Example 2: FEXBash glxgears

Previous APP_FILENAME will contain `/usr/bin/glxgears`
APP_CONFIG_NAME will contain `glxgears`

We didn't have this exposed any other way before.
2022-08-20 09:58:23 -07:00
Ryan Houdek b8e66e56a0 Config: Remove log message about Config file existing without Config json object 2022-08-20 09:58:04 -07:00
Ryan Houdek fbb008e510 Merge pull request #1929 from Sonicadvance1/support_variadic_struct_packing
Thunks/X11: Support Variadic stack packing
2022-08-20 06:31:22 -07:00
Ryan Houdek 04678f8404 Merge pull request #1885 from FEX-Emu/skmp/simpler-memory-stealing
Allocator: Simplify StealMemory, make it less chatty with kernel space
2022-08-20 06:20:36 -07:00
Stefanos Kornilios Mitsis Poiitidis 1b5aaf1fb8 Allocator: Simplify StealMemory, make it less chatty with kernel space 2022-08-20 13:10:50 +03:00
Ryan Houdek d8e4873f43 Thunks/X11: Support Variadic stack packing
Found an issue with wine + DXVK + thunks where these were passing in
more than 7 arguments and crashing.

Create some assembly to support any size of variadic stack packing.
Only implemented for AArch64 for now.
2022-08-19 21:52:19 -07:00
Ryan Houdek 998a3d8353 Merge pull request #1928 from Sonicadvance1/more_x11_thunks
Thunks/X11: Adds missing XLibint functions
2022-08-19 17:08:40 -07:00
Ryan Houdek 336dedbed5 Merge pull request #1927 from Sonicadvance1/non_fatal_get_fdpath
FDUtils: Don't make unknown get_fdpath fatal
2022-08-19 17:08:32 -07:00
Ryan Houdek 955595be8a FDUtils: Don't make unknown get_fdpath fatal
Encountered this while running wine things.
Non-fatal so don't explode
2022-08-19 02:28:19 -07:00
Ryan Houdek 576bd4f69a Thunks/X11: Adds missing XLibint functions
Some of these were required to get thunking to work with Proton and
DXVK.
2022-08-19 00:03:54 -07:00
Ryan Houdek cd6915917b IRLoader/TestHarnessLoader: Don't build if not building tests
If we're not building tests then just don't build the IRLoader since it
won't get used.
2022-08-17 17:04:13 -07:00
Ryan Houdek 0adbe31112 Merge pull request #1919 from Sonicadvance1/termux_shm_library
Termux: Add android-shmem library
2022-08-15 11:14:27 -07:00
Ryan Houdek d5138f509c Merge pull request #1917 from Sonicadvance1/fix_compile_without_jemalloc
Thunks: Fix compile without jemalloc
2022-08-15 11:14:17 -07:00
Ryan Houdek 1fe6fc3feb Merge pull request #1920 from Sonicadvance1/tests_from_host_features
unittests: Support skipping unit tests based on host feature support
2022-08-15 11:07:02 -07:00
Ryan Houdek c03a7fd482 Merge pull request #1924 from 1ace/fix-git-abbrev
cmake: fix incorrect assumption about the value of git's core.abbrev
2022-08-15 10:44:59 -07:00
Eric Engestrom 31c47f05f4 cmake: fix incorrect assumption about the value of git's core.abbrev
While the default value of `git config core.abbrev` is `7` (used to
truncate the commit hash in several git commands) and most people don't
configure anything else, I happen to prefer when commits hashes stay
valid for a while, so I changed that value to `20`. Because of this,
FEX fails to build in External/FEXCore/Source/Interface/Core/CPUID.cpp:961

Let's be explicit in the git command about what minimum length of commit
hash we expect.
2022-08-15 16:25:27 +01:00
Ryan Houdek edad24479b unittests: Support skipping unit tests based on host feature support
For these unit tests we no longer need to put them in the disabled tests
file. Instead it will be skipped if the host doesn't support the feature
required.
2022-08-14 20:04:26 -07:00
Ryan Houdek bba58732e0 HostFeatures: Add new supported features flag 2022-08-14 19:56:14 -07:00
Ryan Houdek 8d1cc6cf51 Termux: Add android-shmem library
Otherwise we fail at linking.
2022-08-14 18:23:53 -07:00
Ryan Houdek f03d0be5ef Thunks: Fix compile without jemalloc
This only fixes the compile error. Thunks aren't expected to be used
without jemalloc enabled.
2022-08-13 11:29:11 -07:00
Ryan Houdek a2f4f494a9 Merge pull request #1916 from Sonicadvance1/server_socket_path_override
FEXServer: Support socket path override
2022-08-13 07:00:59 -07:00
Ryan Houdek 9de25c200c Merge pull request #1915 from Sonicadvance1/emulate_64bit_getdents
Linux: Emulate classic getdents syscall for x64 and x32
2022-08-13 07:00:52 -07:00
Ryan Houdek fe1f00aadb FEXServer: Support socket path override
This is necessary for the fexserver to function correctly when chrooting
in to our rootfs and doing things.

Requires independent rootfs script modifications which will come with
the next rootfs update.

Problem comes down to a chroot supporting multiple users, where our
typical use case is only one user. Bind the server file to a single
server for the entire chroot session regardless of users, solving this
problem inside the chroot.

Fixes apt-get inside of chroot, which runs as user _apt.
2022-08-12 22:52:30 -07:00
Ryan Houdek a847fac4ca Linux: Emulate classic getdents syscall for x64 and x32
Arm64 doesn't have the classic getdents syscall, only getdents64.
Old glibc versions (like 2.17) don't support getdents64

This fixes 64-bit ls with a centos 7 rootfs. Likely also fixes some other very
old applications.

Packing differences between 32-bit and 64-bit getdents means we need to
template this between the two types, otherwise it's quite similar.

No known applications rely on the 32-bit getdents but worked with test
applications.
2022-08-12 18:21:49 -07:00
Ryan Houdek 0496506fb7 Linux: Define linux_dirent types for x64
These don't match the compatibility defines and they aren't part of the
public interface.
2022-08-12 16:00:06 -07:00
Ryan Houdek e544591c9c Merge pull request #1901 from Sonicadvance1/build_thunks
CI: Build Thunks
2022-08-12 14:47:54 -07:00
Ryan Houdek e5237d1149 Github: Enable Thunkgen tests 2022-08-12 14:31:18 -07:00
Ryan Houdek b9c848c5e9 Thunks: Xext version check define function prototypes
Nothing major here
2022-08-12 14:31:18 -07:00
Ryan Houdek b64e61a793 Thunks: Check for X11 version
There is no define for this so we must generate our own.
Declare a type ourselves if the library is too old
2022-08-12 14:31:18 -07:00
Ryan Houdek e80e2bdafe Docs: Update for release FEX-2208 2022-08-10 08:19:24 -07:00
Stefanos Kornilios Mitsis Poiitidis ac23bce0ba Merge pull request #1898 from Sonicadvance1/fix_appconfig_name_bug
AppConfig: Fix bug with filename
2022-08-10 12:35:16 +03:00
Ryan Houdek 1051cd97cf Merge pull request #1897 from Sonicadvance1/more_gl_thunk_symbols
Thunks: Extends libGL interface to support more functions
2022-08-09 08:18:49 -07:00
Ryan Houdek 45330fdd5d Merge pull request #1900 from Sonicadvance1/Disable_unittestgenerator
Disable UnitTestGenerator
2022-08-09 08:08:36 -07:00
Ryan Houdek bd296d7a11 Merge pull request #1899 from Sonicadvance1/glob_data_dir
CMake: Support multiple json files in the root of Data/
2022-08-09 08:08:29 -07:00
Ryan Houdek 249e19bf2a Thunks: Asound fix building with alsa 1.2.2 2022-08-09 06:18:43 -07:00
Ryan Houdek baf52ed286 CI: Try building thunks 2022-08-09 05:30:53 -07:00
Ryan Houdek dd7e1baa78 Merge pull request #1896 from Sonicadvance1/glxGetProcAddress_query_self
Thunks: Make glXGetProcAddress self-query work
2022-08-09 04:02:46 -07:00
Ryan Houdek 80a209d6dc Thunks: Make glXGetProcAddress self-query work
We can just have it return itself.
Unknown why some games do this, but it happens and is expected to work.
2022-08-09 03:42:27 -07:00
Ryan Houdek ab228c1fcb Merge pull request #1895 from Sonicadvance1/glxGetProcAddress_non_fatal
Thunks: Make unknown glXGetProcAddress non-fatal
2022-08-09 03:40:16 -07:00
Ryan Houdek 0dde233b1e Disable UnitTestGenerator
This is currently unused and can just cause compilation issues.
Disable it until we start hooking up aggressive fuzzing tests.
2022-08-09 03:13:54 -07:00
Ryan Houdek cf0d92d968 AppConfig: Fix bug with filename
Turns out cmake will cut the filename at the first period, not the last.
Fixes an issue if we have application names with multiple periods in it.
2022-08-09 03:01:27 -07:00
Ryan Houdek ce4b052242 CMake: Support multiple json files in the root of Data/
Instead of just ThunksDB
2022-08-09 02:04:28 -07:00
Ryan Houdek 125bb3afe1 Thunks: Extends libGL interface to support more functions
Some of these functions are missing that games rely on.

Only glGetVkProcAddrNV will cause us pain, so avoid it for now.
The rest are simple and at least glXWaitX fixes some games.

Fixes #1893
Fixes #1893
2022-08-09 02:00:52 -07:00
Ryan Houdek 30ee2c18ef Thunks: Support one more argument in argument packing 2022-08-09 01:59:32 -07:00
Ryan Houdek ff1a5dd4c2 Thunks: Make unknown glXGetProcAddress non-fatal
Multiple games will query symbols that are leaked but we don't support.
It is safe to return nullptr in these cases.
Print a warning message still just in-case someone fails hard at GL.
2022-08-09 01:55:25 -07:00
Ryan Houdek b9c9c7d671 Merge pull request #1888 from FEX-Emu/skmp/synchronized-linking
Synchronized Block Linking
2022-08-09 00:48:35 -07:00
Ryan Houdek 62d9961bd1 Merge pull request #1868 from neobrain/feature_thunk_funcptrs_by_signature
Implement signature-based thunking of function pointers
2022-08-08 21:39:46 -07:00
Tony Wasserka 4b7ed9599d Thunks: Clarify return of a function pointer 2022-08-08 16:08:46 +02:00
Tony Wasserka c8951de9dd Thunks/X11: Add thunks for _XInitImageFuncPtrs and XInitImage 2022-08-08 16:08:46 +02:00
Tony Wasserka 63517c377d Thunks: Add a helper to make typed host function pointers guest-callable
This recurring pattern combines the existing helpers GetCallerForHostFunction
and LinkAddressToFunction.
2022-08-08 16:08:46 +02:00
Tony Wasserka 5005ebdfc9 Thunks/gen: Remove ThunkedCallback members not used anymore 2022-08-08 16:08:46 +02:00
Tony Wasserka a2615f0e52 Thunks/gen: Use uniform function naming for stub callbacks 2022-08-08 16:08:46 +02:00
Tony Wasserka 0317d381fe Thunks/X11: Add definitions for more API functions 2022-08-08 16:08:46 +02:00
Tony Wasserka 31b5181bca Thunks/gen: Drop now unneeded callback_unpacks file 2022-08-08 16:08:46 +02:00
Tony Wasserka 7ae59055ff Thunks/X11: Handle function pointer thunking for XInitThreads and _XReply 2022-08-08 16:08:46 +02:00
Tony Wasserka 3524117c23 Thunks: Unindent code blocks from previous patch 2022-08-08 16:08:46 +02:00
Tony Wasserka e8ad1ca0a0 Thunks: Use signature-based thunking of guest function pointers
This changes how host trampolines for guest functions are created. Instead
of doing this purely on the guest-side, it's either the host-side that
creates them in a single step *or* a cooperative two-step initialization
process must be used. In the latter, trampolines are allocated and partially
initialized on the guest and must be finalized on the host before use.
2022-08-08 16:08:46 +02:00
Tony Wasserka 3ac1650001 Thunks/X11: Thunk function pointers set up in XOpenDisplay 2022-08-08 16:08:46 +02:00
Tony Wasserka 4f8ae81562 Thunks/gen: Remove now unneeded hostcall bits 2022-08-08 16:08:46 +02:00
Tony Wasserka 329d624a99 Thunks/gen: Use signature-based thunking of host function pointers 2022-08-08 16:08:46 +02:00
Tony Wasserka de92624c9a Thunks: Use stronger types for interfaces 2022-08-08 16:08:45 +02:00
Tony Wasserka c6a034da40 Thunks: Move PackedArguments to a dedicated header 2022-08-08 16:08:45 +02:00
Tony Wasserka 737e76968a Thunks/X11: Add more symbols 2022-08-08 16:08:45 +02:00
Tony Wasserka 7c6d49155c Thunks: Define GUEST_THUNK_LIBRARY for guest-side thunkgen invocations 2022-08-08 16:08:45 +02:00
Tony Wasserka 4d404ea94d Thunks: Fix tests 2022-08-08 16:08:45 +02:00
Stefanos Kornilios Mitsis Poiitidis dc810c7a1e Fix arm64 build 2022-08-08 05:19:09 +03:00
Stefanos Kornilios Misis Poiitidis e0ee4e71f8 Fixes 2022-08-08 05:01:38 +03:00
Stefanos Kornilios Misis Poiitidis ab7dac90a2 Mask signals around Block Linking 2022-08-08 04:19:01 +03:00
Stefanos Kornilios Misis Poiitidis 3a423fbc41 Synchronized Block Linking 2022-08-08 03:33:23 +03:00
Ryan Houdek 0982ec617d Merge pull request #1887 from Sonicadvance1/fexserver_use_after_free
FEXServer: Fix unsafe vector insert/removal
2022-08-06 01:24:52 -07:00
Ryan Houdek b45f27c8d0 FEXServer: Fix unsafe vector insert/removal
Depending on the operation we will do a vector insert or removal while
iterating over the vector.

Fixes a use after free that asan found when insert caused the vector to
resize.
2022-08-05 23:00:13 -07:00
Ryan Houdek 54f62b6701 Merge pull request #1882 from lioncash/fmt
Externals: Update fmt to 9.0.0
2022-08-05 02:30:03 -07:00
Ryan Houdek 9664d98ba5 Merge pull request #1884 from Sonicadvance1/vulkan_headers
Thunks: Use external Vulkan-Headers
2022-08-05 01:50:57 -07:00
Ryan Houdek 0c318834b5 Thunks: Use external Vulkan-Headers
This will help compiling on older distros which are shipping older
Vulkan-Headers.

We need to catch newer Vulkan features earlier than what distros ship
since it is highly common that users will update their drivers through
means that make their drivers be newer.

For example to build on Ubuntu 20.04 we will support symbols much newer
than what that verison of the distro supports.
We could wrap all uses of newer features behind `#ifdef` checks, or
include the newest version of the loader that FEX itself supports.

The submodule is much cleaner and resolves the issue of drivers
supporting newer vulkan versions than libvulkan-dev.
Again super common with Nvidia blob users, kisak-ppa users, or people
updating mesa directly to be on the bleeding edge.
2022-08-04 22:21:22 -07:00
Ryan Houdek 4c2836f4b3 External: Adds Vulkan-Headers to external 2022-08-04 22:21:22 -07:00
Ryan Houdek a112db169c Merge pull request #1883 from Sonicadvance1/remove_static
cmake: Remove the static-pie compilation option
2022-08-04 20:13:28 -07:00
Ryan Houdek 0dc33ec893 cmake: Remove the static-pie compilation option
Due to glibc issues around static applications doing dlopen this is a
fundamentally broken option and no longer supported by FEX.

Remove the option entirely as to not be confusing.

We kept this around initially for chroot support, but with our RootFS
mounting AArch64 folders inside the chroot this isn't necessary anymore.
2022-08-04 19:22:11 -07:00
Ryan Houdek 0e93ba532f Merge pull request #1881 from lioncash/harness
HarnessHelpers: Handle SSE register offsets in CompareStates
2022-08-04 12:43:14 -07:00
lioncash 9e524a30d6 General: Resolve fmt deprecation warnings
fmt 9.0.0 deprecates implicit conversions of unscoped enum values to
integers to be consistent with scoped enum behavior.

Fairly trivial to resolve
2022-08-04 11:39:51 -04:00
lioncash ecca4e4abc Externals: Update fmt to 9.0.0
Keeps the codebase up to date with the latest major version (we were
previously on 8.1.1)

Changelog can be seen here: https://github.com/fmtlib/fmt/releases/tag/9.0.0
2022-08-04 11:20:55 -04:00
lioncash 51791e9efa HarnessHelpers: Handle SSE register offsets in CompareStates
Last of the things that slipped through while re-adding the ldp/stp
optimization.
2022-08-04 10:56:48 -04:00
Ryan Houdek 5eab087cc2 Merge pull request #1880 from neobrain/feature_thunk_x11_dependency
Thunks: Make GL guest thunks implicitly load libX11.so
2022-08-04 03:27:11 -07:00
Tony Wasserka cd05cdaa57 Thunks: Make GL guest thunks implicitly load libX11.so
Steam's gameoverlayrenderer.so relies on libX11 symbols to be available
without actually loading that library directly. This works on an unthunked
system since libGL.so depends on libGLX.so, which in turn pulls in libX11.so
at load-time. Adding a fake libX11 dependency to the libGL-guest thunks
reproduces this behavior.
2022-08-04 11:46:26 +02:00
Ryan Houdek f5e18ccea6 Merge pull request #1879 from lioncash/spill
Arm64Dispatcher: Amend memcpy in SpillSRA
2022-08-03 17:38:50 -07:00
lioncash 0da6b07770 Arm64Dispatcher: Fix signed/unsigned comparisons
Trivial change. We just need to use size_t instead of int here.
2022-08-03 11:23:40 -04:00
lioncash 05805356fb Arm64Dispatcher: Amend register memcpy in SpillSRA
Ensures that we spill to the correct register offset depending on the
execution mode.
2022-08-03 11:21:55 -04:00
Ryan Houdek a72ebfdec0 Merge pull request #1877 from Sonicadvance1/change_fexbash_ps1
FEXBash: Changes PS1 to hopefully help users
2022-08-01 16:31:33 -07:00
Ryan Houdek 59e5da9c7f FEXBash: Changes PS1 to hopefully help users
Currently FEXBash only outputs `FEXBash>` which gives weird docker
vibes. This can confuse users since they no longer see what folder they
are in.

Change this so it still has the user and path exposed like a typical PS1

eg: `FEXBash-ryanh@ryanh-TR2:/mnt/Work/Work/work/FEXNew/Build>`
2022-08-01 15:39:07 -07:00
Ryan Houdek cfd59db998 Merge pull request #1875 from Sonicadvance1/fexrootfsfetcher_runtime_porgram_checks
FEXRootFSFetcher: Adds runtime checks for image mounting tools
2022-08-01 14:24:36 -07:00
Ryan Houdek 1fa6bf1cc8 Merge pull request #1874 from Sonicadvance1/update_drm_v5.19
drm: Update to v5.19
2022-08-01 14:24:18 -07:00
Ryan Houdek a2f89de71e FEXRootFSFetcher: Adds runtime checks for image mounting tools
If the user doesn't have any of the tools necessary for handling FEX's
images then the tool would spuriously fail with `Couldn't parse rootfs definition URL.`
With zero indication as to why we removed images from the parsed json.

If the user has at least one of these tools installed then they won't
get this error message.
2022-07-31 18:34:56 -07:00
Ryan Houdek 36a27de286 drm: Update to v5.19
Not much changed here.
Some changes to msm which naturally work.
Some i915, and amdgpu churn with no functional change.
2022-07-31 15:45:37 -07:00
Ryan Houdek 9b685ba824 Merge pull request #1869 from Sonicadvance1/update_vixl2
vixl: Update
2022-07-30 11:37:10 -07:00
Ryan Houdek 8e9d5fe6ac vixl: Update
Removes three files from vixl that we aren't using. These take nearly
20 seconds apiece to compile so this just improves compile time
2022-07-29 16:09:44 -07:00
Ryan Houdek 89aa590615 Merge pull request #1866 from lioncash/used
Arm64/JIT: Remove unnecessary [[maybe_unused]] attributes
2022-07-28 12:59:40 -07:00
lioncash 3960e0f2a3 Arm64/JIT: Remove unnecessary [[maybe_unused]] attributes
These are always used inside the function via assignments.
2022-07-28 10:29:32 -04:00
Ryan Houdek c3c52d01d3 Merge pull request #1865 from neobrain/refactor_syscalls_cleanup
Cleanup syscalls headers
2022-07-28 03:20:10 -07:00
Tony Wasserka 0caef59e6a Syscalls: Simplify RegisterSyscall implementation 2022-07-28 11:23:23 +02:00
Tony Wasserka 99e6924ed1 Syscalls: Remove unneeded code 2022-07-28 11:22:39 +02:00
Tony Wasserka 431e1629b0 Syscalls: Clean up registry macros 2022-07-28 11:22:38 +02:00
Tony Wasserka ff9f702e52 Syscalls: Simplify CollectArgsFmtString implementation 2022-07-28 11:17:01 +02:00
Ryan Houdek 6903159f30 Merge pull request #1864 from FEX-Emu/remove_syscall_registration_vector
Syscalls: Removes staging vector usage
2022-07-28 01:41:26 -07:00
Ryan Houdek 504b7a03ad Syscalls: Removes staging vector usage
This removes about half a millisecond from syscall handler registration.
2022-07-28 01:29:26 -07:00
Ryan Houdek 6933c2aaff Merge pull request #1863 from neobrain/feature_thunk_xfree
Thunks/X11: Distinguish between host and guest pointers in XFree
2022-07-27 12:41:22 -07:00
Tony Wasserka 5fd00e31bb Thunks/X11: Distinguish between host and guest pointers in XFree
This function must be able to handle both guest heap pointers *and* host heap
pointers, so it only forwards to the native host library for the latter.

This is because Xlibint users allocate memory using internal macros aliasing
to libc's malloc but then they free using the function XFree. For libX11,
this is not a problem since the allocation happens in a thunked API function
(and hence on the host heap), but if a function from an unthunked library
accesses Xlibint, it will allocate on the guest heap.

One notable example where this was encountered is XF86VidModeGetAllModeLines.
2022-07-27 12:05:07 +02:00
Ryan Houdek a984674dae Merge pull request #1862 from Sonicadvance1/fix_tmp_pollution
JitSymbols: Only initialize perf map file if using
2022-07-26 22:47:00 -07:00
Ryan Houdek 8589119725 Merge pull request #1856 from lioncash/avx-ldp
Arm64Emitter: Re-add use of stp/ldp with hosts that don't support SVE2
2022-07-26 14:21:58 -07:00
lioncash 5e0205378b Allow skipping tests based on desired host features
Necessary for tests that depend on the state of the running context.

Since we support an SSE mode and an AVX mode, the FPR store truncate
test will fail on hosts that don't support AVX as the register offsets
are going to be different between the two. So we can conditionally
enable support for these tests.
2022-07-26 16:56:57 -04:00
lioncash bff2f2e5f9 Context: Expose ability to retrieve host features
Allows the test runner to determine whether or not to use AVX or SSE
state.
2022-07-26 16:56:57 -04:00
lioncash 5f9052a675 CoreState: Use a union for representing xmm data
In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)

However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
2022-07-26 16:56:54 -04:00
Ryan Houdek 8691b3964f JitSymbols: Only initialize perf map file if using
In most cases we aren't using JIT symbols but still creating the perf
map file.

Early check if we should generate the file or not, this way we stop
polluting the /tmp folder.
2022-07-25 14:21:14 -07:00
Ryan Houdek 4dfe0a0595 Merge pull request #1861 from Sonicadvance1/fix_close_range_hang
Linux: Fixes hang in close_range
2022-07-25 12:51:44 -07:00
Ryan Houdek 164299ce25 Merge pull request #1860 from Sonicadvance1/fexconfig_changes
FEXConfig: Some quality of life improvements
2022-07-25 10:43:39 -07:00
Ryan Houdek 9269ca6ebe FEXConfig: Rename Open Default to Open from default location 2022-07-25 10:29:37 -07:00
Ryan Houdek 93fe22e045 FEXConfig: Open the default config on application open.
User will most likely only care about the default configuration.
2022-07-25 10:28:30 -07:00
Ryan Houdek 3849278d5a FEXConfig: Add the ability to close with Ctrl+Q
Instead of only listening to the window close event
2022-07-25 10:26:11 -07:00
Ryan Houdek c0f976e200 FEXLinuxTests: Adds close_range test
Ensures we never get this hang again.
2022-07-25 10:17:43 -07:00
Ryan Houdek 3143749f64 Linux: Fixes hang in close_range
It's common for Linux applications that use close_range to pass in ~0 as
the last FD. This was causing FEX to spin from [2, ~0U] in this loop.
This would take /forever/ to run.

Change over to an ordered map and use the map's range searching and
ranged erase to more quickly remove these elements.

Fixes a hang that occurs with first time Steam setup in
steam-linux-runtime heavy application `steam-runtime-identify-library-abi`.
2022-07-24 19:30:28 -07:00
Ryan Houdek 1704805e32 FEXConfig: Some quality of life improvements
Removes the "Load Default Options" menu option. This option was
confusing for new users and isn't necessary anymore.

Fixes the "Load Default" option so it actually populates the full
configuration layer in the face of partial configuration.

This is a /very/ common use case for new users that ran through
FEXRootFSFetcher, where the only configuration set is the RootFS.
The configuration would be visually confusing since the visual
representation for missing options wouldn't reflect their default
configuration state. "TSO Enabled" is an example where it would appear
disabled in the GUI, but it is default enabled.

Also fixes the issue that the default configuration window would just be
a 320x240 floating window in the center of the screen. This is due to
the window being a floating sub window in the dockspace by default, and
not docked.

Instead just remove the dockspace, it isn't serving us any purpose.
This means the child configuration window now maximizes to the window
size which is the desired behaviour from default.

Additionally only save the config file once. While the msg dialog is
open (2 seconds while it is open, or escape to make it go away
immediately) the program won't save the file again. This fixes an issue
that if you used the shortcut key to save the file, it would save the
file at the refresh rate of your screen. Which is 144hz on my setup, so
it spams my filesystem quite heavily.
2022-07-24 16:07:45 -07:00
Ryan Houdek 3aabe07720 Merge pull request #1859 from Sonicadvance1/fix_fexrootfs_option
FEXRootFSFetcher: Actually wire up -a -x
2022-07-24 11:59:53 -07:00
Ryan Houdek c5815ab59e FEXRootFSFetcher: Actually wire up -a -x
Missed this since I tested the erofs path which only supports -a
2022-07-24 11:48:11 -07:00
Ryan Houdek f6fcfbabd5 Merge pull request #1858 from Sonicadvance1/add_fexrootfsfetch_options
FEXRootFSFetcher: Add some options for automation without user intervention
2022-07-23 20:56:00 -07:00
Ryan Houdek a0835c95e4 FEXRootFSFetcher: Add some options for automation without user intervention
Main options here are `-y` and `-a`.

Passing both will allow the user to download the exact match image, use
it as-is, and set the config to use the rootfs by default.

Additional options `--distro-name` and `--distro-option` allows one to
select a particular distro version.

eg: `FEXRootFSFetcher -y -a --distro-name=ubuntu --distro-version=22.04`
* Will download the first Ubuntu 22.04 compressed in the json list
** Priority of erofs or squashfs depends on order in json and what the
host supports
2022-07-23 14:57:58 -07:00
Ryan Houdek ce0c24eee9 Merge pull request #1855 from Sonicadvance1/fix_wine_telemetry
Telemetry: Support executable names through wineserver
2022-07-22 15:02:04 -07:00
Mai c9f0ecb1e8 Merge pull request #1854 from Sonicadvance1/determine_x1c
CPUID: Detect Cortex-X1C
2022-07-22 17:22:26 -04:00
Ryan Houdek 50febb3459 Telemetry: Support executable names through wineserver
While we were getting the application name for the application layer, we
were failing to store the filename for telemetry.

Save the filename we get for application layers and store it for the
telemetry file.

Otherwise these were just alway ending up as wine or wine-preloader.
2022-07-21 20:12:36 -07:00
Ryan Houdek 601b0f96b1 CPUID: Detect Cortex-X1C
Missed this in the list before since we didn't know the part ID.
2022-07-20 23:20:56 -07:00
Ryan Houdek 018661609a Merge pull request #1841 from Sonicadvance1/restrict_syscall_install_by_arch
Linux: Only install syscall handlers for the arch we launched with
2022-07-20 10:19:29 -07:00
Stefanos Kornilios Mitsis Poiitidis 98d935d972 Merge pull request #1852 from lioncash/json
json_ir_generator: Remove Args() functions from IR structs
2022-07-20 14:03:58 +00:00
Ryan Houdek b07660c4fa Merge pull request #1851 from Sonicadvance1/fix_soma_and_sa_mask
Fix SOMA and sigaction definition
2022-07-19 13:06:45 -07:00
Ryan Houdek 8037231376 Merge pull request #1850 from Sonicadvance1/fix_faulting_instructions
FEXCore: Fix-up edge case behaviour on faulting instructions
2022-07-19 10:23:12 -07:00
Ryan Houdek 6e422a8691 Signals: Steal back the XID signal from glibc
glibc lazily initializes the SETXID signal handler until first thread
creation.
Once we create our first pthread, steal it back from GLIBC after the
fact.

Fixes SOMA again.
2022-07-19 10:18:47 -07:00
Ryan Houdek 3d347ed565 Merge pull request #1840 from Sonicadvance1/add_assume_assert
FEXCore: Adds assume optimizing LogManager function
2022-07-19 10:13:21 -07:00
Ryan Houdek 13446d8b6b Linux: Only install syscall handlers for the arch we launched with
We were registering syscall handlers to the temporary working vectors for
both x86 and x86_64 regardless of which bitness we launched with.

We then only installed the syscall handlers depending on bitness.
This was burning a decent amount of time and some memory on
initialization. Instead just don't register syscall handlers on the
other bitness.

Also stop wasting memory, clear the syscall registration vector after
we've registered everything.

Saves about 28KB of memory per process.
2022-07-19 10:12:09 -07:00
lioncash 08c34bf573 json_ir_generator: Remove Args() functions from IR structs
Since the IR argument can be accessed directly by name, this is no
longer needed as a shorthand.
2022-07-19 11:23:13 -04:00
Mai 3a64ea1650 Merge pull request #1849 from Sonicadvance1/global_main_config
Config: Support a global configuration file
2022-07-19 11:10:49 -04:00
Ryan Houdek db3439b61b Signals: Fixes kernel definition of restorer in sigaction
Even though this member isn't used on ARM, the member is still there.

Fixes crashing in optimized builds and also makes unity's garbage
collector significantly more stable
2022-07-18 16:50:02 -07:00
Ryan Houdek 7814be7467 Update IR tests for new break op 2022-07-17 20:56:31 -07:00
Ryan Houdek 8a7e0432f1 FEXLinuxTests: Adds unit tests for the remaining five fault instructions
We already have one for `into` so we just need to cover the remaining
five instructions.
2022-07-17 20:56:31 -07:00
Ryan Houdek ff1d51c7bd FEXCore: Fix-up edge case behaviour on faulting instructions
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.

One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.

Additionally some of the metadata generated around the signal delegation
wasn't correct.

With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.

I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.

With this change, Elden Ring is now stabilized and works under FEX.
2022-07-17 20:56:31 -07:00
Ryan Houdek f2602e3136 Merge pull request #1824 from Sonicadvance1/into_test
unittests: Adds 32-bit into test
2022-07-17 13:12:59 -07:00
Ryan Houdek 313ce0fed8 unittests: Adds 32-bit into test
This is a very basic test to check for overflow exception.
`into` instruction only exists on 32-bit, so it needs to not be built as
a 64-bit executable.
2022-07-17 12:55:00 -07:00
Ryan Houdek ba0887defa Misc: Convert assert logs to assume+assert that can be
Most of these won't make a performance difference. But we should be
using the assume version everywhere we can.
2022-07-17 12:51:43 -07:00
Ryan Houdek 063841f491 FEXCore: Adds assume optimizing LogManager function
This comes in the form of a new define and a new function.

Assuming assert doesn't quite cover 100% of our asserting logging use
cases so we need to break this out.
If the compiler can't see through side effects at compile time for the
predicate then it'll throw a warning.

This gives us a nice behaviour. In the case that asserts are enabled, we
fall down the assert checking path. So if an assert fails to predicate
we will do an assert log like normal.
The change comes when we build release without asserts, we use the
`__builtin_assume` definition to allow the compiler to optimize around
assumptions. Since these aren't runtime asserts in release build, this
gives us some small optimizations in various locations.

Clang will very specifically do additional optimizations when you
provide it `__builtin_assume` directives that can really be worth it.
2022-07-17 12:50:51 -07:00
Ryan Houdek 4aa5c78150 Config: Support a global configuration file
By default this file ends up in `/usr/share/fex-emu/Config.json`.
2022-07-17 11:03:46 -07:00
Ryan Houdek c1688fa392 Merge pull request #1848 from wannacu/main
improve compile ability for older linux
2022-07-14 23:18:48 -07:00
wannacu 0dd03e9cb6 Add pidfd_open syscall helpers 2022-07-15 13:44:07 +08:00
wannacu 8d60d70553 FEXServer: Use compatible syscall helpers 2022-07-15 13:42:36 +08:00
Ryan Houdek c17340547c Merge pull request #1789 from lioncash/avx-impl
AVX initial groundwork
2022-07-13 14:12:53 -07:00
lioncash 34aa6bf6c7 CoreState: Move flag variables into padding
Allows for a smaller struct size.

This movement of the struct members also requires us to modify the
DeadContextStorePass to take the new locations into account.

On the plus side, we get to remove all of the padding bits in the
CPUState struct, so we can remove the handling for them.
2022-07-13 16:57:16 -04:00
lioncash 82a6ce63b4 Dispatcher: Make use of AVX confix to determine stack sizes
Now that we have the AVX config option in place, we can use it to
determine how to set up the stack for storing and loading all of the
necessary state we need.
2022-07-13 14:55:04 -04:00
Ryan Houdek 3ac6ba0fe2 Merge pull request #1846 from lioncash/x86ir
x86_64: Migrate args over to named IR arguments
2022-07-13 11:39:51 -07:00
lioncash 8b0fb4fa19 HostFeatures: Add querying for AVX support
If we have SVE2 support and a vector length of at least 256, then we can
adequately handle AVX.

However we also add in the ability for application profiles to disable
AVX if necessary for any reason.

Not currently used, but will be in subsequent changes.
2022-07-13 14:30:31 -04:00
lioncash 9d437b8863 CoreState: Expand xmm registers
Expands them to add the high lanes added in AVX
2022-07-13 14:29:48 -04:00
lioncash f72469c80a Dispatcher: Handle restoring ymm high lanes
On Linux the bottom 12 bytes of the legacy FXSAVE context are used to
encode information about any extension blocks that may follow it.

We can use the same approach to encode our high lanes of the AVX
YMM registers into memory and vice-versa.
2022-07-13 14:29:48 -04:00
lioncash eec7972808 UContext: Add definitions for accessing XSAVE data
Adds the layout of extended state for accessing/saving ymm high lanes
2022-07-13 14:29:48 -04:00
Lioncache fc9343e7a4 x86_64/VectorOps: Correct VRev64 op cast
This was using the structure for VDupElement, but happened
to work because the layouts are the same.
2022-07-13 13:46:16 -04:00
Lioncache be8e353d12 x86_64/VectorOps: Move args names over to IR names 2022-07-13 13:45:45 -04:00
Lioncache 53217acdd0 x86_64/MoveOps: Move args names over to IR names 2022-07-13 13:16:37 -04:00
Lioncache c1e0ea5e0a x86_64/MiscOps: Move args names over to IR names 2022-07-13 13:15:16 -04:00
Lioncache 467afc7248 x86_64/FlagOps: Move args names over to IR names 2022-07-13 13:13:37 -04:00
Lioncache 62753071f4 x86_64/EncryptionOps: Move args names over to IR names 2022-07-13 13:13:08 -04:00
Lioncache de32de9edd x86_64/ConversionOps: Move args names over to IR names 2022-07-13 13:10:18 -04:00
Lioncache aa7d2c3f7e x86_64/BranchOps: Move args names over to IR names 2022-07-13 13:07:48 -04:00
Lioncache 2a1dd68583 x86_64/ALUOps: Move args names over to IR names 2022-07-13 13:06:30 -04:00
Ryan Houdek 80909eaa84 Merge pull request #1845 from Sonicadvance1/support_block_rip_synchronization
FEXCore: Support synchronizing RIP on block entry through config
2022-07-13 09:54:24 -07:00
Ryan Houdek c6a557abe2 AppConfig: Add VC redistributable configs
These require the previous commit's RIP block synchronization.

This is due to the fact that these rely on exceptions to occur on
try-catch blocks.

With us ensuring RIP synchronization on block entry, this solves the
problem of these crashing before doing anything.

Wine 6.x didn't require this, but 7.x does.

Fixes Proton Experimental hangs when it is installing the VC runtimes.
2022-07-13 08:49:13 -07:00
Ryan Houdek 8d7b8729f1 FEXCore: Support synchronizing RIP on block entry through config
If an application is doing long jumps on exceptions then we need to
ensure that RIP is synchronized at least to block entry for some amount
of safety.

Due to our block linking which doesn't ensure that RIP is synchronized
on block entry, our exception handling wouldn't see a RIP change, thus
not going down the path that we jump to Dispatcher loop top on RIP
change.

Burn a couple of instructions on block entry to ensure that RIP
synchronized but only on config.
2022-07-13 08:49:13 -07:00
Ryan Houdek 3f9a1c3751 Merge pull request #1844 from lioncash/interp
Interpreter: Move argument names over to IR names
2022-07-12 10:10:25 -07:00
Ryan Houdek 0dfe617d96 Merge pull request #1843 from Sonicadvance1/fix_sra_signal_handler_config
Dispatcher: Fix SRA enabled check in signal delegator handlers
2022-07-10 11:53:55 -07:00
Ryan Houdek 235ad441b4 Dispatcher: Fix SRA enabled check in signal delegator handlers
When some SRA code was being shuffled around, this config variable was
never being set. This caused the guest signal handlers to never expect
SRA to be enabled.

Moves the Dispatcher config object to the actual dispatcher rather than
having it in the arch specific dispatcher class. Then have the signal
delegator check the config directly instead of having this secondary
config value.

Fixes vc_redist_x64 and vc_redist_x86. Probably also fixes a bunch of
other random Proton/Wine things that rely on signal long jump.
2022-07-10 11:39:25 -07:00
lioncash 5cd0f6e7e7 Interpreter/VectorOps: Move argument names over to IR names 2022-06-17 15:31:10 -04:00
lioncash 2627ed3193 Interpreter/MoveOps: Move argument names over to IR names 2022-06-17 14:31:01 -04:00
lioncash e3bb8b43ab Interpreter/MiscOps: Move argument names over to IR names 2022-06-17 14:30:00 -04:00
lioncash 8ba6a3bda9 Interpreter/MemoryOps: Move argument names over to IR names 2022-06-17 14:28:44 -04:00
lioncash d288249609 Interpreter/FlagOps: Move argument names over to IR names 2022-06-17 14:25:39 -04:00
lioncash 33d7c6be9e Interpreter/F80Ops: Move argument names over to IR names 2022-06-17 14:25:09 -04:00
lioncash c13acc5b60 Interpreter/EncryptionOps: Move argument names over to IR names 2022-06-17 14:14:46 -04:00
lioncash fc9be28d51 Interpreter/ConversionOps: Move argument names over to IR names 2022-06-17 14:12:09 -04:00
lioncash 944f400d19 Interpreter/BranchOps: Move argument names over to IR names 2022-06-17 14:08:43 -04:00
lioncash 5aee30a7b5 Interpreter/ALUOps: Move argument names over to IR names 2022-06-17 14:06:18 -04:00
509 changed files with 26064 additions and 11178 deletions

No files matched your search

+46 -3
View File
@@ -40,7 +40,6 @@ jobs:
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
@@ -65,7 +64,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -178,6 +177,51 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXLinuxTests.log || true
- name: Thunkgen tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target thunkgen_tests
- name: Thunkgen Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkgenTests.log || true
- name: Install
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: Test GL No-Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DISPLAY: ":0"
run: cmake --build . --config $BUILD_TYPE --target thunk_functional_tests_nothunks
- name: No thunks Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_NoThunkResults.log || true
- name: Test GL Thunks
if: matrix.arch[1] == 'x64'
working-directory: ${{runner.workspace}}/build
shell: bash
env:
DISPLAY: ":0"
run: cmake --build . --config $BUILD_TYPE --target thunk_functional_tests_thunks
- name: Thunks Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ThunkResults.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
@@ -197,4 +241,3 @@ jobs:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+118
View File
@@ -0,0 +1,118 @@
name: Vixl Simulator run
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
# Only the x86-64 runner is fast enough to run this
arch: [[self-hosted, x64], [self-hosted, ARMv8.4]]
fail-fast: false
steps:
- uses: actions/checkout@v2
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: ASM Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v2'
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+1
View File
@@ -10,3 +10,4 @@ out/
.vscode/
.vs/
*.pyc
.cache
+4
View File
@@ -49,3 +49,7 @@
shallow = true
path = External/robin-map
url = https://github.com/Tessil/robin-map.git
[submodule "External/Vulkan-Headers"]
shallow = true
path = External/Vulkan-Headers
url = https://github.com/KhronosGroup/Vulkan-Headers.git
+5
View File
@@ -0,0 +1,5 @@
{
"ThunksDB": {
"GL": 1
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"ThunksDB": {
"Vulkan": 1
}
}
+21
View File
@@ -0,0 +1,21 @@
if(NOT EXISTS "@CMAKE_BINARY_DIR@/install_manifest.txt")
message(FATAL_ERROR "Cannot find install manifest: @CMAKE_BINARY_DIR@/install_manifest.txt")
endif()
file(READ "@CMAKE_BINARY_DIR@/install_manifest.txt" files)
string(REGEX REPLACE "\n" ";" files "${files}")
foreach(file ${files})
message(STATUS "Uninstalling $ENV{DESTDIR}${file}")
if(IS_SYMLINK "$ENV{DESTDIR}${file}" OR EXISTS "$ENV{DESTDIR}${file}")
exec_program(
"@CMAKE_COMMAND@" ARGS "-E remove \"$ENV{DESTDIR}${file}\""
OUTPUT_VARIABLE rm_out
RETURN_VALUE rm_retval
)
if(NOT "${rm_retval}" STREQUAL 0)
message(FATAL_ERROR "Problem when removing $ENV{DESTDIR}${file}")
endif()
else(IS_SYMLINK "$ENV{DESTDIR}${file}" OR EXISTS "$ENV{DESTDIR}${file}")
message(STATUS "File $ENV{DESTDIR}${file} does not exist.")
endif()
endforeach()
+83 -142
View File
@@ -7,6 +7,7 @@ CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
@@ -20,7 +21,6 @@ option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JI
option(ENABLE_VISUAL_DEBUGGER "Enables the visual debugger for compiling" FALSE)
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
option(ENABLE_WERROR "Enables -Werror" FALSE)
option(ENABLE_STATIC_PIE "Enables static-pie build" FALSE)
option(ENABLE_JEMALLOC "Enables jemalloc allocator" TRUE)
option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
@@ -28,10 +28,36 @@ option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_INTERPRETER "Enables FEX's Interpreter" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
set (X86_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86.cmake" CACHE FILEPATH "Toolchain file for the x86 (cross-)compiler")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
if (ENABLE_FEXCORE_PROFILER)
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
if (FEXCORE_PROFILER_BACKEND STREQUAL "GPUVIS")
add_definitions(-DFEXCORE_PROFILER_BACKEND=1)
else()
message(FATAL_ERROR "Unknown FEXCore profiler backend ${FEXCORE_PROFILER_BACKEND}")
endif()
endif()
# uninstall target
if(NOT TARGET uninstall)
configure_file(
"${CMAKE_CURRENT_SOURCE_DIR}/CMakeFiles/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_BINARY_DIR}/CMakeFiles/cmake_uninstall.cmake"
IMMEDIATE @ONLY)
add_custom_target(uninstall
COMMAND ${CMAKE_COMMAND} -P ${CMAKE_CURRENT_BINARY_DIR}/CMakeFiles/cmake_uninstall.cmake)
endif()
# These options are meant for package management
set (TUNE_CPU "native" CACHE STRING "Override the CPU the build is tuned for")
set (TUNE_ARCH "generic" CACHE STRING "Override the Arch the build is tuned for")
@@ -85,10 +111,9 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
set (X86_TOOLCHAIN_FILE "")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
add_definitions(-D_M_ARM_64=1)
endif()
@@ -140,130 +165,6 @@ if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
set(ENABLE_JEMALLOC FALSE)
endif()
if (ENABLE_STATIC_PIE)
if (_M_ARM_64 AND ENABLE_LLD)
message (FATAL_ERROR "Static linking does not currently work with AArch64+lld. Use GNU ld for now.")
endif()
file(WRITE ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt.c
"int main(int argc, char* argv[])
{
return 0;
}")
# Compile the test application with our LD_OVERRIDE and static-pie options
try_compile(
COMPILE_RESULT
${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp
${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt.c
COMPILE_DEFINITIONS "-fPIE ${LD_OVERRIDE}"
LINK_LIBRARIES "-static-pie ${LD_OVERRIDE}"
COPY_FILE ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt
)
if (${COMPILE_RESULT})
# Read the symbols from the elf
execute_process(COMMAND
readelf -s ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/Determine_iplt
OUTPUT_FILE ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/plt_out.txt
OUTPUT_VARIABLE PLT_SYMBOLS)
# Pull out the __rela_iplt_{start,end} symbols if they exist
execute_process(COMMAND
"grep" "__rela_iplt" ${PROJECT_BINARY_DIR}/CMakeFiles/CMakeTmp/plt_out.txt
OUTPUT_VARIABLE PLT_SYMBOLS)
set (SYMBOLS_FINE TRUE)
set (HAS_IPLT -1)
# Check if we have any symbols in our grep output
# The symbols must either not exist at all OR the symbols are zero
if (PLT_SYMBOLS)
string(FIND ${PLT_SYMBOLS} "__rela_iplt_start" HAS_IPLT)
endif()
if (NOT HAS_IPLT EQUAL -1)
# We have some symbols from readelf. Let's parse the results to check if they are zero
# Format: '35: 0000000000000000 0 NOTYPE LOCAL HIDDEN UND __rela_iplt_start'
string(REPLACE "\n" ";" SYMBOL_LIST ${PLT_SYMBOLS})
foreach (SYMBOL ${SYMBOL_LIST})
# strip any leading and trailing whitespace
string (STRIP ${SYMBOL} SYMBOL)
# Convert string to a list
string(REPLACE " " ";" SYMBOL_VALUES ${SYMBOL}})
# Pull out the address argument
list(GET SYMBOL_VALUES 1 OFFSET)
# Check against integer zero
if (NOT ${OFFSET} EQUAL 0)
# Symbol wasn't zero, this now fails
set (SYMBOLS_FINE FALSE)
endif()
endforeach()
endif()
if (SYMBOLS_FINE)
# We can now exnable static-pie
set (STATIC_PIE_OPTIONS "-static-pie")
# Pthreads has an issue with exposing symbols
# We need to make some concessions to the pthread gods
if (ENABLE_LLD)
set (PTHREAD_LIB
-Wl,--undefined-glob=pthread_*
-Wl,--undefined=__cxa_finalize
-Wl,--undefined=_pthread_cleanup_push_defer
-Wl,--undefined=_pthread_cleanup_pop_restore
-Wl,--undefined=__pthread_cleanup_upto
pthread)
else()
set (PTHREAD_LIB
-Wl,--undefined=pthread_join
-Wl,--undefined=pthread_attr_getdetachstate
-Wl,--undefined=pthread_sigmask
-Wl,--undefined=pthread_mutex_lock
-Wl,--undefined=pthread_cond_init
-Wl,--undefined=pthread_attr_init
-Wl,--undefined=pthread_mutex_unlock
-Wl,--undefined=pthread_mutexattr_destroy
-Wl,--undefined=pthread_detach
-Wl,--undefined=pthread_mutex_init
-Wl,--undefined=pthread_getattr_np
-Wl,--undefined=pthread_cond_timedwait
-Wl,--undefined=pthread_attr_destroy
-Wl,--undefined=pthread_mutexattr_settype
-Wl,--undefined=pthread_rwlock_unlock
-Wl,--undefined=pthread_rwlock_wrlock
-Wl,--undefined=pthread_setspecific
-Wl,--undefined=pthread_create
-Wl,--undefined=pthread_cond_clockwait
-Wl,--undefined=pthread_key_create
-Wl,--undefined=pthread_rwlock_rdlock
-Wl,--undefined=pthread_setname_np
-Wl,--undefined=pthread_cond_signal
-Wl,--undefined=pthread_mutexattr_init
-Wl,--undefined=pthread_attr_setstack
-Wl,--undefined=pthread_self
-Wl,--undefined=pthread_getaffinity_np
-Wl,--undefined=pthread_cond_wait
-Wl,--undefined=pthread_mutex_trylock
-Wl,--undefined=pthread_cond_broadcast
-Wl,--undefined=pthread_cond_destroy
-Wl,--undefined=pthread_getspecific
-Wl,--undefined=pthread_key_delete
-Wl,--undefined=pthread_once
-Wl,--undefined=__cxa_finalize
-Wl,--undefined=_pthread_cleanup_push_defer
-Wl,--undefined=_pthread_cleanup_pop_restore
-Wl,--undefined=__pthread_cleanup_upto
pthread)
endif()
else()
message (FATAL_ERROR "Application has __rela_iplt_{start,end} symbols. Which means static-pie can't be enabled")
endif()
else()
message (FATAL_ERROR "Couldn't compile static-pie test. Static-pie can't be enabled! Is your glibc compiled without static-pie?")
endif()
endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
@@ -488,8 +389,6 @@ if (BUILD_TESTS)
endif()
add_subdirectory(FEXHeaderUtils/)
include_directories(FEXHeaderUtils/)
add_subdirectory(External/FEXCore)
# Binfmt_misc files must be installed prior to Source/ installs
@@ -499,15 +398,20 @@ add_subdirectory(Source/)
add_subdirectory(Data/AppConfig/)
# Install the ThunksDB file
install(
FILES ${CMAKE_CURRENT_SOURCE_DIR}/Data/ThunksDB.json
DESTINATION ${DATA_DIRECTORY}/)
file(GLOB CONFIG_SOURCES CONFIGURE_DEPENDS ${CMAKE_CURRENT_SOURCE_DIR}/Data/*.json)
# Any application configuration json file gets installed
foreach(CONFIG_SRC ${CONFIG_SOURCES})
install(FILES ${CONFIG_SRC}
DESTINATION ${DATA_DIRECTORY}/)
endforeach()
if (BUILD_TESTS)
add_subdirectory(unittests/)
endif()
if (BUILD_THUNKS)
set (FEX_PROJECT_SOURCE_DIR ${PROJECT_SOURCE_DIR})
add_subdirectory(ThunkLibs/Generator)
# Thunk targets for both host libraries and IDE integration
@@ -523,10 +427,31 @@ if (BUILD_THUNKS)
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest"
CMAKE_ARGS
"-DBITNESS=64"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_TOOLCHAIN_FILE}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_64_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
DEPENDS thunkgen
)
ExternalProject_Add(guest-libs-32
PREFIX guest-libs-32
SOURCE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/ThunkLibs/GuestLibs"
BINARY_DIR "Guest_32"
CMAKE_ARGS
"-DBITNESS=32"
"-DCMAKE_BUILD_TYPE=${CMAKE_BUILD_TYPE}"
"-DENABLE_CLANG_THUNKS=${ENABLE_CLANG_THUNKS}"
"-DCMAKE_TOOLCHAIN_FILE:FILEPATH=${X86_32_TOOLCHAIN_FILE}"
"-DCMAKE_INSTALL_PREFIX=${CMAKE_INSTALL_PREFIX}"
"-DSTRUCT_VERIFIER=${CMAKE_SOURCE_DIR}/Scripts/StructPackVerifier.py"
"-DFEX_PROJECT_SOURCE_DIR=${FEX_PROJECT_SOURCE_DIR}"
"-DGENERATOR_EXE=$<TARGET_FILE:thunkgen>"
INSTALL_COMMAND ""
BUILD_ALWAYS ON
@@ -541,6 +466,28 @@ if (BUILD_THUNKS)
)"
DEPENDS guest-libs
)
install(
CODE "MESSAGE(\"-- Installing: guest-libs-32\")"
CODE "
EXECUTE_PROCESS(COMMAND ${CMAKE_COMMAND} --build . --target install
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32
)"
DEPENDS guest-libs-32
)
add_custom_target(uninstall_guest-libs
COMMAND ${CMAKE_COMMAND} "--build" "." "--target" "uninstall"
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest
)
add_custom_target(uninstall_guest-libs-32
COMMAND ${CMAKE_COMMAND} "--build" "." "--target" "uninstall"
WORKING_DIRECTORY ${CMAKE_BINARY_DIR}/Guest_32
)
add_dependencies(uninstall uninstall_guest-libs)
add_dependencies(uninstall uninstall_guest-libs-32)
endif()
set(FEX_VERSION_MAJOR "0")
@@ -593,15 +540,9 @@ endif()
# Package creation
set (CPACK_GENERATOR "DEB")
if (ENABLE_STATIC_PIE)
set (CPACK_PACKAGE_NAME fex-emu-static)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu")
else()
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_DEBIAN_PACKAGE_CONFLICTS "fex-emu-static")
endif()
set (CPACK_PACKAGE_NAME fex-emu)
set (CPACK_PACKAGE_FILE_NAME "${CPACK_PACKAGE_NAME}-${GIT_DESCRIBE_STRING}_${CMAKE_SYSTEM_PROCESSOR}")
set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.org>")
set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.com>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
+1 -1
View File
@@ -55,7 +55,7 @@ further defined and clarified by project maintainers.
## Enforcement
Instances of abusive, harassing, or otherwise unacceptable behavior may be
reported by contacting the project team at team@fex-emu.org. All
reported by contacting the project team at team@fex-emu.com. All
complaints will be reviewed and investigated and will result in a response that
is deemed necessary and appropriate to the circumstances. The project team is
obligated to maintain confidentiality with regard to the reporter of an incident.
+3 -3
View File
@@ -11,15 +11,15 @@ endforeach()
# First generate then install it
foreach(GEN_CONFIG_SRC ${GEN_CONFIG_SOURCES})
# Get the filename only component
get_filename_component(CONFIG_NAME ${GEN_CONFIG_SRC} NAME_WE)
get_filename_component(CONFIG_NAME ${GEN_CONFIG_SRC} NAME_WLE)
# Configure it
configure_file(
${GEN_CONFIG_SRC}
${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json)
${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME})
# Then install the configured json
install(
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}.json
FILES ${CMAKE_BINARY_DIR}/Data/AppConfig/${CONFIG_NAME}
DESTINATION ${DATA_DIRECTORY}/AppConfig/)
endforeach()
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"Config": {
"x86dec_SynchronizeRIPOnAllBlocks": "1"
}
}
+8
View File
@@ -165,6 +165,14 @@
"@PREFIX_LIB@/x86_64-linux-gnu/libXfixes.so.3.1.0"
]
},
"OpenCL": {
"Library" : "libOpenCL-guest.so",
"Overlay": [
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so",
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so.1",
"@PREFIX_LIB@/x86_64-linux-gnu/libOpenCL.so.1.0.0"
]
},
"":{}
}
}
+12 -6
View File
@@ -9,12 +9,19 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
endif()
set(ENABLE_JIT_X86_64 ${_M_X86_64} CACHE BOOL "Enable the x86_64 JIT")
set(ENABLE_JIT_ARM64 ${_M_ARM_64} CACHE BOOL "Enable the ARM64 JIT")
if (ENABLE_VIXL_SIMULATOR)
# If the vixl simulator is enabled then we are using the ARM64 JIT
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" FALSE)
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" TRUE)
else()
option(ENABLE_JIT_X86_64 "Enable the x86_64 JIT" ${_M_X86_64})
option(ENABLE_JIT_ARM64 "Enable the ARM64 JIT" ${_M_ARM_64})
endif()
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
@@ -27,7 +34,6 @@ set(CMAKE_INCLUDE_CURRENT_DIR ON)
include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
@@ -46,14 +52,14 @@ if (OVERRIDE_VERSION STREQUAL "detect")
if (GIT_FOUND)
execute_process(
COMMAND ${GIT_EXECUTABLE} rev-parse --short HEAD
COMMAND ${GIT_EXECUTABLE} rev-parse --short=7 HEAD
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_SHORT_HASH
ERROR_QUIET
OUTPUT_STRIP_TRAILING_WHITESPACE
)
execute_process(
COMMAND ${GIT_EXECUTABLE} describe
COMMAND ${GIT_EXECUTABLE} describe --abbrev=7
WORKING_DIRECTORY "${CMAKE_SOURCE_DIR}"
OUTPUT_VARIABLE GIT_DESCRIBE_STRING
ERROR_QUIET
-10
View File
@@ -321,8 +321,6 @@ def print_ir_structs(defines):
if op.SSAArgNum > 0:
# Add helpers for accessing SSA arguments, given how frequently they're accessed
output_file.write("\t// Get index of argument by name\n")
SSAArg = 0
for arg in op.Arguments:
@@ -330,14 +328,6 @@ def print_ir_structs(defines):
output_file.write("\tstatic constexpr size_t {}_Index = {};\n".format(arg.Name, SSAArg))
SSAArg = SSAArg + 1
output_file.write("\n")
output_file.write("\t[[nodiscard]] OrderedNodeWrapper& Args(size_t Index) {\n")
output_file.write("\t\treturn Header.Args[Index];\n")
output_file.write("\t}\n")
output_file.write("\t[[nodiscard]] const OrderedNodeWrapper& Args(size_t Index) const {\n")
output_file.write("\t\treturn Header.Args[Index];\n")
output_file.write("\t}\n")
output_file.write("};\n")
+10 -5
View File
@@ -135,7 +135,6 @@ set (SRCS
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/StaticRegisterAllocationPass.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
@@ -143,6 +142,7 @@ set (SRCS
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
)
if (ENABLE_INTERPRETER)
@@ -177,6 +177,11 @@ if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
endif()
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
list(APPEND DEFINES -DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_JIT_X86_64)
list(APPEND SRCS
Interface/Core/JIT/x86_64/JIT.cpp
@@ -213,7 +218,7 @@ if (ENABLE_JIT_ARM64)
)
endif()
set (LIBS vixl dl xxhash tiny-json)
set (LIBS fmt::fmt vixl dl xxhash tiny-json FEXHeaderUtils)
if (ENABLE_JEMALLOC)
list (APPEND LIBS FEX_jemalloc)
endif()
@@ -359,14 +364,14 @@ endfunction()
# Build FEXCore_Config static library
add_library(FEXCore_Base STATIC ${FEXCORE_BASE_SRCS})
target_link_libraries(FEXCore_Base fmt::fmt tiny-json)
target_link_libraries(FEXCore_Base ${LIBS})
AddDefaultOptionsToTarget(FEXCore_Base)
function(AddObject Name Type)
add_library(${Name} ${Type} ${SRCS})
add_dependencies(${Name} IR_INC)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
target_link_libraries(${Name} FEXCore_Base)
AddDefaultOptionsToTarget(${Name})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
@@ -374,7 +379,7 @@ endfunction()
function(AddLibrary Name Type)
add_library(${Name} ${Type} $<TARGET_OBJECTS:${PROJECT_NAME}_object>)
target_link_libraries(${Name} FEXCore_Base ${LIBS})
target_link_libraries(${Name} FEXCore_Base)
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
AddDefaultOptionsToTarget(${Name})
+3 -3
View File
@@ -20,12 +20,12 @@ struct BitSet final {
ElementType *Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
LOGMAN_THROW_A_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
void Free() {
@@ -64,7 +64,7 @@ struct BitSetView final {
ElementType *Memory;
void GetView(BitSet<T> &Set, uint64_t ElementOffset) {
LOGMAN_THROW_A_FMT((ElementOffset % MinimumSize) == 0,
LOGMAN_THROW_AA_FMT((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
+5 -2
View File
@@ -7,6 +7,11 @@
namespace FEXCore {
JITSymbols::JITSymbols() : fp{nullptr, std::fclose} {
}
JITSymbols::~JITSymbols() = default;
void JITSymbols::InitFile() {
const auto PerfMap = fmt::format("/tmp/perf-{}.map", getpid());
fp.reset(fopen(PerfMap.c_str(), "wb"));
@@ -16,8 +21,6 @@ namespace FEXCore {
}
}
JITSymbols::~JITSymbols() = default;
void JITSymbols::Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (!fp) return;
+1
View File
@@ -11,6 +11,7 @@ public:
JITSymbols();
~JITSymbols();
void InitFile();
void Register(const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
+32 -18
View File
@@ -82,7 +82,7 @@ namespace JSON {
json_t const* ConfigList = json_getProperty(json, "Config");
if (!ConfigList) {
LogMan::Msg::EFmt("Couldn't get config list");
// This is a non-error if the configuration file exists but no Config section
return;
}
@@ -154,15 +154,20 @@ namespace JSON {
return ConfigDir;
}
std::string GetConfigFileLocation() {
std::string GetConfigFileLocation(bool Global) {
std::string ConfigFile{};
const char *AppConfig = getenv("FEX_APP_CONFIG");
if (AppConfig) {
// App config environment variable overwrites only the config file
ConfigFile = AppConfig;
if (Global) {
ConfigFile = GetConfigDirectory(true) + "Config.json";
}
else {
ConfigFile = GetConfigDirectory(false) + "Config.json";
const char *AppConfig = getenv("FEX_APP_CONFIG");
if (AppConfig) {
// App config environment variable overwrites only the config file
ConfigFile = AppConfig;
}
else {
ConfigFile = GetConfigDirectory(false) + "Config.json";
}
}
return ConfigFile;
}
@@ -206,9 +211,12 @@ namespace JSON {
static std::map<FEXCore::Config::LayerType, std::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static FEXCore::Config::Layer *Meta{};
constexpr std::array<FEXCore::Config::LayerType, 6> LoadOrder = {
constexpr std::array<FEXCore::Config::LayerType, 9> LoadOrder = {
FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN,
FEXCore::Config::LayerType::LAYER_MAIN,
FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_GLOBAL_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_APP,
FEXCore::Config::LayerType::LAYER_ARGUMENTS,
FEXCore::Config::LayerType::LAYER_ENVIRONMENT,
@@ -613,7 +621,7 @@ namespace JSON {
// Application loaders
class MainLoader final : public FEXCore::Config::OptionMapper {
public:
explicit MainLoader();
explicit MainLoader(FEXCore::Config::LayerType Type);
explicit MainLoader(std::string ConfigFile);
void Load() override;
@@ -623,7 +631,7 @@ namespace JSON {
class AppLoader final : public FEXCore::Config::OptionMapper {
public:
explicit AppLoader(const std::string& Filename, bool Global);
explicit AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type);
void Load();
private:
@@ -659,9 +667,9 @@ namespace JSON {
}
}
MainLoader::MainLoader()
: FEXCore::Config::OptionMapper(FEXCore::Config::LayerType::LAYER_MAIN)
, Config{FEXCore::Config::GetConfigFileLocation()} {
MainLoader::MainLoader(FEXCore::Config::LayerType Type)
: FEXCore::Config::OptionMapper(Type)
, Config{FEXCore::Config::GetConfigFileLocation(Type == FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN)} {
}
MainLoader::MainLoader(std::string ConfigFile)
@@ -675,8 +683,10 @@ namespace JSON {
});
}
AppLoader::AppLoader(const std::string& Filename, bool Global)
: FEXCore::Config::OptionMapper(Global ? FEXCore::Config::LayerType::LAYER_GLOBAL_APP : FEXCore::Config::LayerType::LAYER_LOCAL_APP) {
AppLoader::AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type)
: FEXCore::Config::OptionMapper(Type) {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP ||
Type == FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP;
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
// Immediately load so we can reload the meta layer
@@ -735,17 +745,21 @@ namespace JSON {
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateGlobalMainLayer() {
return std::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN);
}
std::unique_ptr<FEXCore::Config::Layer> CreateMainLayer(std::string const *File) {
if (File) {
return std::make_unique<FEXCore::Config::MainLoader>(*File);
}
else {
return std::make_unique<FEXCore::Config::MainLoader>();
return std::make_unique<FEXCore::Config::MainLoader>(FEXCore::Config::LayerType::LAYER_MAIN);
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, bool Global) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Global);
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, FEXCore::Config::LayerType Type) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Type);
}
std::unique_ptr<FEXCore::Config::Layer> CreateEnvironmentLayer(char *const _envp[]) {
+45 -2
View File
@@ -16,10 +16,11 @@
},
"Multiblock": {
"Type": "bool",
"Default": "true",
"Default": "false",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation"
"Controls multiblock code compilation",
"Can cause long JIT compilation times and stutter"
]
},
"MaxInst": {
@@ -49,6 +50,13 @@
"Cache JIT object code to drive.",
"Allows JIT code to be shared between applications"
]
},
"EnableAVX": {
"Type": "bool",
"Default": "true",
"Desc": [
"Determines whether or not we use the expanded register file for AVX or not"
]
}
},
"Emulation": {
@@ -83,6 +91,13 @@
"Folder to find the guest-side thunking libraries."
]
},
"ThunkGuestLibs32": {
"Type": "str",
"Default": "@CMAKE_INSTALL_PREFIX@/share/fex-emu/GuestThunks_32/",
"Desc": [
"Folder to find the 32-bit guest-side thunking libraries."
]
},
"ThunkConfig": {
"Type": "str",
"Default": "",
@@ -309,6 +324,16 @@
"Forces a process to stall out on initialization",
"Useful for a process that keeps restarting and doesn't work"
]
},
"x86dec_SynchronizeRIPOnAllBlocks": {
"Type": "bool",
"Default": "false",
"Desc": [
"An application that uses try-catch or longjump extensively needs the ability to do context aware state flushing",
"In the case of FEX's block-linking, it won't always ensure that RIP is synchronized.",
"If an exception occurs and RIP isn't synchronized, then FEX's exception stack restore may not long jump as expected",
"Can be useful for Wine applications that rely on stack unwinding"
]
}
},
"Misc": {
@@ -334,6 +359,13 @@
"Desc": [
"Loads an AOT IR cache for the loaded executable."
]
},
"ServerSocketPath": {
"Type": "str",
"Default": "",
"Desc": [
"Override for a FEXServer socket path. Only useful for chroots."
]
}
}
},
@@ -351,6 +383,17 @@
"Type": "str",
"Default": ""
},
"APP_CONFIG_NAME": {
"Type": "str",
"Default": "",
"Desc": [
"This is the application config name that has been loaded.",
"This differs from APP_FILENAME in two ways",
"Where APP_FILENAME always points to the executable path that FEX-Emu is executing.",
"This matches what is used to load the AppLayer configuration name.",
"When running through a compatibility layer like wine, this will only be the exe name, instead of wine full path."
]
},
"IS64BIT_MODE": {
"Type": "bool",
"Default": "false"
+8
View File
@@ -110,6 +110,10 @@ namespace FEXCore::Context {
void RegisterExternalSyscallVisitor(FEXCore::Context::Context *CTX, [[maybe_unused]] uint64_t Syscall, [[maybe_unused]] FEXCore::HLE::SyscallVisitor *Visitor) {
}
HostFeatures GetHostFeatures(const FEXCore::Context::Context *CTX) {
return CTX->HostFeatures;
}
void HandleCallback(FEXCore::Context::Context *CTX, FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
CTX->HandleCallback(Thread, RIP);
}
@@ -198,6 +202,10 @@ namespace FEXCore::Context {
return CTX->AddCustomIREntrypoint(Entrypoint, Handler, Creator, Data);
}
void AppendThunkDefinitions(FEXCore::Context::Context *CTX, std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
CTX->AppendThunkDefinitions(Definitions);
}
namespace Debug {
void CompileRIP(FEXCore::Context::Context *CTX, uint64_t RIP) {
CTX->CompileRIP(CTX->ParentThread, RIP);
+29 -9
View File
@@ -1,8 +1,8 @@
#pragma once
#include "Common/JitSymbols.h"
#include "FEXHeaderUtils/ScopedSignalMask.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Interface/Core/X86HelperGen.h"
#include "Interface/Core/ObjectCache/ObjectCacheService.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -10,10 +10,12 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Event.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <stdint.h>
@@ -113,6 +115,8 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(x86dec_SynchronizeRIPOnAllBlocks, X86DEC_SYNCHRONIZERIPONALLBLOCKS);
FEX_CONFIG_OPT(EnableAVX, ENABLEAVX);
} Config;
FEXCore::HostFeatures HostFeatures;
@@ -121,6 +125,7 @@ namespace FEXCore::Context {
FEXCore::Core::InternalThreadState* ParentThread;
std::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
bool NeedToCheckXID{true};
std::mutex IdleWaitMutex;
std::condition_variable IdleWaitCV;
@@ -129,8 +134,8 @@ namespace FEXCore::Context {
Event PauseWait;
bool Running{};
std::shared_mutex CodeInvalidationMutex;
std::shared_mutex CodeInvalidationMutex;
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
@@ -170,13 +175,26 @@ namespace FEXCore::Context {
void RegisterHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
void RegisterFrontendHostSignalHandler(int Signal, HostSignalDelegatorFunction Func, bool Required);
// Must be called from owning thread
static void RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker);
// Wrapper which takes CpuStateFrame instead of InternalThreadState
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
FHU::ScopedSignalMaskWithSharedLock lk(Frame->Thread->CTX->CodeInvalidationMutex);
return Fn(Frame, record);
}
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
static void RemoveThreadCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
RemoveThreadCodeEntry(Frame->Thread, GuestRIP);
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}", Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
FHU::ScopedSignalMaskWithUniqueLock lk(Thread->CTX->CodeInvalidationMutex);
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
// returns false if a handler was already registered
@@ -305,6 +323,8 @@ namespace FEXCore::Context {
IRCaptureCache.SetAOTIRRenamer(CacheRenamer);
}
void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions);
FEXCore::Utils::PooledAllocatorMMap OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorMMap FrontendAllocator;
@@ -350,7 +370,7 @@ namespace FEXCore::Context {
bool StartPaused = false;
bool IsMemoryShared = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::unordered_map<uint64_t, std::tuple<std::function<void(uintptr_t Entrypoint, FEXCore::IR::IREmitter *)>, void *, void *>> CustomIRHandlers;
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
@@ -3,6 +3,7 @@
#include <aarch64/cpu-aarch64.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
@@ -1986,7 +1987,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}",
ToUnderlying(AtomicOp));
return false;
}
@@ -2039,7 +2041,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}",
ToUnderlying(AtomicOp));
return false;
}
@@ -2092,7 +2095,8 @@ uint64_t HandleAtomicLoadstoreExclusive(void *_ucontext, void *_info) {
DesiredFunction = NEGDesired;
break;
default:
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}", AtomicOp);
LogMan::Msg::EFmt("Unhandled JIT SIGBUS Atomic mem op 0x{:02x}",
ToUnderlying(AtomicOp));
return false;
}
@@ -17,23 +17,35 @@
#include <utility>
namespace FEXCore::CPU {
#define STATE x28
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
Arm64Emitter::Arm64Emitter(FEXCore::Context::Context *ctx, size_t size)
: vixl::aarch64::Assembler(size, vixl::aarch64::PositionDependentCode)
: vixl::aarch64::Assembler(size ? (byte*)FEXCore::Allocator::mmap(nullptr, size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0) : reinterpret_cast<byte*>(~0ULL),
size,
vixl::aarch64::PositionDependentCode)
, EmitterCTX {ctx} {
CPU.SetUp();
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
#else
auto Features = vixl::CPUFeatures::InferFromOS();
if (ctx->HostFeatures.SupportsAtomics) {
// Hypervisor can hide this on the c630?
Features.Combine(vixl::CPUFeatures::Feature::kLORegions);
}
#endif
SetCPUFeatures(Features);
}
Arm64Emitter::~Arm64Emitter() {
auto CodeBuffer = GetBuffer();
if (CodeBuffer->GetCapacity()) {
FEXCore::Allocator::munmap(CodeBuffer->GetStartAddress<void*>(), CodeBuffer->GetCapacity());
}
}
void Arm64Emitter::LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad) {
bool Is64Bit = Reg.IsX();
int Segments = Is64Bit ? 4 : 2;
@@ -209,19 +221,30 @@ void Arm64Emitter::SpillStaticRegs(bool FPRs, uint32_t GPRSpillMask, uint32_t FP
}
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg1.GetCode()) & FPRSpillMask) &&
((1U << Reg2.GetCode()) & FPRSpillMask)) {
stp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
if (((1U << Reg.GetCode()) & FPRSpillMask) != 0) {
mov(TMP4, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
st1b(Reg.Z().VnB(), PRED_TMP_32B, SVEMemOperand(STATE, TMP4));
}
}
else if (((1U << Reg1.GetCode()) & FPRSpillMask)) {
str(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRSpillMask)) {
str(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
} else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
if (((1U << Reg1.GetCode()) & FPRSpillMask) &&
((1U << Reg2.GetCode()) & FPRSpillMask)) {
stp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRSpillMask)) {
str(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRSpillMask)) {
str(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0])));
}
}
}
}
@@ -246,19 +269,37 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
if (FPRs) {
for (size_t i = 0; i < SRAFPR.size(); i+=2) {
auto Reg1 = SRAFPR[i];
auto Reg2 = SRAFPR[i+1];
if (EmitterCTX->HostFeatures.SupportsAVX) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
ptrue(PRED_TMP_16B.VnB(), SVE_VL16);
ptrue(PRED_TMP_32B.VnB(), SVE_VL32);
if (((1U << Reg1.GetCode()) & FPRFillMask) &&
((1U << Reg2.GetCode()) & FPRFillMask)) {
ldp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
for (size_t i = 0; i < SRAFPR.size(); i++) {
const auto Reg = SRAFPR[i];
if (((1U << Reg.GetCode()) & FPRFillMask) != 0) {
mov(TMP4, offsetof(Core::CpuStateFrame, State.xmm.avx.data[i][0]));
ld1b(Reg.Z().VnB(), PRED_TMP_32B.Zeroing(), SVEMemOperand(STATE, TMP4));
}
}
else if (((1U << Reg1.GetCode()) & FPRFillMask)) {
ldr(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRFillMask)) {
ldr(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm[i+1][0])));
} else {
for (size_t i = 0; i < SRAFPR.size(); i += 2) {
const auto Reg1 = SRAFPR[i];
const auto Reg2 = SRAFPR[i + 1];
if (((1U << Reg1.GetCode()) & FPRFillMask) &&
((1U << Reg2.GetCode()) & FPRFillMask)) {
ldp(Reg1.Q(), Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
}
else if (((1U << Reg1.GetCode()) & FPRFillMask)) {
ldr(Reg1.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i][0])));
}
else if (((1U << Reg2.GetCode()) & FPRFillMask)) {
ldr(Reg2.Q(), MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[i+1][0])));
}
}
}
}
@@ -266,20 +307,31 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
void Arm64Emitter::PushDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (RA64.size() + 1) * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRSize = RAFPR.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
sub(sp, sp, SPOffset);
int i = 0;
for (auto RA : RAFPR)
{
str(RA.Q(), MemOperand(sp, i * 8));
i+=2;
if (CanUseSVE) {
for (const auto& RA : RAFPR) {
mov(TMP4, i * 8);
st1b(RA.Z().VnB(), PRED_TMP_32B, SVEMemOperand(sp, TMP4));
i += 4;
}
} else {
for (const auto& RA : RAFPR) {
str(RA.Q(), MemOperand(sp, i * 8));
i += 2;
}
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
for (const auto& RA : RA64) {
str(RA, MemOperand(sp, i * 8));
i++;
}
@@ -289,18 +341,29 @@ void Arm64Emitter::PushDynamicRegsAndLR() {
}
void Arm64Emitter::PopDynamicRegsAndLR() {
uint64_t SPOffset = AlignUp((RA64.size() + 1) * 8 + RAFPR.size() * 16, 16);
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto GPRSize = (RA64.size() + 1) * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRSize = RAFPR.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
int i = 0;
for (auto RA : RAFPR)
{
ldr(RA.Q(), MemOperand(sp, i * 8));
i+=2;
if (CanUseSVE) {
for (const auto& RA : RAFPR) {
mov(TMP4, i * 8);
ld1b(RA.Z().VnB(), PRED_TMP_32B.Zeroing(), SVEMemOperand(sp, TMP4));
i += 4;
}
} else {
for (const auto& RA : RAFPR) {
ldr(RA.Q(), MemOperand(sp, i * 8));
i += 2;
}
}
#if 0 // All GPRs should be caller saved
for (auto RA : RA64)
{
for (const auto& RA : RA64) {
ldr(RA, MemOperand(sp, i * 8));
i++;
}
@@ -8,6 +8,10 @@
#include <aarch64/cpu-aarch64.h>
#include <aarch64/operands-aarch64.h>
#include <platform-vixl.h>
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#include <aarch64/simulator-constants-aarch64.h>
#endif
#include <FEXCore/Config/Config.h>
@@ -58,15 +62,41 @@ const std::array<aarch64::VRegister, 12> RAFPR = {
v8, v9, v10, v11, v12, v13, v14, v15
};
// Contains the address to the currently available CPU state
#define STATE x28
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
#define TMP1 x0
#define TMP2 x1
#define TMP3 x2
#define TMP4 x3
// Vector temporaries
#define VTMP1 v1
#define VTMP2 v2
#define VTMP3 v3
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
#define PRED_TMP_16B p6
#define PRED_TMP_32B p7
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public vixl::aarch64::Assembler {
protected:
Arm64Emitter(FEXCore::Context::Context *ctx, size_t size);
~Arm64Emitter();
FEXCore::Context::Context *EmitterCTX;
vixl::aarch64::CPU CPU;
void LoadConstant(vixl::aarch64::Register Reg, uint64_t Constant, bool NOPPad = false);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
@@ -83,6 +113,71 @@ protected:
void PopCalleeSavedRegisters();
void Align16B();
#ifdef VIXL_SIMULATOR
// Generates a vixl simulator runtime call.
//
// This matches behaviour of vixl's macro assembler, but we need to reimplement it since we aren't using the macro assembler.
// This isn't too complex with how vixl emits this.
//
// Emit:
// 1) hlt(kRuntimeCallOpcode)
// 2) Simulator wrapper handler
// 3) Function to call
// 4) Style of the function call (Call versus tail-call)
template<typename R, typename... P>
void GenerateRuntimeCall(R (*Function)(P...)) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
uintptr_t FunctionAddress = reinterpret_cast<uintptr_t>(Function);
hlt(kRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
// Runtime function address to call
dc(FunctionAddress);
// Call type
dc32(kCallRuntime);
}
template<typename R, typename... P>
void GenerateIndirectRuntimeCall(vixl::aarch64::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
hlt(kIndirectRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
// Register that contains the function to call
dc(Reg.GetCode());
// Call type
dc32(kCallRuntime);
}
template<>
void GenerateIndirectRuntimeCall<float, __uint128_t>(vixl::aarch64::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(Simulator::RuntimeCallStructHelper<float, __uint128_t>::Wrapper));
hlt(kIndirectRuntimeCallOpcode);
// Simulator wrapper address pointer.
dc(SimulatorWrapperAddress);
// Register that contains the function to call
dc(Reg.GetCode());
// Call type
dc32(kCallRuntime);
}
#endif
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
@@ -124,7 +124,7 @@ static inline void SetArmReg(void* ucontext, uint32_t id, uint64_t val) {
static inline __uint128_t GetArmFPR(void* ucontext, uint32_t id) {
auto MContext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&MContext->__reserved[0]);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
LOGMAN_THROW_AA_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
return HostState->FPRs[id];
}
@@ -143,7 +143,7 @@ static inline void BackupContext(void* ucontext, T *Backup) {
// Host FPR state starts at _mcontext->reserved[0];
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
LOGMAN_THROW_AA_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
Backup->FPSR = HostState->FPSR;
Backup->FPCR = HostState->FPCR;
memcpy(&Backup->FPRs[0], &HostState->FPRs[0], 32 * sizeof(__uint128_t));
@@ -163,7 +163,7 @@ static inline void RestoreContext(void* ucontext, T *Backup) {
auto _mcontext = GetMContext(ucontext);
HostFPRState *HostState = reinterpret_cast<HostFPRState*>(&_mcontext->__reserved[0]);
LOGMAN_THROW_A_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
LOGMAN_THROW_AA_FMT(HostState->Head.Magic == FPR_MAGIC, "Wrong FPR Magic: 0x{:08x}", HostState->Head.Magic);
memcpy(&HostState->FPRs[0], &Backup->FPRs[0], 32 * sizeof(__uint128_t));
HostState->FPCR = Backup->FPCR;
HostState->FPSR = Backup->FPSR;
+3 -3
View File
@@ -58,8 +58,8 @@ auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::mmap(nullptr, Buffer.Size, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (ThreadState->CTX->Config.GlobalJITNaming()) {
ThreadState->CTX->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
@@ -84,4 +84,4 @@ bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
}
}
}
}
+8 -3
View File
@@ -8,11 +8,11 @@ $end_info$
#include "Common/StringConv.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include "Utils/FileLoading.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXHeaderUtils/Syscalls.h>
#include "git_version.h"
@@ -39,6 +39,7 @@ namespace ProductNames {
static const char ARM_A78C[] = "Cortex-A78C";
static const char ARM_A710[] = "Cortex-A710";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X1C[] = "Cortex-X1C";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
@@ -83,7 +84,10 @@ static uint32_t CalculateNumberOfCPUs() {
return CPUs;
}
// TODO: Replace usages with CTX->HostFeatures.EnableAVX
// when AVX implementations are further along.
constexpr uint32_t SUPPORTS_AVX = 0;
// #define CPUID_AMD
#ifdef CPUID_AMD
constexpr uint32_t FAMILY_IDENTIFIER =
@@ -151,7 +155,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 35> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 36> CPUMIDRs = {{
// Typically big CPU cores
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
@@ -161,6 +165,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
{0x41, 0xd49, 1, ProductNames::ARM_N2}, // N2
{0x41, 0xd48, 1, ProductNames::ARM_X2}, // X2
{0x41, 0xd47, 1, ProductNames::ARM_A710}, // A710
{0x41, 0xd4C, 1, ProductNames::ARM_X1C}, // X1C
{0x41, 0xd44, 1, ProductNames::ARM_X1}, // X1
{0x41, 0xd42, 1, ProductNames::ARM_A78AE}, // A78AE
{0x41, 0xd41, 1, ProductNames::ARM_A78}, // A78
@@ -416,7 +421,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
Res.ecx =
(1 << 0) | // SSE3
(1 << 1) | // PCLMULQDQ
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
(1 << 2) | // DS area supports 64bit layout
(1 << 3) | // MWait
(0 << 4) | // DS-CPL
+78 -25
View File
@@ -44,6 +44,7 @@ $end_info$
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Threads.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXHeaderUtils/Syscalls.h>
#include <FEXHeaderUtils/TodoDefines.h>
@@ -154,6 +155,16 @@ namespace FEXCore::Context {
if (Config.CacheObjectCodeCompilation() != FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
CodeObjectCacheService = std::make_unique<FEXCore::CodeSerialize::CodeObjectSerializeService>(this);
}
if (!Config.EnableAVX) {
HostFeatures.SupportsAVX = false;
}
if (Config.BlockJITNaming() ||
Config.GlobalJITNaming() ||
Config.LibraryJITNaming()) {
// Only initialize symbols file if enabled. Ensures we don't pollute /tmp with empty files.
Symbols.InitFile();
}
}
Context::~Context() {
@@ -184,9 +195,11 @@ namespace FEXCore::Context {
greg = 0;
}
for (auto& xmm : NewThreadState.xmm) {
for (auto& xmm : NewThreadState.xmm.avx.data) {
xmm[0] = 0xDEADBEEFULL;
xmm[1] = 0xBAD0DAD1ULL;
xmm[2] = 0xDEADCAFEULL;
xmm[3] = 0xBAD2CAD3ULL;
}
memset(NewThreadState.flags, 0, Core::CPUState::NUM_EFLAG_BITS);
NewThreadState.flags[1] = 1;
@@ -209,7 +222,7 @@ namespace FEXCore::Context {
#if (_M_X86_64 && JIT_X86_64)
FEXCore::CPU::InitializeX86JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetX86JITBackendFeatures();
#elif (_M_ARM_64 && JIT_ARM64)
#elif (_M_ARM_64 && JIT_ARM64) || defined(VIXL_SIMULATOR)
FEXCore::CPU::InitializeArm64JITSignalHandlers(this);
BackendFeatures = FEXCore::CPU::GetArm64JITBackendFeatures();
#else
@@ -226,16 +239,16 @@ namespace FEXCore::Context {
DispatcherConfig.StaticRegisterAllocation = Config.StaticRegisterAllocation && BackendFeatures.SupportsStaticRegisterAllocation;
#if (_M_X86_64)
Dispatcher = FEXCore::CPU::Dispatcher::CreateX86(this, DispatcherConfig);
#elif (_M_ARM_64)
#if JIT_ARM64
Dispatcher = FEXCore::CPU::Dispatcher::CreateArm64(this, DispatcherConfig);
#elif JIT_X86_64
Dispatcher = FEXCore::CPU::Dispatcher::CreateX86(this, DispatcherConfig);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled with an unknown target");
#endif
// Initialize common signal handlers
auto PauseHandler = [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
return Thread->CTX->Dispatcher->HandleSignalPause(Thread, Signal, info, ucontext);
};
@@ -495,10 +508,21 @@ namespace FEXCore::Context {
ExecutionThreadHandler *Arg = reinterpret_cast<ExecutionThreadHandler*>(FEXCore::Allocator::malloc(sizeof(ExecutionThreadHandler)));
Arg->This = this;
Arg->Thread = Thread;
Thread->StartPaused = NeedToCheckXID;
Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
// Wait for the thread to have started
Thread->ThreadWaiting.Wait();
if (NeedToCheckXID) {
// The first time an application creates a thread, GLIBC installs their SETXID signal handler.
// FEX needs to capture all signals and defer them to the guest.
// Once FEX creates its first guest thread, overwrite the GLIBC SETXID handler *again* to ensure
// FEX maintains control of the signal handler on this signal.
NeedToCheckXID = false;
SignalDelegation->CheckXIDHandler();
Thread->StartRunning.NotifyAll();
}
}
void Context::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
@@ -546,11 +570,11 @@ namespace FEXCore::Context {
break;
#endif
case FEXCore::Config::CONFIG_IRJIT:
Thread->PassManager->InsertRegisterAllocationPass(DoSRA);
Thread->PassManager->InsertRegisterAllocationPass(DoSRA, HostFeatures.SupportsAVX);
#if (_M_X86_64 && JIT_X86_64)
Thread->CPUBackend = FEXCore::CPU::CreateX86JITCore(this, Thread);
#elif (_M_ARM_64 && JIT_ARM64)
#elif (_M_ARM_64 && JIT_ARM64) || defined(VIXL_SIMULATOR)
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
#else
ERROR_AND_DIE_FMT("FEXCore has been compiled without a viable JIT core");
@@ -651,6 +675,8 @@ namespace FEXCore::Context {
}
void Context::ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) {
FEXCORE_PROFILE_INSTANT("ClearCodeCache");
{
// Ensure the Code Object Serialization service has fully serialized this thread's data before clearing the cache
// Use the thread's object cache ref counter for this
@@ -717,7 +743,9 @@ namespace FEXCore::Context {
}
}
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
Context::GenerateIRResult Context::GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo) {
FEXCORE_PROFILE_SCOPED("GenerateIR");
Thread->OpDispatcher->ReownOrClaimBuffer();
Thread->OpDispatcher->ResetWorkingList();
@@ -726,7 +754,7 @@ namespace FEXCore::Context {
std::shared_lock lk(CustomIRMutex);
auto Handler = CustomIRHandlers.find(GuestRIP);
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
@@ -762,6 +790,13 @@ namespace FEXCore::Context {
// Reset any block-specific state
Thread->OpDispatcher->StartNewBlock();
if (Config.x86dec_SynchronizeRIPOnAllBlocks) {
// Ensure the RIP is synchronized to the context on block entry.
// In the case of block linking, the RIP may not have synchronized.
auto NewRIP = Thread->OpDispatcher->_EntrypointOffset(Block.Entry - GuestRIP, GPRSize);
Thread->OpDispatcher->_StoreContext(GPRSize, IR::GPRClass, NewRIP, offsetof(FEXCore::Core::CPUState, rip));
}
uint64_t InstsInBlock = Block.NumInstructions;
for (size_t i = 0; i < InstsInBlock; ++i) {
@@ -775,7 +810,7 @@ namespace FEXCore::Context {
if (ExtendedDebugInfo) {
Thread->OpDispatcher->_GuestOpcode(Block.Entry + BlockInstructionsLength - GuestRIP);
}
if (Config.SMCChecks == FEXCore::Config::CONFIG_SMC_FULL) {
auto ExistingCodePtr = reinterpret_cast<uint64_t*>(Block.Entry + BlockInstructionsLength);
@@ -788,7 +823,7 @@ namespace FEXCore::Context {
Thread->OpDispatcher->SetTrueJumpTarget(InvalidateCodeCond, CodeWasChangedBlock);
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_RemoveThreadCodeEntry();
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(Block.Entry + BlockInstructionsLength - GuestRIP, GPRSize));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -842,7 +877,7 @@ namespace FEXCore::Context {
}
}
}
Thread->OpDispatcher->Finalize();
Thread->FrontendDecoder->DelayedDisownBuffer();
@@ -981,6 +1016,7 @@ namespace FEXCore::Context {
}
uintptr_t Context::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
@@ -1090,7 +1126,7 @@ namespace FEXCore::Context {
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
if (Thread != Thread->CTX->ParentThread || StartPaused) {
if (Thread != Thread->CTX->ParentThread || StartPaused || Thread->StartPaused) {
// Parent thread doesn't need to wait to run
Thread->StartRunning.Wait();
}
@@ -1144,24 +1180,30 @@ namespace FEXCore::Context {
for (auto it = lower; it != upper; it++) {
for (auto Address: it->second) {
Context::RemoveThreadCodeEntry(Thread, Address);
Context::ThreadRemoveCodeEntry(Thread, Address);
}
it->second.clear();
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
static void InvalidateGuestCodeRangeInternal(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(CTX->ThreadCreationMutex);
for (auto &Thread : CTX->Threads) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
std::unique_lock CodeInvalidationLock(CTX->CodeInvalidationMutex);
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
InvalidateGuestCodeRange(CTX, Start, Length);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
}
void InvalidateGuestCodeRange(FEXCore::Context::Context *CTX, uint64_t Start, uint64_t Length, std::function<void(uint64_t start, uint64_t Length)> CallAfter) {
FHU::ScopedSignalMaskWithUniqueLock CodeInvalidationLock(CTX->CodeInvalidationMutex);
InvalidateGuestCodeRangeInternal(CTX, Start, Length);
CallAfter(Start, Length);
}
@@ -1170,8 +1212,6 @@ namespace FEXCore::Context {
IsMemoryShared = true;
if (Config.TSOAutoMigration) {
LogMan::Msg::IFmt("Migrating to shared memory mode");
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
LogMan::Throw::AFmt(Threads.size() == 1, "First MarkMemoryShared called must be before creating any threads");
@@ -1191,7 +1231,15 @@ namespace FEXCore::Context {
CTX->MarkMemoryShared();
}
void Context::RemoveThreadCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
void Context::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
std::shared_lock lk(Thread->CTX->CodeInvalidationMutex);
Thread->LookupCache->AddBlockLink(GuestDestination, HostLink, delinker);
}
void Context::ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
LogMan::Throw::AFmt(Thread->CTX->CodeInvalidationMutex.try_lock() == false, "CodeInvalidationMutex needs to be unique_locked here");
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->DebugStore.erase(GuestRIP);
@@ -1230,7 +1278,7 @@ namespace FEXCore::Context {
Thread->CurrentFrame->State.rip = RIP;
// Erase the RIP from all the storage backings if it exists
RemoveThreadCodeEntry(Thread, RIP);
ThreadRemoveCodeEntry(Thread, RIP);
// We don't care if compilation passes or not
CompileBlock(Thread->CurrentFrame, RIP);
@@ -1288,8 +1336,13 @@ namespace FEXCore::Context {
}
}
void Context::AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, std::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) {
Thread->FrontendDecoder->SetExternalBranches(ExternalBranches);
Thread->FrontendDecoder->SetSectionMaxAddress(SectionMaxAddress);
}
}
}
@@ -38,12 +38,20 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE x28
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 8192;
Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: FEXCore::CPU::Dispatcher(ctx), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
, config(config) {
: FEXCore::CPU::Dispatcher(ctx, config), Arm64Emitter(ctx, MAX_DISPATCHER_CODE_SIZE)
#ifdef VIXL_SIMULATOR
, Simulator {&Decoder}
#endif
{
#ifdef VIXL_SIMULATOR
// Hardcode a 256-bit vector width if we are running in the simulator.
Simulator.SetVectorLengthInBits(256);
#endif
SetAllowAssembler(true);
DispatchPtr = GetCursorAddress<AsmDispatch>();
@@ -179,7 +187,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
ret();
}
#ifdef VIXL_SIMULATOR
// VIXL simulator can't run syscalls.
constexpr bool SignalSafeCompile = false;
#else
constexpr bool SignalSafeCompile = true;
#endif
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
@@ -207,8 +220,12 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
mov(x0, STATE);
mov(x1, lr);
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
blr(x3);
ldr(x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uintptr_t, void *, void *>(x2);
#else
blr(x2);
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
@@ -267,8 +284,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
ldr(x3, &l_CompileBlock);
// X2 contains our guest RIP
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void *, uint64_t, void *>(x3);
#else
blr(x3); // { CTX, Frame, RIP}
#endif
if (SignalSafeCompile) {
// Now restore the signal mask
@@ -301,7 +321,7 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = GetCursorAddress<uint64_t>();
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
@@ -310,26 +330,29 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
}
{
// Guest Overflow handler
// Guest SIGTRAP handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = GetCursorAddress<uint64_t>();
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
LoadConstant(w1, 1);
strb(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException));
LoadConstant(w1, X86State::X86_TRAPNO_OF);
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo));
LoadConstant(w1, 0x80);
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code));
LoadConstant(x1, 0);
str(w1, STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code));
brk(0);
}
{
// Guest Overflow handler
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGSEGV = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs();
// hlt/udf = SIGILL
// brk = SIGTRAP
// ??? = SIGSEGV
// Force a SIGSEGV by loading zero
LoadConstant(x1, 0);
ldr(x1, MemOperand(x1));
}
@@ -347,7 +370,11 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
ldr(x0, &l_CTX);
mov(x1, STATE);
ldr(x2, &l_Sleep);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void *, void *>(x2);
#else
blr(x2);
#endif
PauseReturnInstruction = GetCursorAddress<uint64_t>();
// Fault to start running again
@@ -410,11 +437,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
#else
blr(x3);
#endif
FillStaticRegs();
// Result is now in x0
@@ -429,11 +459,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
#else
blr(x3);
#endif
FillStaticRegs();
// Result is now in x0
@@ -448,11 +481,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
#else
blr(x3);
#endif
FillStaticRegs();
// Result is now in x0
@@ -467,11 +503,14 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR();
SpillStaticRegs();
ldr(x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(x3);
#else
blr(x3);
#endif
FillStaticRegs();
// Result is now in x0
@@ -502,13 +541,27 @@ Arm64Dispatcher::Arm64Dispatcher(FEXCore::Context::Context *ctx, const Dispatche
}
}
#ifdef VIXL_SIMULATOR
void Arm64Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.RunFrom(reinterpret_cast<Instruction const*>(DispatchPtr));
}
void Arm64Dispatcher::ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.WriteXRegister(1, RIP);
Simulator.RunFrom(reinterpret_cast<Instruction const*>(CallbackPtr));
}
#endif
// Used by GenerateGDBPauseCheck, GenerateInterpreterTrampoline, destination buffer is set before use
static thread_local vixl::aarch64::Assembler emit((uint8_t*)&emit, 1);
size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) {
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxGDBPauseCheckSize);
vixl::CodeBufferCheckScope scope(&emit, MaxGDBPauseCheckSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
aarch64::Label RunBlock;
@@ -543,8 +596,8 @@ size_t Arm64Dispatcher::GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t Gues
}
size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "GenerateInterpreterTrampoline dispatcher does not support SRA");
LOGMAN_THROW_AA_FMT(!config.StaticRegisterAllocation, "GenerateInterpreterTrampoline dispatcher does not support SRA");
*emit.GetBuffer() = vixl::CodeBuffer(CodeBuffer, MaxInterpreterTrampolineSize);
vixl::CodeBufferCheckScope scope(&emit, MaxInterpreterTrampolineSize, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
@@ -570,7 +623,7 @@ size_t Arm64Dispatcher::GenerateInterpreterTrampoline(uint8_t *CodeBuffer) {
}
void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {
for(int i = 0; i < SRA64.size(); i++) {
for (size_t i = 0; i < SRA64.size(); i++) {
if (IgnoreMask & (1U << SRA64[i].GetCode())) {
// Skip this one, it's already spilled
continue;
@@ -578,14 +631,21 @@ void Arm64Dispatcher::SpillSRA(FEXCore::Core::InternalThreadState *Thread, void
Thread->CurrentFrame->State.gregs[i] = ArchHelpers::Context::GetArmReg(ucontext, SRA64[i].GetCode());
}
for(int i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&Thread->CurrentFrame->State.xmm[i][0], &FPR, sizeof(__uint128_t));
if (EmitterCTX->HostFeatures.SupportsAVX) {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][0], &FPR, sizeof(__uint128_t));
}
} else {
for (size_t i = 0; i < SRAFPR.size(); i++) {
auto FPR = ArchHelpers::Context::GetArmFPR(ucontext, SRAFPR[i].GetCode());
memcpy(&Thread->CurrentFrame->State.xmm.sse.data[i][0], &FPR, sizeof(__uint128_t));
}
}
}
void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
@@ -594,8 +654,9 @@ void Arm64Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thr
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.GuestSignal_SIGILL = GuestSignal_SIGILL;
Common.GuestSignal_SIGTRAP = GuestSignal_SIGTRAP;
Common.GuestSignal_SIGSEGV = GuestSignal_SIGSEGV;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
auto &AArch64 = Thread->CurrentFrame->Pointers.AArch64;
@@ -610,4 +671,4 @@ std::unique_ptr<Dispatcher> Dispatcher::CreateArm64(FEXCore::Context::Context *C
return std::make_unique<Arm64Dispatcher>(CTX, Config);
}
}
}
@@ -3,6 +3,10 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#ifdef VIXL_SIMULATOR
#include <aarch64/simulator-aarch64.h>
#endif
namespace FEXCore::Context {
struct Context;
}
@@ -20,6 +24,11 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
size_t GenerateGDBPauseCheck(uint8_t *CodeBuffer, uint64_t GuestRIP) override;
size_t GenerateInterpreterTrampoline(uint8_t *CodeBuffer) override;
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) override;
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) override;
#endif
protected:
void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) override;
@@ -29,7 +38,11 @@ class Arm64Dispatcher final : public Dispatcher, public Arm64Emitter {
uint64_t LDIVHandlerAddress{};
uint64_t LUREMHandlerAddress{};
uint64_t LREMHandlerAddress{};
DispatcherConfig config;
#ifdef VIXL_SIMULATOR
vixl::aarch64::Decoder Decoder;
vixl::aarch64::Simulator Simulator;
#endif
};
}
@@ -99,6 +99,7 @@ void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread,
SignalFrames.pop();
}
const bool IsAVXEnabled = CTX->Config.EnableAVX;
uintptr_t NewSP = OldSP;
auto Context = reinterpret_cast<ArchHelpers::Context::ContextBackup*>(NewSP);
@@ -159,10 +160,22 @@ void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread,
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86_64::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
auto *xstate = reinterpret_cast<x86_64::xstate*>(guest_uctx->uc_mcontext.fpregs);
auto *fpstate = &xstate->fpstate;
// Copy float registers
memcpy(Frame->State.mm, fpstate->_st, sizeof(Frame->State.mm));
memcpy(Frame->State.xmm, fpstate->_xmm, sizeof(Frame->State.xmm));
if (IsAVXEnabled) {
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&Frame->State.xmm.avx.data[i][0], &fpstate->_xmm[i], sizeof(__uint128_t));
}
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&Frame->State.xmm.avx.data[i][2], &xstate->ymmh.ymmh_space[i], sizeof(__uint128_t));
}
} else {
memcpy(Frame->State.xmm.sse.data, fpstate->_xmm, sizeof(Frame->State.xmm.sse.data));
}
// FCW store default
Frame->State.FCW = fpstate->fcw;
@@ -199,12 +212,20 @@ void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread,
Frame->State.flags[9] = 1;
Frame->State.rip = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP];
Frame->State.cs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS];
Frame->State.ds = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS];
Frame->State.es = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES];
Frame->State.fs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS];
Frame->State.gs = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS];
Frame->State.ss = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS];
Frame->State.cs_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS];
Frame->State.ds_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS];
Frame->State.es_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES];
Frame->State.fs_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS];
Frame->State.gs_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS];
Frame->State.ss_idx = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS];
Frame->State.cs_cached = Frame->State.gdt[Frame->State.cs_idx >> 3].base;
Frame->State.ds_cached = Frame->State.gdt[Frame->State.ds_idx >> 3].base;
Frame->State.es_cached = Frame->State.gdt[Frame->State.es_idx >> 3].base;
Frame->State.fs_cached = Frame->State.gdt[Frame->State.fs_idx >> 3].base;
Frame->State.gs_cached = Frame->State.gdt[Frame->State.gs_idx >> 3].base;
Frame->State.ss_cached = Frame->State.gdt[Frame->State.ss_idx >> 3].base;
#define COPY_REG(x) \
Frame->State.gregs[X86State::REG_##x] = guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_##x];
COPY_REG(RDI);
@@ -216,7 +237,8 @@ void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread,
COPY_REG(RCX);
COPY_REG(RSP);
#undef COPY_REG
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(guest_uctx->uc_mcontext.fpregs);
auto *xstate = reinterpret_cast<x86::xstate*>(guest_uctx->uc_mcontext.fpregs);
auto *fpstate = &xstate->fpstate;
// Copy float registers
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
@@ -225,7 +247,16 @@ void Dispatcher::RestoreThreadState(FEXCore::Core::InternalThreadState *Thread,
}
// Extended XMM state
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
if (IsAVXEnabled) {
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&fpstate->_xmm[i], &Frame->State.xmm.avx.data[i][0], sizeof(__uint128_t));
}
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&xstate->ymmh.ymmh_space[i], &Frame->State.xmm.avx.data[i][2], sizeof(__uint128_t));
}
} else {
memcpy(Frame->State.xmm.sse.data, fpstate->_xmm, sizeof(Frame->State.xmm.sse.data));
}
// FCW store default
Frame->State.FCW = fpstate->fcw;
@@ -274,6 +305,26 @@ static uint32_t ConvertSignalToError(int Signal, siginfo_t *HostSigInfo) {
return 0;
}
template <typename T>
static void SetXStateInfo(T* xstate, bool is_avx_enabled) {
auto* fpstate = &xstate->fpstate;
fpstate->sw_reserved.magic1 = x86_64::fpx_sw_bytes::FP_XSTATE_MAGIC;
fpstate->sw_reserved.extended_size = is_avx_enabled ? sizeof(T) : 0;
fpstate->sw_reserved.xfeatures |= x86_64::fpx_sw_bytes::FEATURE_FP |
x86_64::fpx_sw_bytes::FEATURE_SSE;
if (is_avx_enabled) {
fpstate->sw_reserved.xfeatures |= x86_64::fpx_sw_bytes::FEATURE_YMM;
}
fpstate->sw_reserved.xstate_size = fpstate->sw_reserved.extended_size;
if (is_avx_enabled) {
xstate->xstate_hdr.xfeatures = 0;
}
}
bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext, GuestSigAction *GuestAction, stack_t *GuestStack) {
auto ContextBackup = StoreThreadState(Thread, Signal, ucontext);
@@ -293,13 +344,14 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
uint64_t NewGuestSP = OldGuestSP;
// Pulling from context here
bool Is64BitMode = CTX->Config.Is64BitMode;
uint64_t SignalReturn = CTX->X86CodeGen.SignalReturn;
const bool Is64BitMode = CTX->Config.Is64BitMode;
const bool IsAVXEnabled = CTX->Config.EnableAVX;
const uint64_t SignalReturn = CTX->X86CodeGen.SignalReturn;
// Spill the SRA regardless of signal handler type
// We are going to be returning to the top of the dispatcher which will fill again
// Otherwise we might load garbage
if (SRAEnabled) {
if (config.StaticRegisterAllocation) {
if (Thread->CPUBackend->IsAddressInCodeBuffer(OldPC)) {
uint32_t IgnoreMask{};
#ifdef _M_ARM_64
@@ -374,8 +426,13 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
if (GuestAction->sa_flags & SA_SIGINFO) {
// Setup ucontext a bit
if (Is64BitMode) {
NewGuestSP -= sizeof(FEXCore::x86_64::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86_64::_libc_fpstate));
if (IsAVXEnabled) {
NewGuestSP -= sizeof(x86_64::xstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(x86_64::xstate));
} else {
NewGuestSP -= sizeof(x86_64::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(x86_64::_libc_fpstate));
}
uint64_t FPStateLocation = NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86_64::ucontext_t);
@@ -397,8 +454,9 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
guest_uctx->uc_flags = FEXCore::x86_64::UC_FP_XSTATE;
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(FPStateLocation);
FEXCore::x86_64::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86_64::_libc_fpstate*>(FPStateLocation);
guest_uctx->uc_mcontext.fpregs = reinterpret_cast<x86_64::_libc_fpstate*>(FPStateLocation);
auto *xstate = reinterpret_cast<x86_64::xstate*>(FPStateLocation);
SetXStateInfo(xstate, IsAVXEnabled);
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_RIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_EFL] = 0;
@@ -415,6 +473,7 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
// Overwrite si_code
guest_siginfo->si_code = Thread->CurrentFrame->SynchronousFaultData.si_code;
Signal = Frame->SynchronousFaultData.Signal;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86_64::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
@@ -443,9 +502,21 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
COPY_REG(RSP);
#undef COPY_REG
auto* fpstate = &xstate->fpstate;
// Copy float registers
memcpy(fpstate->_st, Frame->State.mm, sizeof(Frame->State.mm));
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
if (IsAVXEnabled) {
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&fpstate->_xmm[i], &Frame->State.xmm.avx.data[i][0], sizeof(__uint128_t));
}
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
memcpy(&xstate->ymmh.ymmh_space[i], &Frame->State.xmm.avx.data[i][2], sizeof(__uint128_t));
}
} else {
memcpy(fpstate->_xmm, Frame->State.xmm.sse.data, sizeof(Frame->State.xmm.sse.data));
}
// FCW store default
fpstate->fcw = Frame->State.FCW;
@@ -470,8 +541,13 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
else {
ContextBackup->Flags |= ArchHelpers::Context::ContextFlags::CONTEXT_FLAG_32BIT;
NewGuestSP -= sizeof(FEXCore::x86::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(FEXCore::x86::_libc_fpstate));
if (IsAVXEnabled) {
NewGuestSP -= sizeof(x86::xstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(x86::xstate));
} else {
NewGuestSP -= sizeof(x86::_libc_fpstate);
NewGuestSP = AlignDown(NewGuestSP, alignof(x86::_libc_fpstate));
}
uint64_t FPStateLocation = NewGuestSP;
NewGuestSP -= sizeof(FEXCore::x86::ucontext_t);
@@ -494,27 +570,30 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
// Pointer to where the fpreg memory is
guest_uctx->uc_mcontext.fpregs = static_cast<uint32_t>(FPStateLocation);
FEXCore::x86::_libc_fpstate *fpstate = reinterpret_cast<FEXCore::x86::_libc_fpstate*>(FPStateLocation);
auto *xstate = reinterpret_cast<x86::xstate*>(FPStateLocation);
SetXStateInfo(xstate, IsAVXEnabled);
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS] = Frame->State.cs_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS] = Frame->State.fs_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS] = Frame->State.gs_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS] = Frame->State.ss_idx;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_GS] = Frame->State.gs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_FS] = Frame->State.fs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ES] = Frame->State.es;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_DS] = Frame->State.ds;
if (ContextBackup->FaultToTopAndGeneratedException) {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = Frame->SynchronousFaultData.TrapNo;
guest_siginfo->si_code = Frame->SynchronousFaultData.si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = Frame->SynchronousFaultData.err_code;
Signal = Frame->SynchronousFaultData.Signal;
}
else {
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_TRAPNO] = ConvertSignalToTrapNo(Signal, HostSigInfo);
guest_siginfo->si_code = HostSigInfo->si_code;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_ERR] = ConvertSignalToError(Signal, HostSigInfo);
}
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EIP] = Frame->State.rip;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_CS] = Frame->State.cs;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_EFL] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_UESP] = 0;
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_SS] = Frame->State.ss;
#define COPY_REG(x) \
guest_uctx->uc_mcontext.gregs[FEXCore::x86::FEX_REG_##x] = Frame->State.gregs[X86State::REG_##x];
@@ -528,6 +607,8 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
COPY_REG(RSP);
#undef COPY_REG
auto *fpstate = &xstate->fpstate;
// Copy float registers
for (size_t i = 0; i < Core::CPUState::NUM_MMS; ++i) {
// 32-bit st register size is only 10 bytes. Not padded to 16byte like x86-64
@@ -536,7 +617,16 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
// Extended XMM state
fpstate->status = FEXCore::x86::fpstate_magic::MAGIC_XFPSTATE;
memcpy(fpstate->_xmm, Frame->State.xmm, sizeof(Frame->State.xmm));
if (IsAVXEnabled) {
for (size_t i = 0; i < std::size(Frame->State.xmm.avx.data); i++) {
memcpy(&fpstate->_xmm[i], &Frame->State.xmm.avx.data[i][0], sizeof(__uint128_t));
}
for (size_t i = 0; i < std::size(Frame->State.xmm.avx.data); i++) {
memcpy(&xstate->ymmh.ymmh_space[i], &Frame->State.xmm.avx.data[i][2], sizeof(__uint128_t));
}
} else {
memcpy(fpstate->_xmm, Frame->State.xmm.sse.data, sizeof(Frame->State.xmm.sse.data));
}
// FCW store default
fpstate->fcw = Frame->State.FCW;
@@ -619,14 +709,14 @@ bool Dispatcher::HandleGuestSignal(FEXCore::Core::InternalThreadState *Thread, i
else {
NewGuestSP -= 4;
*(uint32_t*)NewGuestSP = SignalReturn;
LOGMAN_THROW_A_FMT(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
LOGMAN_THROW_AA_FMT(SignalReturn < 0x1'0000'0000ULL, "This needs to be below 4GB");
Frame->State.gregs[FEXCore::X86State::REG_RSP] = NewGuestSP;
}
// The guest starts its signal frame with a zero initialized FPU
// Set that up now. Little bit costly but it's a requirement
// This state will be restored on rt_sigreturn
memset(Frame->State.xmm, 0, sizeof(Frame->State.xmm));
memset(Frame->State.xmm.avx.data, 0, sizeof(Frame->State.xmm));
memset(Frame->State.mm, 0, sizeof(Frame->State.mm));
Frame->State.FCW = 0x37F;
Frame->State.FTW = 0xFFFF;
@@ -665,11 +755,11 @@ bool Dispatcher::HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, i
// Store our thread state so we can come back to this
StoreThreadState(Thread, Signal, ucontext);
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
if (config.StaticRegisterAllocation && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadPauseHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
if (config.StaticRegisterAllocation) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
@@ -698,11 +788,11 @@ bool Dispatcher::HandleSignalPause(FEXCore::Core::InternalThreadState *Thread, i
Thread->CurrentFrame->SignalHandlerRefCounter = 0;
// Set the new PC
if (SRAEnabled && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
if (config.StaticRegisterAllocation && Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// We are in jit, SRA must be spilled
ArchHelpers::Context::SetPc(ucontext, ThreadStopHandlerAddressSpillSRA);
} else {
if (SRAEnabled) {
if (config.StaticRegisterAllocation) {
// We are in non-jit, SRA is already spilled
LOGMAN_THROW_A_FMT(!IsAddressInDispatcher(ArchHelpers::Context::GetPc(ucontext)),
"Signals in dispatcher have unsynchronized context");
@@ -32,7 +32,7 @@ struct DispatcherConfig {
class Dispatcher {
public:
virtual ~Dispatcher() = default;
/**
* @name Dispatch Helper functions
* @{ */
@@ -44,8 +44,9 @@ public:
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t UnimplementedInstructionAddress{};
uint64_t OverflowExceptionInstructionAddress{};
uint64_t GuestSignal_SIGILL{};
uint64_t GuestSignal_SIGTRAP{};
uint64_t GuestSignal_SIGSEGV{};
uint64_t IntCallbackReturnAddress{};
uint64_t PauseReturnInstruction{};
@@ -74,28 +75,29 @@ public:
static std::unique_ptr<Dispatcher> CreateX86(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
static std::unique_ptr<Dispatcher> CreateArm64(FEXCore::Context::Context *CTX, const DispatcherConfig &Config);
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
virtual void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
DispatchPtr(Frame);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
virtual void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
protected:
Dispatcher(FEXCore::Context::Context *ctx)
Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &Config)
: CTX {ctx}
, config {Config}
{}
ArchHelpers::Context::ContextBackup* StoreThreadState(FEXCore::Core::InternalThreadState *Thread, int Signal, void *ucontext);
void RestoreThreadState(FEXCore::Core::InternalThreadState *Thread, void *ucontext);
std::stack<uint64_t, std::vector<uint64_t>> SignalFrames;
bool SRAEnabled = false;
virtual void SpillSRA(FEXCore::Core::InternalThreadState *Thread, void *ucontext, uint32_t IgnoreMask) {}
FEXCore::Context::Context *CTX;
DispatcherConfig config;
static void SleepThread(FEXCore::Context::Context *ctx, FEXCore::Core::CpuStateFrame *Frame);
@@ -28,12 +28,12 @@ static constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096;
#define STATE r14
X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherConfig &config)
: Dispatcher(ctx)
: Dispatcher(ctx, config)
, Xbyak::CodeGenerator(MAX_DISPATCHER_CODE_SIZE,
FEXCore::Allocator::mmap(nullptr, MAX_DISPATCHER_CODE_SIZE, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0),
nullptr) {
LOGMAN_THROW_A_FMT(!config.StaticRegisterAllocation, "X86 dispatcher does not support SRA");
LOGMAN_THROW_AA_FMT(!config.StaticRegisterAllocation, "X86 dispatcher does not support SRA");
using namespace Xbyak;
using namespace Xbyak::util;
@@ -347,24 +347,29 @@ X86Dispatcher::X86Dispatcher(FEXCore::Context::Context *ctx, const DispatcherCon
{
// Guest SIGILL handler
// Needs to be distinct from the SignalHandlerReturnAddress
UnimplementedInstructionAddress = getCurr<uint64_t>();
GuestSignal_SIGILL = getCurr<uint64_t>();
ud2();
}
{
// Guest Overflow handler
// Guest SIGTRAP handler
// Needs to be distinct from the SignalHandlerReturnAddress
OverflowExceptionInstructionAddress = getCurr<uint64_t>();
GuestSignal_SIGTRAP = getCurr<uint64_t>();
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
add(byte STATE_PTR(CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException), 1);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.TrapNo), X86State::X86_TRAPNO_OF);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.err_code), 0);
mov(dword STATE_PTR(CpuStateFrame, SynchronousFaultData.si_code), 0x80);
int3();
}
{
// Guest SIGSEGV handler
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGSEGV = getCurr<uint64_t>();
// ud2 = SIGILL
// int3 = SIGTRAP
// hlt = SIGSEGV
hlt();
}
@@ -477,8 +482,9 @@ void X86Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Threa
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddress;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddress;
Common.UnimplementedInstructionHandler = UnimplementedInstructionAddress;
Common.OverflowExceptionHandler = OverflowExceptionInstructionAddress;
Common.GuestSignal_SIGILL = GuestSignal_SIGILL;
Common.GuestSignal_SIGTRAP = GuestSignal_SIGTRAP;
Common.GuestSignal_SIGSEGV = GuestSignal_SIGSEGV;
Common.SignalReturnHandler = SignalHandlerReturnAddress;
auto &Interpreter = Thread->CurrentFrame->Pointers.Interpreter;
+34 -24
View File
@@ -18,6 +18,7 @@ $end_info$
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/Utils/Telemetry.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <set>
@@ -188,7 +189,7 @@ Decoder::~Decoder() {
uint8_t Decoder::ReadByte() {
uint8_t Byte = InstStream[InstructionSize];
LOGMAN_THROW_A_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
LOGMAN_THROW_AA_FMT(InstructionSize < MAX_INST_SIZE, "Max instruction size exceeded!");
Instruction[InstructionSize] = Byte;
InstructionSize++;
return Byte;
@@ -200,14 +201,7 @@ uint8_t Decoder::PeekByte(uint8_t Offset) const {
}
uint64_t Decoder::ReadData(uint8_t Size) {
if (Size == 0) {
return 0;
}
if (Size > sizeof(uint64_t)) {
LOGMAN_MSG_A_FMT("Unknown data size to read");
return 0;
}
LOGMAN_THROW_AA_FMT(Size != 0 && Size <= sizeof(uint64_t), "Unknown data size to read");
uint64_t Res = 0;
std::memcpy(&Res, &InstStream[InstructionSize], Size);
@@ -347,13 +341,15 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
Operand->Data.SIB.Index = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X ? 1 : 0, SIB.index, false, false, false, false, 0b100);
Operand->Data.SIB.Base = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
LOGMAN_THROW_A_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
LOGMAN_THROW_AA_FMT(Displacement <= 4, "Number of bytes should be <= 4 for literal src");
uint64_t Literal = ReadData(Displacement);
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
if (Displacement) {
uint64_t Literal = ReadData(Displacement);
if (Displacement == 1) {
Literal = static_cast<int8_t>(Literal);
}
Operand->Data.SIB.Offset = Literal;
}
Operand->Data.SIB.Offset = Literal;
}
else if (ModRM.mod == 0) {
// Explained in Table 1-14. "Operand Addressing Using ModRM and SIB Bytes"
@@ -404,7 +400,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
return false;
}
LOGMAN_THROW_A_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
LOGMAN_THROW_AA_FMT(!(Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 && Info->Type <= FEXCore::X86Tables::TYPE_GROUP_P),
"Group Ops should have been decoded before this!");
uint8_t DestSize{};
@@ -455,8 +451,13 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DestSize = 2;
}
else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_128BIT);
DestSize = 16;
if (Options.L) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_256BIT);
DestSize = 32;
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_128BIT);
DestSize = 16;
}
}
else if (HasNarrowingDisplacement &&
(DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_DEF ||
@@ -487,7 +488,14 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
}
else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
if (Options.L) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
}
}
else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_256BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
}
else if (HasNarrowingDisplacement &&
(SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_DEF ||
@@ -528,7 +536,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (HAS_NON_XMM_SUBFLAG(Info->Flags, FEXCore::X86Tables::InstFlags::FLAGS_SF_REX_IN_BYTE)) {
LOGMAN_THROW_A_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
LOGMAN_THROW_AA_FMT(!HasMODRM, "This instruction shouldn't have ModRM!");
// If the REX is in the byte that means the lower nibble of the OP contains the destination GPR
// This also means that the destination is always a GPR on these ones
@@ -637,7 +645,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
}
if (Bytes != 0) {
LOGMAN_THROW_A_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
LOGMAN_THROW_AA_FMT(Bytes <= 8, "Number of bytes should be <= 8 for literal src");
DecodeInst->Src[CurrentSrc].Data.Literal.Size = Bytes;
@@ -662,7 +670,7 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->Src[CurrentSrc].Data.Literal.Value = Literal;
}
LOGMAN_THROW_A_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining",
LOGMAN_THROW_AA_FMT(Bytes == 0, "Inst at 0x{:x}: 0x{:04x} '{}' Had an instruction of size {} with {} remaining",
DecodeInst->PC, DecodeInst->OP, DecodeInst->TableInfo->Name ?: "UND", InstructionSize, Bytes);
DecodeInst->InstSize = InstructionSize;
return true;
@@ -688,7 +696,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
return false;
}
LOGMAN_THROW_A_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX,
LOGMAN_THROW_AA_FMT(Info->Type != FEXCore::X86Tables::TYPE_REX_PREFIX,
"REX PREFIX should have been decoded before this!");
if (Info->Type >= FEXCore::X86Tables::TYPE_GROUP_1 &&
@@ -745,7 +753,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
3,
};
uint8_t Field = RegToField[ModRM.reg];
LOGMAN_THROW_A_FMT(Field != 255, "Invalid field selected!");
LOGMAN_THROW_AA_FMT(Field != 255, "Invalid field selected!");
LocalOp = (Field << 3) | ModRM.rm;
return NormalOp(&SecondModRMTableOps[LocalOp], LocalOp);
@@ -781,6 +789,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
options.L = (Byte1 & 0b100) != 0;
}
else { // 0xC4 = Three byte VEX
const uint8_t Byte2 = ReadByte();
@@ -788,6 +797,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.X shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
@@ -1137,6 +1147,7 @@ const uint8_t *Decoder::AdjustAddrForSpecialRegion(uint8_t const* _InstStream, u
}
void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage) {
FEXCORE_PROFILE_SCOPED("DecodeInstructions");
Blocks.clear();
BlocksToDecode.clear();
HasBlocks.clear();
@@ -1171,7 +1182,6 @@ void Decoder::DecodeInstructionsAtEntry(uint8_t const* _InstStream, uint64_t PC,
std::set<uint64_t> CodePages = { CurrentCodePage };
AddContainedCodePage(PC, CurrentCodePage, FHU::FEX_PAGE_SIZE);
while (!BlocksToDecode.empty()) {
auto BlockDecodeIt = BlocksToDecode.begin();
+1
View File
@@ -49,6 +49,7 @@ private:
struct DecodedHeader {
uint8_t vvvv; // Encoded operand in a VEX prefix.
bool w; // VEX.W bit.
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Context::Context *CTX;
+29 -5
View File
@@ -259,7 +259,7 @@ struct FEX_PACKED GDBContextDefinition {
uint32_t fctrl;
uint32_t fstat;
uint32_t dummies[6];
uint64_t xmm[Core::CPUState::NUM_XMMS][2];
uint64_t xmm[Core::CPUState::NUM_XMMS][4];
uint32_t mxcsr;
};
@@ -306,7 +306,7 @@ std::string GdbServer::readRegs() {
GDB.fstat |= static_cast<uint32_t>(state.flags[FEXCore::X86State::X87FLAG_C2_LOC]) << 10;
GDB.fstat |= static_cast<uint32_t>(state.flags[FEXCore::X86State::X87FLAG_C3_LOC]) << 14;
memcpy(&GDB.xmm[0], &state.xmm[0], sizeof(GDB.xmm));
memcpy(&GDB.xmm[0], &state.xmm.avx.data[0], sizeof(GDB.xmm));
return encodeHex((unsigned char *)&GDB, sizeof(GDBContextDefinition));
}
@@ -382,9 +382,9 @@ GdbServer::HandledPacketType GdbServer::readReg(const std::string& packet) {
}
else if (addr >= offsetof(GDBContextDefinition, xmm[0][0]) &&
addr < offsetof(GDBContextDefinition, xmm[16][0])) {
const auto XmmIndex = (addr - offsetof(GDBContextDefinition, xmm[0][0])) / Core::CPUState::XMM_REG_SIZE;
const auto *Data = (unsigned char *)&state.xmm[XmmIndex][0];
return {encodeHex(Data, Core::CPUState::XMM_REG_SIZE), HandledPacketType::TYPE_ACK};
const auto XmmIndex = (addr - offsetof(GDBContextDefinition, xmm[0][0])) / Core::CPUState::XMM_AVX_REG_SIZE;
const auto *Data = (unsigned char *)&state.xmm.avx.data[XmmIndex][0];
return {encodeHex(Data, Core::CPUState::XMM_AVX_REG_SIZE), HandledPacketType::TYPE_ACK};
}
else if (addr == offsetof(GDBContextDefinition, mxcsr)) {
uint32_t Empty{};
@@ -490,6 +490,30 @@ std::string buildTargetXML() {
reg("mxcsr", "int", 32);
xml << "</feature>\n";
xml << "<feature name='org.gnu.gdb.i386.avx'>";
xml <<
R"(<vector id="v4f" type="ieee_single" count="4"/>
<vector id="v2d" type="ieee_double" count="2"/>
<vector id="v16i8" type="int8" count="16"/>
<vector id="v8i16" type="int16" count="8"/>
<vector id="v4i32" type="int32" count="4"/>
<vector id="v2i64" type="int64" count="2"/>
<union id="vec128">
<field name="v4_float" type="v4f"/>
<field name="v2_double" type="v2d"/>
<field name="v16_int8" type="v16i8"/>
<field name="v8_int16" type="v8i16"/>
<field name="v4_int32" type="v4i32"/>
<field name="v2_int64" type="v2i64"/>
<field name="uint128" type="uint128"/>
</union>
)";
for (size_t i = 0; i < Core::CPUState::NUM_XMMS; i++) {
reg(fmt::format("ymm{}h", i), "vec128", 128);
}
xml << "</feature>\n";
xml << "</target>";
xml << std::flush;
+61 -28
View File
@@ -1,7 +1,7 @@
#include "Interface/Core/CPUID.h"
#include "Interface/Core/HostFeatures.h"
#include <FEXCore/Core/HostFeatures.h>
#ifdef _M_ARM_64
#if defined(_M_ARM_64) || defined(VIXL_SIMULATOR)
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
@@ -50,8 +50,12 @@ static uint32_t GetDCZID() {
HostFeatures::HostFeatures() {
#ifdef _M_ARM_64
#if defined(_M_ARM_64) || defined(VIXL_SIMULATOR)
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
#else
auto Features = vixl::CPUFeatures::InferFromOS();
#endif
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
@@ -61,7 +65,26 @@ HostFeatures::HostFeatures() {
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
SupportsRCPC = Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsTSOImm9 = Features.Has(vixl::CPUFeatures::Feature::kRCpcImm);
SupportsPMULL_128Bit = Features.Has(vixl::CPUFeatures::Feature::kPmull1Q);
Supports3DNow = true;
SupportsSSE4A = true;
#ifdef VIXL_SIMULATOR
// Hardcode enable SVE with 256-bit wide registers.
SupportsAVX = true;
#else
SupportsAVX = Features.Has(vixl::CPUFeatures::Feature::kSVE2) &&
vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
#endif
SupportsSHA = true;
SupportsBMI1 = true;
SupportsBMI2 = true;
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
@@ -71,31 +94,6 @@ HostFeatures::HostFeatures() {
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#endif
#ifdef _M_X86_64
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
SupportsRCPC = true;
SupportsTSOImm9 = true;
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#else
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
constexpr uint32_t ExceptionEnableTraps =
@@ -116,6 +114,40 @@ HostFeatures::HostFeatures() {
SetFPCR(OriginalFPCR);
#endif
#endif
#if defined(_M_X86_64) && !defined(VIXL_SIMULATOR)
Xbyak::util::Cpu Features{};
SupportsAES = Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = Features.has(Xbyak::util::Cpu::tRDRAND) && Features.has(Xbyak::util::Cpu::tRDSEED);
SupportsRCPC = true;
SupportsTSOImm9 = true;
Supports3DNow = Features.has(Xbyak::util::Cpu::t3DN) && Features.has(Xbyak::util::Cpu::tE3DN);
SupportsSSE4A = Features.has(Xbyak::util::Cpu::tSSE4a);
SupportsAVX = true;
SupportsSHA = Features.has(Xbyak::util::Cpu::tSHA);
SupportsBMI1 = Features.has(Xbyak::util::Cpu::tBMI1);
SupportsBMI2 = Features.has(Xbyak::util::Cpu::tBMI2);
SupportsPMULL_128Bit = Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
// First ensure we support a new enough extended CPUID function range
__cpuid(0x8000'0000, eax, ebx, ecx, edx);
if (eax >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
__cpuid(0x8000'0008, eax, ebx, ecx, edx);
SupportsCLZERO = ebx & 1;
}
SupportsFlushInputsToZero = true;
SupportsFloatExceptions = true;
#endif
#ifdef VIXL_SIMULATOR
// simulator doesn't support dc(ZVA)
SupportsCLZERO = false;
#else
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
@@ -125,5 +157,6 @@ HostFeatures::HostFeatures() {
// This means we can use the instruction
SupportsCLZERO = DCZID_Bytes == CPUIDEmu::CACHELINE_SIZE;
}
#endif
}
}
+245 -225
View File
@@ -20,7 +20,7 @@ DEF_OP(TruncElementPair) {
switch (IROp->Size) {
case 4: {
uint64_t *Src = GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t *Src = GetSrc<uint64_t*>(Data->SSAData, Op->Pair);
uint64_t Result{};
Result = Src[0] & ~0U;
Result |= Src[1] << 32;
@@ -69,11 +69,11 @@ DEF_OP(CycleCounter) {
DEF_OP(Add) {
auto Op = IROp->C<IR::IROp_Add>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
auto Func = [](auto a, auto b) { return a + b; };
auto *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
auto *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
const auto Func = [](auto a, auto b) { return a + b; };
switch (OpSize) {
DO_OP(4, uint32_t, Func)
@@ -84,11 +84,11 @@ DEF_OP(Add) {
DEF_OP(Sub) {
auto Op = IROp->C<IR::IROp_Sub>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
auto Func = [](auto a, auto b) { return a - b; };
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
const auto Func = [](auto a, auto b) { return a - b; };
switch (OpSize) {
DO_OP(4, uint32_t, Func)
@@ -99,9 +99,9 @@ DEF_OP(Sub) {
DEF_OP(Neg) {
auto Op = IROp->C<IR::IROp_Neg>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
const uint64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Src);
switch (OpSize) {
case 4:
GD = -static_cast<int32_t>(Src);
@@ -115,10 +115,10 @@ DEF_OP(Neg) {
DEF_OP(Mul) {
auto Op = IROp->C<IR::IROp_Mul>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 4:
@@ -138,10 +138,10 @@ DEF_OP(Mul) {
DEF_OP(UMul) {
auto Op = IROp->C<IR::IROp_UMul>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 4:
@@ -161,9 +161,9 @@ DEF_OP(UMul) {
DEF_OP(Div) {
auto Op = IROp->C<IR::IROp_Div>();
uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint8_t OpSize = IROp->Size;
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 1:
@@ -179,7 +179,7 @@ DEF_OP(Div) {
GD = static_cast<int64_t>(Src1) / static_cast<int64_t>(Src2);
break;
case 16: {
__int128_t Tmp = *GetSrc<__int128_t*>(Data->SSAData, Op->Header.Args[0]) / *GetSrc<__int128_t*>(Data->SSAData, Op->Header.Args[1]);
__int128_t Tmp = *GetSrc<__int128_t*>(Data->SSAData, Op->Src1) / *GetSrc<__int128_t*>(Data->SSAData, Op->Src2);
memcpy(GDP, &Tmp, 16);
break;
}
@@ -189,10 +189,10 @@ DEF_OP(Div) {
DEF_OP(UDiv) {
auto Op = IROp->C<IR::IROp_UDiv>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 1:
@@ -208,7 +208,7 @@ DEF_OP(UDiv) {
GD = static_cast<uint64_t>(Src1) / static_cast<uint64_t>(Src2);
break;
case 16: {
__uint128_t Tmp = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]) / *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
__uint128_t Tmp = *GetSrc<__uint128_t*>(Data->SSAData, Op->Src1) / *GetSrc<__uint128_t*>(Data->SSAData, Op->Src2);
memcpy(GDP, &Tmp, 16);
break;
}
@@ -218,10 +218,10 @@ DEF_OP(UDiv) {
DEF_OP(Rem) {
auto Op = IROp->C<IR::IROp_Rem>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 1:
@@ -237,7 +237,7 @@ DEF_OP(Rem) {
GD = static_cast<int64_t>(Src1) % static_cast<int64_t>(Src2);
break;
case 16: {
__int128_t Tmp = *GetSrc<__int128_t*>(Data->SSAData, Op->Header.Args[0]) % *GetSrc<__int128_t*>(Data->SSAData, Op->Header.Args[1]);
__int128_t Tmp = *GetSrc<__int128_t*>(Data->SSAData, Op->Src1) % *GetSrc<__int128_t*>(Data->SSAData, Op->Src2);
memcpy(GDP, &Tmp, 16);
break;
}
@@ -247,10 +247,10 @@ DEF_OP(Rem) {
DEF_OP(URem) {
auto Op = IROp->C<IR::IROp_URem>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 1:
@@ -266,7 +266,7 @@ DEF_OP(URem) {
GD = static_cast<uint64_t>(Src1) % static_cast<uint64_t>(Src2);
break;
case 16: {
__uint128_t Tmp = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]) % *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
__uint128_t Tmp = *GetSrc<__uint128_t*>(Data->SSAData, Op->Src1) % *GetSrc<__uint128_t*>(Data->SSAData, Op->Src2);
memcpy(GDP, &Tmp, 16);
break;
}
@@ -276,10 +276,10 @@ DEF_OP(URem) {
DEF_OP(MulH) {
auto Op = IROp->C<IR::IROp_MulH>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 4: {
@@ -298,10 +298,10 @@ DEF_OP(MulH) {
DEF_OP(UMulH) {
auto Op = IROp->C<IR::IROp_UMulH>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
switch (OpSize) {
case 4:
GD = static_cast<uint64_t>(Src1) * static_cast<uint64_t>(Src2);
@@ -324,11 +324,11 @@ DEF_OP(UMulH) {
DEF_OP(Or) {
auto Op = IROp->C<IR::IROp_Or>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
auto Func = [](auto a, auto b) { return a | b; };
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
const auto Func = [](auto a, auto b) { return a | b; };
switch (OpSize) {
DO_OP(1, uint8_t, Func)
@@ -342,11 +342,11 @@ DEF_OP(Or) {
DEF_OP(And) {
auto Op = IROp->C<IR::IROp_And>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
auto Func = [](auto a, auto b) { return a & b; };
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
const auto Func = [](auto a, auto b) { return a & b; };
switch (OpSize) {
DO_OP(1, uint8_t, Func)
@@ -361,8 +361,8 @@ DEF_OP(Andn) {
auto Op = IROp->C<IR::IROp_Andn>();
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
constexpr auto Func = [](auto a, auto b) {
using Type = decltype(a);
return static_cast<Type>(a & static_cast<Type>(~b));
@@ -379,11 +379,11 @@ DEF_OP(Andn) {
DEF_OP(Xor) {
auto Op = IROp->C<IR::IROp_Xor>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
auto Func = [](auto a, auto b) { return a ^ b; };
void *Src1 = GetSrc<void*>(Data->SSAData, Op->Src1);
void *Src2 = GetSrc<void*>(Data->SSAData, Op->Src2);
const auto Func = [](auto a, auto b) { return a ^ b; };
switch (OpSize) {
DO_OP(1, uint8_t, Func)
@@ -396,11 +396,11 @@ DEF_OP(Xor) {
DEF_OP(Lshl) {
auto Op = IROp->C<IR::IROp_Lshl>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Mask = OpSize * 8 - 1;
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const uint8_t Mask = OpSize * 8 - 1;
switch (OpSize) {
case 4:
GD = static_cast<uint32_t>(Src1) << (Src2 & Mask);
@@ -414,11 +414,11 @@ DEF_OP(Lshl) {
DEF_OP(Lshr) {
auto Op = IROp->C<IR::IROp_Lshr>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Mask = OpSize * 8 - 1;
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const uint8_t Mask = OpSize * 8 - 1;
switch (OpSize) {
case 4:
GD = static_cast<uint32_t>(Src1) >> (Src2 & Mask);
@@ -432,11 +432,11 @@ DEF_OP(Lshr) {
DEF_OP(Ashr) {
auto Op = IROp->C<IR::IROp_Ashr>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint8_t Mask = OpSize * 8 - 1;
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const uint8_t Mask = OpSize * 8 - 1;
switch (OpSize) {
case 4:
GD = (uint32_t)(static_cast<int32_t>(Src1) >> (Src2 & Mask));
@@ -450,12 +450,12 @@ DEF_OP(Ashr) {
DEF_OP(Ror) {
auto Op = IROp->C<IR::IROp_Ror>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
auto Ror = [] (auto In, auto R) {
auto RotateMask = sizeof(In) * 8 - 1;
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src2);
const auto Ror = [] (auto In, auto R) {
const auto RotateMask = sizeof(In) * 8 - 1;
R &= RotateMask;
return (In >> R) | (In << (sizeof(In) * 8 - R));
};
@@ -474,11 +474,11 @@ DEF_OP(Ror) {
DEF_OP(Extr) {
auto Op = IROp->C<IR::IROp_Extr>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
auto Extr = [] (auto Src1, auto Src2, uint8_t lsb) -> decltype(Src1) {
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Upper);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Lower);
const auto Extr = [] (auto Src1, auto Src2, uint8_t lsb) -> decltype(Src1) {
__uint128_t Result{};
Result = Src1;
Result <<= sizeof(Src1) * 8;
@@ -500,7 +500,7 @@ DEF_OP(Extr) {
}
DEF_OP(PDep) {
const auto Op = IROp->C<IR::IROp_PExt>();
const auto Op = IROp->C<IR::IROp_PDep>();
const auto OpSize = IROp->Size;
if (OpSize != 4 && OpSize != 8) {
@@ -508,10 +508,10 @@ DEF_OP(PDep) {
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Input)
: *GetSrc<uint64_t*>(Data->SSAData, Op->Input);
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Mask)
: *GetSrc<uint64_t*>(Data->SSAData, Op->Mask);
uint64_t Result = 0;
for (uint64_t Index = 0; Mask > 0; Index++) {
@@ -532,10 +532,10 @@ DEF_OP(PExt) {
return;
}
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(0))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(0));
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Args(1))
: *GetSrc<uint64_t*>(Data->SSAData, Op->Args(1));
const uint64_t Input = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Input)
: *GetSrc<uint64_t*>(Data->SSAData, Op->Input);
uint64_t Mask = OpSize == 4 ? *GetSrc<uint32_t*>(Data->SSAData, Op->Mask)
: *GetSrc<uint64_t*>(Data->SSAData, Op->Mask);
uint64_t Result = 0;
for (uint64_t Offset = 0; Mask > 0; Offset++) {
@@ -549,39 +549,39 @@ DEF_OP(PExt) {
DEF_OP(LDiv) {
auto Op = IROp->C<IR::IROp_LDiv>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
switch (OpSize) {
case 2: {
uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
int16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[2]);
int32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
int32_t Res = Source / Divisor;
const uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Lower);
const uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Upper);
const int16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Divisor);
const int32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
const int32_t Res = Source / Divisor;
// We only store the lower bits of the result
GD = static_cast<int16_t>(Res);
break;
}
case 4: {
uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
int32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[2]);
int64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
int64_t Res = Source / Divisor;
const uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Lower);
const uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Upper);
const int32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Divisor);
const int64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
const int64_t Res = Source / Divisor;
// We only store the lower bits of the result
GD = static_cast<int32_t>(Res);
break;
}
case 8: {
uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
int64_t Divisor = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[2]);
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
const uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Lower);
const uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Upper);
const int64_t Divisor = *GetSrc<int64_t*>(Data->SSAData, Op->Divisor);
const __int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
const __int128_t Res = Source / Divisor;
// We only store the lower bits of the result
memcpy(GDP, &Res, OpSize);
@@ -593,39 +593,39 @@ DEF_OP(LDiv) {
DEF_OP(LUDiv) {
auto Op = IROp->C<IR::IROp_LUDiv>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
// Each source is OpSize in size
// So you can have up to a 128bit divide from x86-64
switch (OpSize) {
case 2: {
uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[2]);
uint32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
uint32_t Res = Source / Divisor;
const uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Lower);
const uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Upper);
const uint16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Divisor);
const uint32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
const uint32_t Res = Source / Divisor;
// We only store the lower bits of the result
GD = static_cast<uint16_t>(Res);
break;
}
case 4: {
uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[2]);
uint64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
uint64_t Res = Source / Divisor;
const uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Lower);
const uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Upper);
const uint32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Divisor);
const uint64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
const uint64_t Res = Source / Divisor;
// We only store the lower bits of the result
GD = static_cast<uint32_t>(Res);
break;
}
case 8: {
uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Divisor = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[2]);
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source / Divisor;
const uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Lower);
const uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Upper);
const uint64_t Divisor = *GetSrc<uint64_t*>(Data->SSAData, Op->Divisor);
const __uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
const __uint128_t Res = Source / Divisor;
// We only store the lower bits of the result
memcpy(GDP, &Res, OpSize);
@@ -637,39 +637,39 @@ DEF_OP(LUDiv) {
DEF_OP(LRem) {
auto Op = IROp->C<IR::IROp_LRem>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
// Each source is OpSize in size
// So you can have up to a 128bit Remainder from x86-64
switch (OpSize) {
case 2: {
uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
int16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[2]);
int32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
int32_t Res = Source % Divisor;
const uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Lower);
const uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Upper);
const int16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Divisor);
const int32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
const int32_t Res = Source % Divisor;
// We only store the lower bits of the result
GD = static_cast<int16_t>(Res);
break;
}
case 4: {
uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
int32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[2]);
int64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
int64_t Res = Source % Divisor;
const uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Lower);
const uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Upper);
const int32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Divisor);
const int64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
const int64_t Res = Source % Divisor;
// We only store the lower bits of the result
GD = static_cast<int32_t>(Res);
break;
}
case 8: {
uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
int64_t Divisor = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[2]);
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
const uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Lower);
const uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Upper);
const int64_t Divisor = *GetSrc<int64_t*>(Data->SSAData, Op->Divisor);
const __int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
const __int128_t Res = Source % Divisor;
// We only store the lower bits of the result
memcpy(GDP, &Res, OpSize);
break;
@@ -680,39 +680,39 @@ DEF_OP(LRem) {
DEF_OP(LURem) {
auto Op = IROp->C<IR::IROp_LURem>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
// Each source is OpSize in size
// So you can have up to a 128bit Remainder from x86-64
switch (OpSize) {
case 2: {
uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[1]);
uint16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[2]);
uint32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
uint32_t Res = Source % Divisor;
const uint16_t SrcLow = *GetSrc<uint16_t*>(Data->SSAData, Op->Lower);
const uint16_t SrcHigh = *GetSrc<uint16_t*>(Data->SSAData, Op->Upper);
const uint16_t Divisor = *GetSrc<uint16_t*>(Data->SSAData, Op->Divisor);
const uint32_t Source = (static_cast<uint32_t>(SrcHigh) << 16) | SrcLow;
const uint32_t Res = Source % Divisor;
// We only store the lower bits of the result
GD = static_cast<uint16_t>(Res);
break;
}
case 4: {
uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[1]);
uint32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[2]);
uint64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
uint64_t Res = Source % Divisor;
const uint32_t SrcLow = *GetSrc<uint32_t*>(Data->SSAData, Op->Lower);
const uint32_t SrcHigh = *GetSrc<uint32_t*>(Data->SSAData, Op->Upper);
const uint32_t Divisor = *GetSrc<uint32_t*>(Data->SSAData, Op->Divisor);
const uint64_t Source = (static_cast<uint64_t>(SrcHigh) << 32) | SrcLow;
const uint64_t Res = Source % Divisor;
// We only store the lower bits of the result
GD = static_cast<uint32_t>(Res);
break;
}
case 8: {
uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Divisor = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[2]);
__uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__uint128_t Res = Source % Divisor;
const uint64_t SrcLow = *GetSrc<uint64_t*>(Data->SSAData, Op->Lower);
const uint64_t SrcHigh = *GetSrc<uint64_t*>(Data->SSAData, Op->Upper);
const uint64_t Divisor = *GetSrc<uint64_t*>(Data->SSAData, Op->Divisor);
const __uint128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
const __uint128_t Res = Source % Divisor;
// We only store the lower bits of the result
memcpy(GDP, &Res, OpSize);
break;
@@ -723,62 +723,62 @@ DEF_OP(LURem) {
DEF_OP(Not) {
auto Op = IROp->C<IR::IROp_Not>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
const uint64_t mask[9]= { 0, 0xFF, 0xFFFF, 0, 0xFFFFFFFF, 0, 0, 0, 0xFFFFFFFFFFFFFFFFULL };
uint64_t Mask = mask[OpSize];
const uint64_t Mask = mask[OpSize];
GD = (~Src) & Mask;
}
DEF_OP(Popcount) {
auto Op = IROp->C<IR::IROp_Popcount>();
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
GD = std::popcount(Src);
}
DEF_OP(FindLSB) {
auto Op = IROp->C<IR::IROp_FindLSB>();
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Result = FindFirstSetBit(Src);
const uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
const uint64_t Result = FindFirstSetBit(Src);
GD = Result - 1;
}
DEF_OP(FindMSB) {
auto Op = IROp->C<IR::IROp_FindMSB>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]))) - 1; break;
case 2: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]))) - 1; break;
case 4: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]))) - 1; break;
case 8: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]))) - 1; break;
case 1: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint8_t*>(Data->SSAData, Op->Src))) - 1; break;
case 2: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint16_t*>(Data->SSAData, Op->Src))) - 1; break;
case 4: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint32_t*>(Data->SSAData, Op->Src))) - 1; break;
case 8: GD = (OpSize * 8 - std::countl_zero(*GetSrc<uint64_t*>(Data->SSAData, Op->Src))) - 1; break;
default: LOGMAN_MSG_A_FMT("Unknown FindMSB size: {}", OpSize); break;
}
}
DEF_OP(FindTrailingZeros) {
auto Op = IROp->C<IR::IROp_FindTrailingZeros>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1: {
auto Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Src);
GD = std::countr_zero(Src);
break;
}
case 2: {
auto Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Src);
GD = std::countr_zero(Src);
break;
}
case 4: {
auto Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Src);
GD = std::countr_zero(Src);
break;
}
case 8: {
auto Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
GD = std::countr_zero(Src);
break;
}
@@ -788,26 +788,26 @@ DEF_OP(FindTrailingZeros) {
DEF_OP(CountLeadingZeroes) {
auto Op = IROp->C<IR::IROp_CountLeadingZeroes>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 1: {
auto Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint8_t*>(Data->SSAData, Op->Src);
GD = std::countl_zero(Src);
break;
}
case 2: {
auto Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint16_t*>(Data->SSAData, Op->Src);
GD = std::countl_zero(Src);
break;
}
case 4: {
auto Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint32_t*>(Data->SSAData, Op->Src);
GD = std::countl_zero(Src);
break;
}
case 8: {
auto Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
GD = std::countl_zero(Src);
break;
}
@@ -817,12 +817,12 @@ DEF_OP(CountLeadingZeroes) {
DEF_OP(Rev) {
auto Op = IROp->C<IR::IROp_Rev>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
switch (OpSize) {
case 2: GD = BSwap16(*GetSrc<uint16_t*>(Data->SSAData, Op->Header.Args[0])); break;
case 4: GD = BSwap32(*GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[0])); break;
case 8: GD = BSwap64(*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0])); break;
case 2: GD = BSwap16(*GetSrc<uint16_t*>(Data->SSAData, Op->Src)); break;
case 4: GD = BSwap32(*GetSrc<uint32_t*>(Data->SSAData, Op->Src)); break;
case 8: GD = BSwap64(*GetSrc<uint64_t*>(Data->SSAData, Op->Src)); break;
default: LOGMAN_MSG_A_FMT("Unknown REV size: {}", OpSize); break;
}
}
@@ -830,34 +830,36 @@ DEF_OP(Rev) {
DEF_OP(Bfi) {
auto Op = IROp->C<IR::IROp_Bfi>();
uint64_t SourceMask = (1ULL << Op->Width) - 1;
if (Op->Width == 64)
if (Op->Width == 64) {
SourceMask = ~0ULL;
uint64_t DestMask = ~(SourceMask << Op->lsb);
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
uint64_t Res = (Src1 & DestMask) | ((Src2 & SourceMask) << Op->lsb);
}
const uint64_t DestMask = ~(SourceMask << Op->lsb);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Dest);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
const uint64_t Res = (Src1 & DestMask) | ((Src2 & SourceMask) << Op->lsb);
GD = Res;
}
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
LOGMAN_THROW_AA_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
uint64_t SourceMask = (1ULL << Op->Width) - 1;
if (Op->Width == 64)
if (Op->Width == 64) {
SourceMask = ~0ULL;
}
SourceMask <<= Op->lsb;
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Src);
GD = (Src & SourceMask) >> Op->lsb;
}
DEF_OP(Sbfe) {
auto Op = IROp->C<IR::IROp_Sbfe>();
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for SBFE: {}", IROp->Size);
int64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t ShiftLeftAmount = (64 - (Op->Width + Op->lsb));
uint64_t ShiftRightAmount = ShiftLeftAmount + Op->lsb;
LOGMAN_THROW_AA_FMT(IROp->Size <= 8, "OpSize is too large for SBFE: {}", IROp->Size);
int64_t Src = *GetSrc<int64_t*>(Data->SSAData, Op->Src);
const uint64_t ShiftLeftAmount = (64 - (Op->Width + Op->lsb));
const uint64_t ShiftRightAmount = ShiftLeftAmount + Op->lsb;
Src <<= ShiftLeftAmount;
Src >>= ShiftRightAmount;
GD = Src;
@@ -865,20 +867,20 @@ DEF_OP(Sbfe) {
DEF_OP(Select) {
auto Op = IROp->C<IR::IROp_Select>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp2);
uint64_t ArgTrue;
uint64_t ArgFalse;
if (OpSize == 4) {
ArgTrue = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[2]);
ArgFalse = *GetSrc<uint32_t*>(Data->SSAData, Op->Header.Args[3]);
ArgTrue = *GetSrc<uint32_t*>(Data->SSAData, Op->TrueVal);
ArgFalse = *GetSrc<uint32_t*>(Data->SSAData, Op->FalseVal);
} else {
ArgTrue = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[2]);
ArgFalse = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[3]);
ArgTrue = *GetSrc<uint64_t*>(Data->SSAData, Op->TrueVal);
ArgFalse = *GetSrc<uint64_t*>(Data->SSAData, Op->FalseVal);
}
bool CompResult;
@@ -892,57 +894,75 @@ DEF_OP(Select) {
}
DEF_OP(VExtractToGPR) {
auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto OpSize = IROp->Size;
uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Header.Args[0]);
constexpr auto AVXRegSize = Core::CPUState::XMM_AVX_REG_SIZE;
constexpr auto SSERegSize = Core::CPUState::XMM_SSE_REG_SIZE;
constexpr auto SSEBitSize = SSERegSize * 8;
LOGMAN_THROW_A_FMT(IROp->Size <= 16, "OpSize is too large for VExtractToGPR: {}", IROp->Size);
const auto ElementSize = Op->Header.ElementSize;
const auto ElementSizeBits = ElementSize * 8;
const auto Shift = ElementSizeBits * Op->Index;
if (SourceSize == 16) {
__uint128_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
const uint32_t SourceSize = GetOpSize(Data->CurrentIR, Op->Vector);
LOGMAN_THROW_AA_FMT(OpSize <= AVXRegSize,
"OpSize is too large for VExtractToGPR: {}", OpSize);
if (SourceSize >= SSERegSize) {
__uint128_t SourceMask = (1ULL << ElementSizeBits) - 1;
if (ElementSize == 8) {
SourceMask = ~0ULL;
}
__uint128_t Src = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
Src >>= Shift;
Src &= SourceMask;
memcpy(GDP, &Src, Op->Header.ElementSize);
const auto Src = *GetSrc<InterpVector256*>(Data->SSAData, Op->Vector);
const auto GetResult = [&] {
if (Shift >= SSEBitSize) {
const auto NormalizedShift = Shift - SSEBitSize;
return (Src.Upper >> NormalizedShift) & SourceMask;
} else {
return (Src.Lower >> Shift) & SourceMask;
}
};
const auto Result = GetResult();
memcpy(GDP, &Result, ElementSize);
}
else {
uint64_t SourceMask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
uint64_t Shift = Op->Header.ElementSize * Op->Index * 8;
if (Op->Header.ElementSize == 8)
uint64_t SourceMask = (1ULL << ElementSizeBits) - 1;
if (ElementSize == 8) {
SourceMask = ~0ULL;
}
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
Src >>= Shift;
Src &= SourceMask;
GD = Src;
const uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Vector);
const uint64_t Result = (Src >> Shift) & SourceMask;
GD = Result;
}
}
DEF_OP(Float_ToGPR_ZS) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_ZS>();
uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
const uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // int64_t <- float
int64_t Dst = (int64_t)std::trunc(*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]));
const int64_t Dst = (int64_t)std::trunc(*GetSrc<float*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0808: { // int64_t <- double
int64_t Dst = (int64_t)std::trunc(*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]));
const int64_t Dst = (int64_t)std::trunc(*GetSrc<double*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0404: { // int32_t <- float
int32_t Dst = (int32_t)std::trunc(*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]));
const int32_t Dst = (int32_t)std::trunc(*GetSrc<float*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0408: { // int32_t <- double
int32_t Dst = (int32_t)std::trunc(*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]));
const int32_t Dst = (int32_t)std::trunc(*GetSrc<double*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
@@ -951,25 +971,25 @@ DEF_OP(Float_ToGPR_ZS) {
DEF_OP(Float_ToGPR_S) {
auto Op = IROp->C<IR::IROp_Float_ToGPR_S>();
uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
const uint16_t Conv = (IROp->Size << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // int64_t <- float
int64_t Dst = (int64_t)std::nearbyint(*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]));
const int64_t Dst = (int64_t)std::nearbyint(*GetSrc<float*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0808: { // int64_t <- double
int64_t Dst = (int64_t)std::nearbyint(*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]));
const int64_t Dst = (int64_t)std::nearbyint(*GetSrc<double*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0404: { // int32_t <- float
int32_t Dst = (int32_t)std::nearbyint(*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]));
const int32_t Dst = (int32_t)std::nearbyint(*GetSrc<float*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
case 0x0408: { // int32_t <- double
int32_t Dst = (int32_t)std::nearbyint(*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]));
const int32_t Dst = (int32_t)std::nearbyint(*GetSrc<double*>(Data->SSAData, Op->Scalar));
memcpy(GDP, &Dst, IROp->Size);
break;
}
@@ -980,9 +1000,9 @@ DEF_OP(FCmp) {
auto Op = IROp->C<IR::IROp_FCmp>();
uint32_t ResultFlags{};
if (Op->ElementSize == 4) {
float Src1 = *GetSrc<float*>(Data->SSAData, Op->Header.Args[0]);
float Src2 = *GetSrc<float*>(Data->SSAData, Op->Header.Args[1]);
bool Unordered = std::isnan(Src1) || std::isnan(Src2);
const float Src1 = *GetSrc<float*>(Data->SSAData, Op->Scalar1);
const float Src2 = *GetSrc<float*>(Data->SSAData, Op->Scalar2);
const bool Unordered = std::isnan(Src1) || std::isnan(Src2);
if (Op->Flags & (1 << IR::FCMP_FLAG_LT)) {
if (Unordered || (Src1 < Src2)) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
@@ -1000,9 +1020,9 @@ DEF_OP(FCmp) {
}
}
else {
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
bool Unordered = std::isnan(Src1) || std::isnan(Src2);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Scalar1);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Scalar2);
const bool Unordered = std::isnan(Src1) || std::isnan(Src2);
if (Op->Flags & (1 << IR::FCMP_FLAG_LT)) {
if (Unordered || (Src1 < Src2)) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
@@ -42,7 +42,7 @@ DEF_OP(ExitFunction) {
uintptr_t* ContextPtr = reinterpret_cast<uintptr_t*>(Data->State->CurrentFrame);
void *ContextData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src = GetSrc<void*>(Data->SSAData, Op->NewRIP);
memcpy(ContextData, Src, OpSize);
@@ -51,22 +51,22 @@ DEF_OP(ExitFunction) {
DEF_OP(Jump) {
auto Op = IROp->C<IR::IROp_Jump>();
uintptr_t ListBegin = Data->CurrentIR->GetListData();
uintptr_t DataBegin = Data->CurrentIR->GetData();
const uintptr_t ListBegin = Data->CurrentIR->GetListData();
const uintptr_t DataBegin = Data->CurrentIR->GetData();
Data->BlockIterator = IR::NodeIterator(ListBegin, DataBegin, Op->Header.Args[0]);
Data->BlockIterator = IR::NodeIterator(ListBegin, DataBegin, Op->TargetBlock);
Data->BlockResults.Redo = true;
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
uintptr_t ListBegin = Data->CurrentIR->GetListData();
uintptr_t DataBegin = Data->CurrentIR->GetData();
const uintptr_t ListBegin = Data->CurrentIR->GetListData();
const uintptr_t DataBegin = Data->CurrentIR->GetData();
bool CompResult;
uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp1);
uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp2);
const uint64_t Src1 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp1);
const uint64_t Src2 = *GetSrc<uint64_t*>(Data->SSAData, Op->Cmp2);
if (Op->CompareSize == 4)
CompResult = IsConditionTrue<uint32_t, int32_t, float>(Op->Cond.Val, Src1, Src2);
@@ -127,7 +127,7 @@ DEF_OP(Thunk) {
auto Op = IROp->C<IR::IROp_Thunk>();
auto thunkFn = Data->State->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
thunkFn(*GetSrc<void**>(Data->SSAData, Op->Header.Args[0]));
thunkFn(*GetSrc<void**>(Data->SSAData, Op->ArgPtr));
}
DEF_OP(ValidateCode) {
@@ -141,15 +141,15 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveThreadCodeEntry) {
Data->State->CTX->RemoveThreadCodeEntry(Data->State, Data->CurrentEntry);
DEF_OP(ThreadRemoveCodeEntry) {
Data->State->CTX->ThreadRemoveCodeEntryFromJit(Data->State->CurrentFrame, Data->CurrentEntry);
}
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
uint64_t *DstPtr = GetDest<uint64_t*>(Data->SSAData, Node);
uint64_t Arg = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Leaf = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const uint64_t Arg = *GetSrc<uint64_t*>(Data->SSAData, Op->Function);
const uint64_t Leaf = *GetSrc<uint64_t*>(Data->SSAData, Op->Leaf);
auto Results = Data->State->CTX->CPUID.RunFunction(Arg, Leaf);
memcpy(DstPtr, &Results, sizeof(uint32_t) * 4);
@@ -13,53 +13,77 @@ $end_info$
namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
const auto ElementSize = Op->Header.ElementSize;
const auto ElementSizeBits = ElementSize * 8;
constexpr auto SSEBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
uint64_t Offset = Op->DestIdx * Op->Header.ElementSize * 8;
__uint128_t Mask = (1ULL << (Op->Header.ElementSize * 8)) - 1;
if (Op->Header.ElementSize == 8) {
const uint64_t Offset = Op->DestIdx * ElementSizeBits;
const auto InUpperLane = Offset >= SSEBitSize;
__uint128_t Mask = (1ULL << ElementSizeBits) - 1;
if (ElementSize == 8) {
Mask = ~0ULL;
}
Src2 = Src2 & Mask;
Mask <<= Offset;
const auto Src1 = *GetSrc<InterpVector256*>(Data->SSAData, Op->DestVector);
const auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Src);
const auto Scalar = Src2 & Mask;
const auto ScaledOffset = InUpperLane ? Offset - SSEBitSize
: Offset;
// Now shift into place and set all bits but
// the ones where we're going to insert our value.
Mask <<= ScaledOffset;
Mask = ~Mask;
__uint128_t Dst = Src1 & Mask;
Dst |= Src2 << Offset;
const auto Dst = [&] {
if (InUpperLane) {
return InterpVector256{
.Lower = Src1.Lower,
.Upper = (Src1.Upper & Mask) | (Scalar << ScaledOffset),
};
} else {
return InterpVector256{
.Lower = (Src1.Lower & Mask) | (Scalar << ScaledOffset),
.Upper = Src1.Upper,
};
}
}();
memcpy(GDP, &Dst, OpSize);
}
DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[0]), Op->Header.ElementSize);
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Src), Op->Header.ElementSize);
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
float Dst = (float)*GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
const float Dst = (float)*GetSrc<int32_t*>(Data->SSAData, Op->Src);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0408: { // Float <- int64_t
float Dst = (float)*GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
const float Dst = (float)*GetSrc<int64_t*>(Data->SSAData, Op->Src);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0804: { // Double <- int32_t
double Dst = (double)*GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
const double Dst = (double)*GetSrc<int32_t*>(Data->SSAData, Op->Src);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
case 0x0808: { // Double <- int64_t
double Dst = (double)*GetSrc<int64_t*>(Data->SSAData, Op->Header.Args[0]);
const double Dst = (double)*GetSrc<int64_t*>(Data->SSAData, Op->Src);
memcpy(GDP, &Dst, Op->Header.ElementSize);
break;
}
@@ -68,15 +92,15 @@ DEF_OP(Float_FromGPR_S) {
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // Double <- Float
double Dst = (double)*GetSrc<float*>(Data->SSAData, Op->Header.Args[0]);
const double Dst = (double)*GetSrc<float*>(Data->SSAData, Op->Scalar);
memcpy(GDP, &Dst, 8);
break;
}
case 0x0408: { // Float <- Double
float Dst = (float)*GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
const float Dst = (float)*GetSrc<double*>(Data->SSAData, Op->Scalar);
memcpy(GDP, &Dst, 4);
break;
}
@@ -86,68 +110,78 @@ DEF_OP(Float_FToF) {
DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
const uint8_t ElementSize = Op->Header.ElementSize;
const uint8_t Elements = OpSize / ElementSize;
auto Func = [](auto a, auto min, auto max) { return a; };
switch (Op->Header.ElementSize) {
const auto Func = [](auto a, auto min, auto max) { return a; };
switch (ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, float, int32_t, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, double, int64_t, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_Vector_FToZS>();
const uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
const uint8_t ElementSize = Op->Header.ElementSize;
const uint8_t Elements = OpSize / ElementSize;
auto Func = [](auto a, auto min, auto max) { return std::trunc(a); };
switch (Op->Header.ElementSize) {
const auto Func = [](auto a, auto min, auto max) { return std::trunc(a); };
switch (ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, int32_t, float, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, int64_t, double, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
const uint8_t ElementSize = Op->Header.ElementSize;
const uint8_t Elements = OpSize / ElementSize;
auto Func = [](auto a, auto min, auto max) { return std::nearbyint(a); };
switch (Op->Header.ElementSize) {
const auto Func = [](auto a, auto min, auto max) { return std::nearbyint(a); };
switch (ElementSize) {
DO_VECTOR_1SRC_2TYPE_OP(4, int32_t, float, Func, 0, 0)
DO_VECTOR_1SRC_2TYPE_OP(8, int64_t, double, Func, 0, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToF) {
auto Op = IROp->C<IR::IROp_Vector_FToF>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_Vector_FToF>();
const uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t ElementSize = Op->Header.ElementSize;
const uint16_t Conv = (ElementSize << 8) | Op->SrcElementSize;
auto Func = [](auto a, auto min, auto max) { return a; };
const auto Func = [](auto a, auto min, auto max) { return a; };
switch (Conv) {
case 0x0804: { // Double <- float
// Only the lower elements from the source
@@ -165,52 +199,55 @@ DEF_OP(Vector_FToF) {
DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(float, double, Func, 0, 0)
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Conversion Type : 0x{:04x}", Conv); break;
default:
LOGMAN_MSG_A_FMT("Unknown Conversion Type : 0x{:04x}", Conv);
break;
}
memcpy(GDP, Tmp, OpSize);
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_Vector_FToI>();
const uint8_t OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Tmp[16]{};
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE]{};
uint8_t Elements = OpSize / Op->Header.ElementSize;
auto Func_Nearest = [](auto a) { return std::rint(a); };
auto Func_Neg = [](auto a) { return std::floor(a); };
auto Func_Pos = [](auto a) { return std::ceil(a); };
auto Func_Trunc = [](auto a) { return std::trunc(a); };
auto Func_Host = [](auto a) { return std::rint(a); };
const uint8_t ElementSize = Op->Header.ElementSize;
const uint8_t Elements = OpSize / ElementSize;
const auto Func_Nearest = [](auto a) { return std::rint(a); };
const auto Func_Neg = [](auto a) { return std::floor(a); };
const auto Func_Pos = [](auto a) { return std::ceil(a); };
const auto Func_Trunc = [](auto a) { return std::trunc(a); };
const auto Func_Host = [](auto a) { return std::rint(a); };
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Nearest)
DO_VECTOR_1SRC_OP(8, double, Func_Nearest)
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Neg)
DO_VECTOR_1SRC_OP(8, double, Func_Neg)
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Pos)
DO_VECTOR_1SRC_OP(8, double, Func_Pos)
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Trunc)
DO_VECTOR_1SRC_OP(8, double, Func_Trunc)
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_1SRC_OP(4, float, Func_Host)
DO_VECTOR_1SRC_OP(8, double, Func_Host)
}
@@ -360,7 +360,7 @@ namespace FEXCore::CPU {
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Vector);
// Pseudo-code
// Dst = InvMixColumns(STATE)
@@ -371,8 +371,8 @@ DEF_OP(AESImc) {
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->State);
auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Key);
// Pseudo-code
// STATE = Src1
@@ -391,8 +391,8 @@ DEF_OP(AESEnc) {
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->State);
auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Key);
// Pseudo-code
// STATE = Src1
@@ -409,8 +409,8 @@ DEF_OP(AESEncLast) {
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->State);
auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Key);
// Pseudo-code
// STATE = Src1
@@ -429,8 +429,8 @@ DEF_OP(AESDec) {
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
__uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
__uint128_t Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[1]);
auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->State);
auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Key);
// Pseudo-code
// STATE = Src1
@@ -447,7 +447,7 @@ DEF_OP(AESDecLast) {
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
const uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->Src);
// Pseudo-code
// X3 = Src1[127:96]
+83 -96
View File
@@ -20,99 +20,90 @@ DEF_OP(F80LOADFCW) {
DEF_OP(F80ADD) {
auto Op = IROp->C<IR::IROp_F80Add>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FADD(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FADD(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SUB) {
auto Op = IROp->C<IR::IROp_F80Sub>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSUB(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FSUB(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80MUL) {
auto Op = IROp->C<IR::IROp_F80Mul>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FMUL(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FMUL(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80DIV) {
auto Op = IROp->C<IR::IROp_F80Div>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FDIV(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FDIV(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FYL2X) {
auto Op = IROp->C<IR::IROp_F80FYL2X>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FYL2X(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FYL2X(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80ATAN) {
auto Op = IROp->C<IR::IROp_F80ATAN>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FATAN(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FATAN(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FPREM1) {
auto Op = IROp->C<IR::IROp_F80FPREM1>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FREM1(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FREM1(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80FPREM) {
auto Op = IROp->C<IR::IROp_F80FPREM>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FREM(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FREM(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SCALE) {
auto Op = IROp->C<IR::IROp_F80SCALE>();
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSCALE(Src1, Src2);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
const auto Tmp = X80SoftFloat::FSCALE(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80CVT) {
auto Op = IROp->C<IR::IROp_F80CVT>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
switch (OpSize) {
case 4: {
@@ -131,9 +122,9 @@ DEF_OP(F80CVT) {
DEF_OP(F80CVTINT) {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
switch (OpSize) {
case 2: {
@@ -160,13 +151,13 @@ DEF_OP(F80CVTTO) {
switch (Op->SrcSize) {
case 4: {
float Src = *GetSrc<float *>(Data->SSAData, Op->Header.Args[0]);
float Src = *GetSrc<float *>(Data->SSAData, Op->X80Src);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
case 8: {
double Src = *GetSrc<double *>(Data->SSAData, Op->Header.Args[0]);
double Src = *GetSrc<double *>(Data->SSAData, Op->X80Src);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
@@ -180,13 +171,13 @@ DEF_OP(F80CVTTOINT) {
switch (Op->SrcSize) {
case 2: {
int16_t Src = *GetSrc<int16_t*>(Data->SSAData, Op->Header.Args[0]);
int16_t Src = *GetSrc<int16_t*>(Data->SSAData, Op->Src);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
}
case 4: {
int32_t Src = *GetSrc<int32_t*>(Data->SSAData, Op->Header.Args[0]);
int32_t Src = *GetSrc<int32_t*>(Data->SSAData, Op->Src);
X80SoftFloat Tmp = Src;
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
break;
@@ -197,77 +188,73 @@ DEF_OP(F80CVTTOINT) {
DEF_OP(F80ROUND) {
auto Op = IROp->C<IR::IROp_F80Round>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FRNDINT(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FRNDINT(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80F2XM1) {
auto Op = IROp->C<IR::IROp_F80F2XM1>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::F2XM1(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::F2XM1(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80TAN) {
auto Op = IROp->C<IR::IROp_F80TAN>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FTAN(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FTAN(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SQRT) {
auto Op = IROp->C<IR::IROp_F80SQRT>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSQRT(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FSQRT(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80SIN) {
auto Op = IROp->C<IR::IROp_F80SIN>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FSIN(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FSIN(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80COS) {
auto Op = IROp->C<IR::IROp_F80COS>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FCOS(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FCOS(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80XTRACT_EXP) {
auto Op = IROp->C<IR::IROp_F80XTRACT_EXP>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FXTRACT_EXP(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FXTRACT_EXP(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80XTRACT_SIG) {
auto Op = IROp->C<IR::IROp_F80XTRACT_SIG>();
X80SoftFloat Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Tmp;
Tmp = X80SoftFloat::FXTRACT_SIG(Src);
const auto Src = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src);
const auto Tmp = X80SoftFloat::FXTRACT_SIG(Src);
memcpy(GDP, &Tmp, sizeof(X80SoftFloat));
}
DEF_OP(F80CMP) {
auto Op = IROp->C<IR::IROp_F80Cmp>();
uint32_t ResultFlags{};
X80SoftFloat Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]);
X80SoftFloat Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[1]);
const auto Src1 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src1);
const auto Src2 = *GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src2);
bool eq, lt, nan;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Op->Flags & (1 << IR::FCMP_FLAG_LT) &&
@@ -288,7 +275,7 @@ DEF_OP(F80CMP) {
DEF_OP(F80BCDLOAD) {
auto Op = IROp->C<IR::IROp_F80BCDLoad>();
uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
const uint8_t *Src1 = GetSrc<uint8_t*>(Data->SSAData, Op->X80Src);
uint64_t BCD{};
// We walk through each uint8_t and pull out the BCD encoding
// Each 4bit split is a digit
@@ -323,7 +310,7 @@ DEF_OP(F80BCDLOAD) {
DEF_OP(F80BCDSTORE) {
auto Op = IROp->C<IR::IROp_F80BCDStore>();
X80SoftFloat Src1 = X80SoftFloat::FRNDINT(*GetSrc<X80SoftFloat*>(Data->SSAData, Op->Header.Args[0]));
X80SoftFloat Src1 = X80SoftFloat::FRNDINT(*GetSrc<X80SoftFloat*>(Data->SSAData, Op->X80Src));
bool Negative = Src1.Sign;
// Clear the Sign bit
@@ -358,74 +345,74 @@ DEF_OP(F80BCDSTORE) {
DEF_OP(F64SIN) {
auto Op = IROp->C<IR::IROp_F64SIN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = sin(Src);
const double Src = *GetSrc<double*>(Data->SSAData, Op->Src);
const double Tmp = sin(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64COS) {
auto Op = IROp->C<IR::IROp_F64COS>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = cos(Src);
const double Src = *GetSrc<double*>(Data->SSAData, Op->Src);
const double Tmp = cos(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64TAN) {
auto Op = IROp->C<IR::IROp_F64TAN>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = tan(Src);
const double Src = *GetSrc<double*>(Data->SSAData, Op->Src);
const double Tmp = tan(Src);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64F2XM1) {
auto Op = IROp->C<IR::IROp_F64F2XM1>();
double Src = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Tmp = exp2(Src) - 1.0;
const double Src = *GetSrc<double*>(Data->SSAData, Op->Src);
const double Tmp = exp2(Src) - 1.0;
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64ATAN) {
auto Op = IROp->C<IR::IROp_F64ATAN>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = atan2(Src1, Src2);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Src1);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Src2);
const double Tmp = atan2(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM) {
auto Op = IROp->C<IR::IROp_F64FPREM>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = fmod(Src1, Src2);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Src1);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Src2);
const double Tmp = fmod(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FPREM1) {
auto Op = IROp->C<IR::IROp_F64FPREM1>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = remainder(Src1, Src2);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Src1);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Src2);
const double Tmp = remainder(Src1, Src2);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64FYL2X) {
auto Op = IROp->C<IR::IROp_F64FYL2X>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double Tmp = Src2 * log2(Src1);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Src);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Src2);
const double Tmp = Src2 * log2(Src1);
memcpy(GDP, &Tmp, sizeof(double));
}
DEF_OP(F64SCALE) {
auto Op = IROp->C<IR::IROp_F64SCALE>();
double Src1 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[0]);
double Src2 = *GetSrc<double*>(Data->SSAData, Op->Header.Args[1]);
double trunc = (double)(int64_t)(Src2); //truncate
double Tmp = Src1 * exp2(trunc);
const double Src1 = *GetSrc<double*>(Data->SSAData, Op->Src1);
const double Src2 = *GetSrc<double*>(Data->SSAData, Op->Src2);
const double trunc = (double)(int64_t)(Src2); //truncate
const double Tmp = Src1 * exp2(trunc);
memcpy(GDP, &Tmp, sizeof(double));
}
@@ -14,7 +14,7 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
GD = (*GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]) >> Op->Flag) & 1;
GD = (*GetSrc<uint64_t*>(Data->SSAData, Op->Value) >> Op->Flag) & 1;
}
#undef DEF_OP
@@ -159,21 +159,26 @@
break; \
}
struct InterpVector256 {
__uint128_t Lower;
__uint128_t Upper;
};
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.ID().Value];
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::NodeID Op) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Op.Value];
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<__uint128_t*>(SSAData)[Src.ID().Value];
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Src.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
@@ -146,7 +146,7 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
@@ -1,5 +1,6 @@
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUID.h"
#include "InterpreterDefines.h"
#include "InterpreterOps.h"
#include "F80Ops.h"
@@ -121,7 +122,7 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
// Conversion ops
@@ -153,8 +154,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
@@ -180,13 +179,10 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
// Move ops
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
// Vector ops
REGISTER_OP(VECTORZERO, VectorZero);
REGISTER_OP(VECTORIMM, VectorImm);
REGISTER_OP(SPLATVECTOR2, SplatVector);
REGISTER_OP(SPLATVECTOR4, SplatVector);
REGISTER_OP(VMOV, VMov);
REGISTER_OP(VAND, VAnd);
REGISTER_OP(VBIC, VBic);
@@ -245,18 +241,13 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VUSHRS, VUShrS);
REGISTER_OP(VSSHRS, VSShrS);
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VINSSCALARELEMENT, VInsScalarElement);
REGISTER_OP(VEXTRACTELEMENT, VExtractElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
REGISTER_OP(VUSHRNI, VUShrNI);
REGISTER_OP(VUSHRNI2, VUShrNI2);
REGISTER_OP(VBITCAST, VBitcast);
REGISTER_OP(VSXTL, VSXTL);
REGISTER_OP(VSXTL2, VSXTL2);
REGISTER_OP(VUXTL, VUXTL);
@@ -336,29 +327,34 @@ void InterpreterOps::Op_NoOp(FEXCore::IR::IROp_Header *IROp, IROpData *Data, IR:
void InterpreterOps::InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *CurrentIR) {
volatile void *StackEntry = alloca(0);
uintptr_t ListSize = CurrentIR->GetSSACount();
const uintptr_t ListSize = CurrentIR->GetSSACount();
static_assert(sizeof(FEXCore::IR::IROp_Header) == 4);
static_assert(sizeof(FEXCore::IR::OrderedNode) == 16);
auto BlockEnd = CurrentIR->GetBlocks().end();
InterpreterOps::IROpData OpData{};
OpData.State = Frame->Thread;
OpData.SSAData = alloca(ListSize * 16);
OpData.CurrentEntry = Frame->State.rip;
OpData.CurrentIR = CurrentIR;
OpData.StackEntry = StackEntry;
OpData.BlockIterator = CurrentIR->GetBlocks().begin();
constexpr size_t ListEntrySizeInBytes = sizeof(InterpVector256);
const size_t SSADataSize = ListSize * ListEntrySizeInBytes;
// Clear them all to zero. Required for Zero-extend semantics
memset(OpData.SSAData, 0, ListSize * 16);
InterpreterOps::IROpData OpData{
.State = Frame->Thread,
.CurrentEntry = Frame->State.rip,
.CurrentIR = CurrentIR,
.StackEntry = StackEntry,
.SSAData = alloca(SSADataSize),
.BlockResults = {},
.BlockIterator = CurrentIR->GetBlocks().begin(),
};
// Clear all SSAData entries to zero. Required for Zero-extend semantics
memset(OpData.SSAData, 0, SSADataSize);
while (1) {
using namespace FEXCore::IR;
auto [BlockNode, BlockHeader] = OpData.BlockIterator();
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
// Reset the block results per block
memset(&OpData.BlockResults, 0, sizeof(OpData.BlockResults));
@@ -49,7 +49,7 @@ namespace FEXCore::CPU {
public:
static void InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *IR);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
static bool GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info);
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
@@ -151,7 +151,7 @@ namespace FEXCore::CPU {
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -181,8 +181,6 @@ namespace FEXCore::CPU {
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -207,7 +205,6 @@ namespace FEXCore::CPU {
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
@@ -264,18 +261,13 @@ namespace FEXCore::CPU {
DEF_OP(VUShrS);
DEF_OP(VSShrS);
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
DEF_OP(VUShrNI);
DEF_OP(VUShrNI2);
DEF_OP(VBitcast);
DEF_OP(VSXTL);
DEF_OP(VSXTL2);
DEF_OP(VUXTL);
@@ -25,93 +25,139 @@ static inline void CacheLineFlush(char *Addr) {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_LoadContext>();
const auto OpSize = IROp->Size;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Src = ContextPtr + Op->Offset;
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->Offset;
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(ContextPtr); \
y const *MemData = reinterpret_cast<y const*>(Src); \
GD = *MemData; \
break; \
}
switch (OpSize) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16: {
void const *MemData = reinterpret_cast<void const*>(ContextPtr);
case 16:
case 32: {
void const *MemData = reinterpret_cast<void const*>(Src);
memcpy(GDP, MemData, OpSize);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
break;
}
#undef LOAD_CTX
}
DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_StoreContext>();
const auto OpSize = IROp->Size;
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->Offset;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Dst = ContextPtr + Op->Offset;
void *MemData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *MemData = reinterpret_cast<void*>(Dst);
void *Src = GetSrc<void*>(Data->SSAData, Op->Value);
memcpy(MemData, Src, OpSize);
}
DEF_OP(LoadRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
DEF_OP(StoreRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Src = ContextPtr + Op->Offset;
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(ContextPtr); \
y const *MemData = reinterpret_cast<y const*>(Src); \
GD = *MemData; \
break; \
}
switch (IROp->Size) {
switch (OpSize) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16: {
void const *MemData = reinterpret_cast<void const*>(ContextPtr);
memcpy(GDP, MemData, IROp->Size);
case 16:
case 32: {
void const *MemData = reinterpret_cast<void const*>(Src);
memcpy(GDP, MemData, OpSize);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
break;
}
#undef LOAD_CTX
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Dst = ContextPtr + Op->Offset;
void *MemData = reinterpret_cast<void*>(Dst);
void *Src = GetSrc<void*>(Data->SSAData, Op->Value);
memcpy(MemData, Src, OpSize);
}
DEF_OP(LoadContextIndexed) {
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
const auto Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Index);
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Src = ContextPtr + Op->BaseOffset + (Index * Op->Stride);
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(Src); \
GD = *MemData; \
break; \
}
switch (OpSize) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16:
case 32: {
void const *MemData = reinterpret_cast<void const*>(Src);
memcpy(GDP, MemData, OpSize);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", OpSize);
break;
}
#undef LOAD_CTX
}
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
uint64_t Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[1]);
const auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
const auto OpSize = IROp->Size;
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += Op->BaseOffset;
ContextPtr += Index * Op->Stride;
const auto Index = *GetSrc<uint64_t*>(Data->SSAData, Op->Index);
void *MemData = reinterpret_cast<void*>(ContextPtr);
void *Src = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
memcpy(MemData, Src, IROp->Size);
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Dst = ContextPtr + Op->BaseOffset + (Index * Op->Stride);
void *MemData = reinterpret_cast<void*>(Dst);
void *Src = GetSrc<void*>(Data->SSAData, Op->Value);
memcpy(MemData, Src, OpSize);
}
DEF_OP(SpillRegister) {
@@ -134,7 +180,7 @@ DEF_OP(LoadFlag) {
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
uint8_t Arg = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
uint8_t Arg = *GetSrc<uint8_t*>(Data->SSAData, Op->Value);
uintptr_t ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
ContextPtr += offsetof(FEXCore::Core::CPUState, flags[0]);
@@ -144,8 +190,8 @@ DEF_OP(StoreFlag) {
}
DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_LoadMem>();
const auto OpSize = IROp->Size;
uint8_t const *MemData = *GetSrc<uint8_t const**>(Data->SSAData, Op->Addr);
@@ -158,7 +204,8 @@ DEF_OP(LoadMem) {
case IR::MEM_OFFSET_SXTW.Val: MemData += (int32_t)Offset; break;
}
}
memset(GDP, 0, 16);
memset(GDP, 0, Core::CPUState::XMM_AVX_REG_SIZE);
switch (OpSize) {
case 1: {
auto D = reinterpret_cast<const std::atomic<uint8_t>*>(MemData);
@@ -180,16 +227,15 @@ DEF_OP(LoadMem) {
GD = D->load();
break;
}
default:
memcpy(GDP, MemData, IROp->Size);
memcpy(GDP, MemData, OpSize);
break;
}
}
DEF_OP(StoreMem) {
auto Op = IROp->C<IR::IROp_StoreMem>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_StoreMem>();
const auto OpSize = IROp->Size;
uint8_t *MemData = *GetSrc<uint8_t **>(Data->SSAData, Op->Addr);
@@ -221,41 +267,11 @@ DEF_OP(StoreMem) {
}
default:
memcpy(MemData, GetSrc<void*>(Data->SSAData, Op->Value), IROp->Size);
memcpy(MemData, GetSrc<void*>(Data->SSAData, Op->Value), OpSize);
break;
}
}
DEF_OP(VLoadMemElement) {
auto Op = IROp->C<IR::IROp_VLoadMemElement>();
void const *MemData = *GetSrc<void const**>(Data->SSAData, Op->Header.Args[0]);
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[1]), 16);
memcpy(reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(GDP) + (Op->Header.ElementSize * Op->Index)),
MemData, Op->Header.ElementSize);
}
DEF_OP(VStoreMemElement) {
#define STORE_DATA(x, y) \
case x: { \
y *MemData = *GetSrc<y**>(Data->SSAData, Op->Header.Args[0]); \
memcpy(MemData, &GetSrc<y*>(Data->SSAData, Op->Header.Args[1])[Op->Index], sizeof(y)); \
break; \
}
auto Op = IROp->C<IR::IROp_VStoreMemElement>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
STORE_DATA(1, uint8_t)
STORE_DATA(2, uint16_t)
STORE_DATA(4, uint32_t)
STORE_DATA(8, uint64_t)
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size"); break;
}
#undef STORE_DATA
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -19,13 +19,6 @@ $end_info$
#include <sys/random.h>
namespace FEXCore::CPU {
[[noreturn]]
static void StopThread(FEXCore::Core::InternalThreadState *Thread) {
Thread->CTX->StopThread(Thread);
LOGMAN_MSG_A_FMT("unreachable");
FEX_UNREACHABLE;
}
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(Fence) {
@@ -46,14 +39,26 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case FEXCore::IR::Break_Halt: // HLT
StopThread(Data->State);
Data->State->CurrentFrame->SynchronousFaultData.FaultToTopAndGeneratedException = 1;
Data->State->CurrentFrame->SynchronousFaultData.Signal = Op->Reason.Signal;
Data->State->CurrentFrame->SynchronousFaultData.TrapNo = Op->Reason.TrapNumber;
Data->State->CurrentFrame->SynchronousFaultData.err_code = Op->Reason.ErrorRegister;
Data->State->CurrentFrame->SynchronousFaultData.si_code = Op->Reason.si_code;
switch (Op->Reason.Signal) {
case SIGILL:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
break;
case FEXCore::IR::Break_InvalidInstruction:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGILL);
case SIGTRAP:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGTRAP);
break;
case SIGSEGV:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGSEGV);
break;
default:
FHU::Syscalls::tgkill(Data->State->ThreadManager.PID, Data->State->ThreadManager.TID, SIGTRAP);
break;
default: LOGMAN_MSG_A_FMT("Unknown Break Reason: {}", Op->Reason); break;
}
}
@@ -88,7 +93,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
uint8_t GuestRounding = *GetSrc<uint8_t*>(Data->SSAData, Op->Header.Args[0]);
const auto GuestRounding = *GetSrc<uint8_t*>(Data->SSAData, Op->RoundMode);
#ifdef _M_ARM_64
uint64_t HostRounding{};
__asm volatile(R"(
@@ -128,16 +133,16 @@ DEF_OP(SetRoundingMode) {
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
uint8_t OpSize = IROp->Size;
const uint8_t OpSize = IROp->Size;
if (OpSize <= 8) {
uint64_t Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Header.Args[0]);
const auto Src = *GetSrc<uint64_t*>(Data->SSAData, Op->Value);
LogMan::Msg::IFmt(">>>> Value in Arg: 0x{:x}, {}", Src, Src);
}
else if (OpSize == 16) {
__uint128_t Src = *GetSrc<__uint128_t*>(Data->SSAData, Op->Header.Args[0]);
uint64_t Src0 = Src;
uint64_t Src1 = Src >> 64;
const auto Src = *GetSrc<__uint128_t*>(Data->SSAData, Op->Value);
const uint64_t Src0 = Src;
const uint64_t Src1 = Src >> 64;
LogMan::Msg::IFmt(">>>> Value[0] in Arg: 0x{:x}, {}", Src0, Src0);
LogMan::Msg::IFmt(" Value[1] in Arg: 0x{:x}, {}", Src1, Src1);
}
@@ -14,15 +14,15 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void InterpreterOps::Op_##x(IR::IROp_Header *IROp, IROpData *Data, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
uintptr_t Src = GetSrc<uintptr_t>(Data->SSAData, Op->Header.Args[0]);
const auto Src = GetSrc<uintptr_t>(Data->SSAData, Op->Pair);
memcpy(GDP,
reinterpret_cast<void*>(Src + Op->Header.Size * Op->Element), Op->Header.Size);
}
DEF_OP(CreateElementPair) {
auto Op = IROp->C<IR::IROp_CreateElementPair>();
void *Src_Lower = GetSrc<void*>(Data->SSAData, Op->Header.Args[0]);
void *Src_Upper = GetSrc<void*>(Data->SSAData, Op->Header.Args[1]);
const void *Src_Lower = GetSrc<void*>(Data->SSAData, Op->Lower);
const void *Src_Upper = GetSrc<void*>(Data->SSAData, Op->Upper);
uint8_t *Dst = GetDest<uint8_t*>(Data->SSAData, Node);
@@ -30,13 +30,6 @@ DEF_OP(CreateElementPair) {
memcpy(Dst + IROp->ElementSize, Src_Upper, IROp->ElementSize);
}
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
uint8_t OpSize = IROp->Size;
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Header.Args[0]), OpSize);
}
#undef DEF_OP
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
+76 -22
View File
@@ -14,7 +14,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -1084,8 +1084,8 @@ DEF_OP(Bfi) {
DEF_OP(Bfe) {
auto Op = IROp->C<IR::IROp_Bfe>();
LOGMAN_THROW_A_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
LOGMAN_THROW_A_FMT(Op->Width != 0, "Invalid BFE width of 0");
LOGMAN_THROW_AA_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
LOGMAN_THROW_AA_FMT(Op->Width != 0, "Invalid BFE width of 0");
auto Dst = GetReg<RA_64>(Node);
ubfx(Dst, GetReg<RA_64>(Op->Src.ID()), Op->lsb, Op->Width);
@@ -1168,25 +1168,79 @@ DEF_OP(Select) {
}
DEF_OP(VExtractToGPR) {
auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VExtractToGPR>();
const auto OpSize = IROp->Size;
switch (OpSize) {
case 1:
umov(GetReg<RA_32>(Node), GetSrc(Op->Vector.ID()).V16B(), Op->Index);
break;
case 2:
umov(GetReg<RA_32>(Node), GetSrc(Op->Vector.ID()).V8H(), Op->Index);
break;
case 4:
umov(GetReg<RA_32>(Node), GetSrc(Op->Vector.ID()).V4S(), Op->Index);
break;
case 8:
umov(GetReg<RA_64>(Node), GetSrc(Op->Vector.ID()).V2D(), Op->Index);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", OpSize);
break;
constexpr auto AVXRegBitSize = Core::CPUState::XMM_AVX_REG_SIZE * 8;
constexpr auto SSERegBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto ElementSizeBits = Op->Header.ElementSize * 8;
const auto Offset = ElementSizeBits * Op->Index;
const auto Is256Bit = Offset >= SSERegBitSize;
const auto Vector = GetSrc(Op->Vector.ID());
const auto PerformMove = [&](const aarch64::VRegister& reg, int index) {
switch (OpSize) {
case 1:
umov(GetReg<RA_32>(Node), reg.V16B(), index);
break;
case 2:
umov(GetReg<RA_32>(Node), reg.V8H(), index);
break;
case 4:
umov(GetReg<RA_32>(Node), reg.V4S(), index);
break;
case 8:
umov(GetReg<RA_64>(Node), reg.V2D(), index);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", OpSize);
break;
}
};
if (Offset < SSERegBitSize) {
// Desired data lies within the lower 128-bit lane, so we
// can treat the operation as a 128-bit operation, even
// when acting on larger register sizes.
PerformMove(Vector, Op->Index);
} else {
LOGMAN_THROW_AA_FMT(HostSupportsSVE,
"Host doesn't support SVE. Cannot perform 256-bit operation.");
LOGMAN_THROW_AA_FMT(Is256Bit,
"Can't perform 256-bit extraction with op side: {}", OpSize);
LOGMAN_THROW_AA_FMT(Offset < AVXRegBitSize,
"Trying to extract element outside bounds of register. Offset={}, Index={}",
Offset, Op->Index);
// We need to use the upper 128-bit lane, so lets move it down.
// Inverting our dedicated predicate for 128-bit operations selects
// all of the top lanes. We can then compact those into a temporary.
const auto CompactPred = p0;
not_(CompactPred.VnB(), PRED_TMP_32B.Zeroing(), PRED_TMP_16B.VnB());
compact(VTMP1.Z().VnD(), CompactPred, Vector.Z().VnD());
// Sanitize the zero-based index to work on the now-moved
// upper half of the vector.
const auto SanitizedIndex = [OpSize, Op] {
switch (OpSize) {
case 1:
return Op->Index - 16;
case 2:
return Op->Index - 8;
case 4:
return Op->Index - 4;
case 8:
return Op->Index - 2;
default:
LOGMAN_MSG_A_FMT("Unhandled OpSize: {}", OpSize);
return 0;
}
}();
// Move the value from the now-low-lane data.
PerformMove(VTMP1, SanitizedIndex);
}
}
@@ -1249,7 +1303,7 @@ DEF_OP(FCmp) {
bool set = false;
if (Op->Flags & (1 << IR::FCMP_FLAG_EQ)) {
LOGMAN_THROW_A_FMT(IR::FCMP_FLAG_EQ == 0, "IR::FCMP_FLAG_EQ must equal 0");
LOGMAN_THROW_AA_FMT(IR::FCMP_FLAG_EQ == 0, "IR::FCMP_FLAG_EQ must equal 0");
// EQ or unordered
cset(Dst, Condition::eq); // Z = 1
csinc(Dst, Dst, xzr, Condition::vc); // IF !V ? Z : 1
@@ -81,7 +81,7 @@ bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uin
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
LOGMAN_THROW_AA_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
// Size is the size of each pair element
@@ -19,7 +19,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
// First we must reset the stack
@@ -40,7 +40,7 @@ DEF_OP(CallbackReturn) {
ResetStack();
// We can now lower the ref counter again
ldr(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
sub(w2, w2, 1);
str(w2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SignalHandlerRefCounter)));
@@ -197,7 +197,11 @@ DEF_OP(Syscall) {
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SyscallHandlerFunc)));
mov(x1, STATE);
mov(x2, sp);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, void*, void*, void*>(x3);
#else
blr(x3);
#endif
add(sp, sp, SPOffset);
@@ -239,7 +243,6 @@ DEF_OP(InlineSyscall) {
bool Intersects{};
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
std::vector<vixl::aarch64::Register> IntersectRegs(FEXCore::HLE::SyscallArguments::MAX_ARGS);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
@@ -381,7 +384,11 @@ DEF_OP(Thunk) {
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
LoadConstant(x2, (uintptr_t)thunkFn);
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -436,7 +443,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveThreadCodeEntry) {
DEF_OP(ThreadRemoveCodeEntry) {
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
@@ -446,9 +453,13 @@ DEF_OP(RemoveThreadCodeEntry) {
mov(x0, STATE);
LoadConstant(x1, Entry);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)));
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT)));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, void*, void*>(x2);
#else
blr(x2);
#endif
FillStaticRegs();
// Fix the stack and any values that were stepped on
@@ -459,6 +470,7 @@ DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR();
SpillStaticRegs();
// x0 = CPUID Handler
// x1 = CPUID Function
@@ -467,10 +479,13 @@ DEF_OP(CPUID) {
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction)));
mov(x1, GetReg<RA_64>(Op->Function.ID()));
mov(x2, GetReg<RA_64>(Op->Leaf.ID()));
SpillStaticRegs();
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, void*, uint64_t, uint64_t>(x3);
#else
blr(x3);
FillStaticRegs();
#endif
FillStaticRegs();
PopDynamicRegsAndLR();
// Results are in x0, x1
@@ -492,7 +507,7 @@ void Arm64JITCore::RegisterBranchHandlers() {
REGISTER_OP(INLINESYSCALL, InlineSyscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -10,28 +10,117 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
mov(GetDst(Node), GetSrc(Op->DestVector.ID()));
switch (Op->Header.ElementSize) {
case 1: {
ins(GetDst(Node).V16B(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
const auto DestIdx = Op->DestIdx;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto DestVector = GetSrc(Op->DestVector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto ElementSizeBits = ElementSize * 8;
const auto Offset = ElementSizeBits * DestIdx;
const auto SSEBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto InUpperLane = Offset >= SSEBitSize;
// This is going to be a little gross. Pls forgive me.
// Since SVE has the whole vector length agnostic programming
// thing going on, we can't exactly freely insert entries into
// arbitrary locations in the vector.
//
// SVE *does* have INSR, however this only shifts the entire
// vector to the left by an element size and inserts a value
// at the beginning of the vector. Not *quite* what we need.
// (though INSR *is* very useful for other things).
//
// The idea is (in the case of the upper lane), move the upper
// lane down, insert into it and recombine with the lower lane.
//
// In the case of the lower lane, insert and then recombine with
// the upper lane.
if (InUpperLane) {
// Move the upper lane down for the insertion.
const auto CompactPred = p0;
not_(CompactPred.VnB(), PRED_TMP_32B.Zeroing(), PRED_TMP_16B.VnB());
compact(VTMP1.Z().VnD(), CompactPred, DestVector.Z().VnD());
}
case 2: {
ins(GetDst(Node).V8H(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
// Put data in place for destructive SPLICE below.
mov(Dst.Z().VnD(), DestVector.Z().VnD());
// Inserts the GPR value into the given V register.
// Also automatically adjusts the index in the case of using the
// moved upper lane.
const auto Insert = [&](const aarch64::VRegister& reg, int index) {
switch (ElementSize) {
case 1:
if (InUpperLane) {
index -= 16;
}
ins(reg.V16B(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 2:
if (InUpperLane) {
index -= 8;
}
ins(reg.V8H(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 4:
if (InUpperLane) {
index -= 4;
}
ins(reg.V4S(), index, GetReg<RA_32>(Op->Src.ID()));
break;
case 8:
if (InUpperLane) {
index -= 2;
}
ins(reg.V2D(), index, GetReg<RA_64>(Op->Src.ID()));
break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
};
if (InUpperLane) {
Insert(VTMP1, DestIdx);
splice(Dst.Z().VnD(), PRED_TMP_16B, Dst.Z().VnD(), VTMP1.Z().VnD());
} else {
Insert(Dst, DestIdx);
splice(Dst.Z().VnD(), PRED_TMP_16B, Dst.Z().VnD(), DestVector.Z().VnD());
}
case 4: {
ins(GetDst(Node).V4S(), Op->DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
} else {
mov(Dst, DestVector);
switch (ElementSize) {
case 1: {
ins(Dst.V16B(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 2: {
ins(Dst.V8H(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 4: {
ins(Dst.V4S(), DestIdx, GetReg<RA_32>(Op->Src.ID()));
break;
}
case 8: {
ins(Dst.V2D(), DestIdx, GetReg<RA_64>(Op->Src.ID()));
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
case 8: {
ins(GetDst(Node).V2D(), Op->DestIdx, GetReg<RA_64>(Op->Src.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
}
}
@@ -57,8 +146,11 @@ DEF_OP(VCastFromGPR) {
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
const uint16_t ElementSize = Op->Header.ElementSize;
const uint16_t Conv = (ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
scvtf(GetDst(Node).S(), GetReg<RA_32>(Op->Src.ID()));
@@ -76,6 +168,10 @@ DEF_OP(Float_FromGPR_S) {
scvtf(GetDst(Node).D(), GetReg<RA_64>(Op->Src.ID()));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled conversion mask: Mask=0x{:04x}, ElementSize={}, SrcElementSize={}",
Conv, ElementSize, Op->SrcElementSize);
break;
}
}
@@ -96,116 +192,379 @@ DEF_OP(Float_FToF) {
}
DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
switch (Op->Header.ElementSize) {
case 4:
scvtf(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
scvtf(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", Op->Header.ElementSize);
const auto Op = IROp->C<IR::IROp_Vector_SToF>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
scvtf(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
scvtf(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
scvtf(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", ElementSize);
break;
}
} else {
switch (ElementSize) {
case 2:
scvtf(Dst.V8H(), Vector.V8H());
break;
case 4:
scvtf(Dst.V4S(), Vector.V4S());
break;
case 8:
scvtf(Dst.V2D(), Vector.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", ElementSize);
break;
}
}
}
DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
switch (Op->Header.ElementSize) {
case 4:
fcvtzs(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
break;
case 8:
fcvtzs(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", Op->Header.ElementSize);
const auto Op = IROp->C<IR::IROp_Vector_FToZS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
fcvtzs(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
fcvtzs(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
fcvtzs(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", ElementSize);
break;
}
} else {
switch (ElementSize) {
case 2:
fcvtzs(Dst.V8H(), Vector.V8H());
break;
case 4:
fcvtzs(Dst.V4S(), Vector.V4S());
break;
case 8:
fcvtzs(Dst.V2D(), Vector.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", ElementSize);
break;
}
}
}
DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
fcvtzs(GetDst(Node).V4S(), GetDst(Node).V4S());
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
fcvtzs(GetDst(Node).V2D(), GetDst(Node).V2D());
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", Op->Header.ElementSize);
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (ElementSize) {
case 2:
frinti(Dst.Z().VnH(), Mask, Vector.Z().VnH());
fcvtzs(Dst.Z().VnH(), Mask, Dst.Z().VnH());
break;
case 4:
frinti(Dst.Z().VnS(), Mask, Vector.Z().VnS());
fcvtzs(Dst.Z().VnS(), Mask, Dst.Z().VnS());
break;
case 8:
frinti(Dst.Z().VnD(), Mask, Vector.Z().VnD());
fcvtzs(Dst.Z().VnD(), Mask, Dst.Z().VnD());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", ElementSize);
break;
}
} else {
switch (ElementSize) {
case 2:
frinti(Dst.V8H(), Vector.V8H());
fcvtzs(Dst.V8H(), Dst.V8H());
break;
case 4:
frinti(Dst.V4S(), Vector.V4S());
fcvtzs(Dst.V4S(), Dst.V4S());
break;
case 8:
frinti(Dst.V2D(), Vector.V2D());
fcvtzs(Dst.V2D(), Dst.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", ElementSize);
break;
}
}
}
DEF_OP(Vector_FToF) {
auto Op = IROp->C<IR::IROp_Vector_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const auto Op = IROp->C<IR::IROp_Vector_FToF>();
const auto OpSize = IROp->Size;
switch (Conv) {
case 0x0804: { // Double <- Float
fcvtl(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2S());
break;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
// Curiously, FCVTLT and FCVTNT have no bottom variants,
// and also interesting is that FCVTLT will iterate the
// source vector by accessing each odd element and storing
// them consecutively in the destination.
//
// FCVTNT is somewhat like the opposite. It will read each
// consecutive element, but store each result into every odd
// element in the destination vector.
//
// We need to undo the behavior of FCVTNT with UZP2. In the case
// of FCVTLT, we instead need to set the vector up with ZIP1, so
// that the elements will be processed correctly.
const auto Mask = PRED_TMP_32B.Merging();
switch (Conv) {
case 0x0402: { // Float <- Half
zip1(Dst.Z().VnH(), Vector.Z().VnH(), Vector.Z().VnH());
fcvtlt(Dst.Z().VnS(), Mask, Dst.Z().VnH());
break;
}
case 0x0804: { // Double <- Float
zip1(Dst.Z().VnS(), Vector.Z().VnS(), Vector.Z().VnS());
fcvtlt(Dst.Z().VnD(), Mask, Dst.Z().VnS());
break;
}
case 0x0204: { // Half <- Float
fcvtnt(Dst.Z().VnH(), Mask, Vector.Z().VnS());
uzp2(Dst.Z().VnH(), Dst.Z().VnH(), Dst.Z().VnH());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(Dst.Z().VnS(), Mask, Vector.Z().VnD());
uzp2(Dst.Z().VnS(), Dst.Z().VnS(), Dst.Z().VnS());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv);
break;
}
case 0x0408: { // Float <- Double
fcvtn(GetDst(Node).V2S(), GetSrc(Op->Vector.ID()).V2D());
break;
} else {
switch (Conv) {
case 0x0402: { // Float <- Half
fcvtl(Dst.V4S(), Vector.V4H());
break;
}
case 0x0804: { // Double <- Float
fcvtl(Dst.V2D(), Vector.V2S());
break;
}
case 0x0204: { // Half <- Float
fcvtn(Dst.V4H(), Vector.V4S());
break;
}
case 0x0408: { // Float <- Double
fcvtn(Dst.V2S(), Vector.V2D());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
}
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (Op->Header.ElementSize) {
case 4:
frintn(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
const auto Op = IROp->C<IR::IROp_Vector_FToI>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (ElementSize) {
case 2:
frintn(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintn(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintn(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
break;
case 8:
frintn(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (ElementSize) {
case 2:
frintm(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintm(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintm(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
break;
}
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintm(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (ElementSize) {
case 2:
frintp(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintp(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintp(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
break;
case 8:
frintm(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
case FEXCore::IR::Round_Towards_Zero.Val:
switch (ElementSize) {
case 2:
frintz(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frintz(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frintz(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
break;
}
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (Op->Header.ElementSize) {
case 4:
frintp(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
case FEXCore::IR::Round_Host.Val:
switch (ElementSize) {
case 2:
frinti(Dst.Z().VnH(), Mask, Vector.Z().VnH());
break;
case 4:
frinti(Dst.Z().VnS(), Mask, Vector.Z().VnS());
break;
case 8:
frinti(Dst.Z().VnD(), Mask, Vector.Z().VnD());
break;
}
break;
case 8:
frintp(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
}
} else {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
switch (ElementSize) {
case 2:
frintn(Dst.V8H(), Vector.V8H());
break;
case 4:
frintn(Dst.V4S(), Vector.V4S());
break;
case 8:
frintn(Dst.V2D(), Vector.V2D());
break;
}
break;
}
break;
case FEXCore::IR::Round_Towards_Zero.Val:
switch (Op->Header.ElementSize) {
case 4:
frintz(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
case FEXCore::IR::Round_Negative_Infinity.Val:
switch (ElementSize) {
case 2:
frintm(Dst.V8H(), Vector.V8H());
break;
case 4:
frintm(Dst.V4S(), Vector.V4S());
break;
case 8:
frintm(Dst.V2D(), Vector.V2D());
break;
}
break;
case 8:
frintz(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
case FEXCore::IR::Round_Positive_Infinity.Val:
switch (ElementSize) {
case 2:
frintp(Dst.V8H(), Vector.V8H());
break;
case 4:
frintp(Dst.V4S(), Vector.V4S());
break;
case 8:
frintp(Dst.V2D(), Vector.V2D());
break;
}
break;
}
break;
case FEXCore::IR::Round_Host.Val:
switch (Op->Header.ElementSize) {
case 4:
frinti(GetDst(Node).V4S(), GetSrc(Op->Vector.ID()).V4S());
case FEXCore::IR::Round_Towards_Zero.Val:
switch (ElementSize) {
case 2:
frintz(Dst.V8H(), Vector.V8H());
break;
case 4:
frintz(Dst.V4S(), Vector.V4S());
break;
case 8:
frintz(Dst.V2D(), Vector.V2D());
break;
}
break;
case 8:
frinti(GetDst(Node).V2D(), GetSrc(Op->Vector.ID()).V2D());
case FEXCore::IR::Round_Host.Val:
switch (ElementSize) {
case 2:
frinti(Dst.V8H(), Vector.V8H());
break;
case 4:
frinti(Dst.V4S(), Vector.V4S());
break;
case 8:
frinti(Dst.V2D(), Vector.V2D());
break;
}
break;
}
break;
}
}
}
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -18,7 +18,7 @@ DEF_OP(AESImc) {
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
auto Op = IROp->C<IR::IROp_VAESEnc>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
@@ -27,7 +27,7 @@ DEF_OP(AESEnc) {
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
auto Op = IROp->C<IR::IROp_VAESEncLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
@@ -35,7 +35,7 @@ DEF_OP(AESEncLast) {
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
auto Op = IROp->C<IR::IROp_VAESDec>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
@@ -44,7 +44,7 @@ DEF_OP(AESDec) {
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
auto Op = IROp->C<IR::IROp_VAESDecLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
@@ -52,7 +52,7 @@ DEF_OP(AESDecLast) {
}
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
@@ -69,9 +69,8 @@ DEF_OP(AESKeyGenAssist) {
if (Op->RCON) {
tbl(VTMP1.V16B(), VTMP1.V16B(), VTMP3.V16B());
LoadConstant(TMP1.W(), Op->RCON);
ins(VTMP2.V4S(), 1, TMP1.W());
ins(VTMP2.V4S(), 3, TMP1.W());
LoadConstant(TMP1, static_cast<uint64_t>(Op->RCON) << 32);
dup(VTMP2.V2D(), TMP1);
eor(GetDst(Node).V16B(), VTMP1.V16B(), VTMP2.V16B());
}
else {
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()), Op->Flag, 1);
+101 -28
View File
@@ -27,6 +27,8 @@ $end_info$
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/Profiler.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
@@ -78,7 +80,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -93,7 +95,12 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
uxth(w0, GetReg<RA_32>(IROp->Args[0].ID()));
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<void, uint16_t>(x1);
#else
blr(x1);
#endif
PopDynamicRegsAndLR();
@@ -108,7 +115,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
fmov(v0.S(), GetSrc(IROp->Args[0].ID()).S()) ;
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, float>(x0);
#else
blr(x0);
#endif
PopDynamicRegsAndLR();
@@ -127,7 +138,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, double>(x0);
#else
blr(x0);
#endif
PopDynamicRegsAndLR();
@@ -152,7 +167,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
mov(w0, GetReg<RA_32>(IROp->Args[0].ID()));
}
ldr(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, uint32_t>(x1);
#else
blr(x1);
#endif
PopDynamicRegsAndLR();
@@ -173,7 +192,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<float, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -192,7 +215,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<double, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -209,7 +236,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<double, double>(x0);
#else
blr(x0);
#endif
PopDynamicRegsAndLR();
@@ -228,7 +259,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
mov(v0.D(), GetSrc(IROp->Args[0].ID()).D());
mov(v1.D(), GetSrc(IROp->Args[1].ID()).D());
ldr(x0, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<double, double, double>(x0);
#else
blr(x0);
#endif
PopDynamicRegsAndLR();
@@ -248,7 +283,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -266,7 +305,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -284,7 +327,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -305,8 +352,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t, uint64_t>(x4);
#else
blr(x4);
#endif
PopDynamicRegsAndLR();
FillStaticRegs();
@@ -323,7 +373,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w1, GetSrc(IROp->Args[0].ID()).V8H(), 4);
ldr(x2, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, uint64_t, uint64_t>(x2);
#else
blr(x2);
#endif
PopDynamicRegsAndLR();
@@ -346,7 +400,11 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
umov(w3, GetSrc(IROp->Args[1].ID()).V8H(), 4);
ldr(x4, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex])));
#ifdef VIXL_SIMULATOR
GenerateIndirectRuntimeCall<__uint128_t, uint64_t, uint64_t, uint64_t, uint64_t>(x4);
#else
blr(x4);
#endif
PopDynamicRegsAndLR();
@@ -361,7 +419,8 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}", FEXCore::IR::GetName(IROp->Op), Info.ABI);
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}",
FEXCore::IR::GetName(IROp->Op), ToUnderlying(Info.ABI));
#endif
break;
}
@@ -395,7 +454,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
vixl::aarch64::CPU::EnsureIAndDCacheCoherency((void*)branch, 24);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
vixl::aarch64::Assembler emit((uint8_t*)(branch), 24);
vixl::CodeBufferCheckScope scope(&emit, 24, vixl::CodeBufferCheckScope::kDontReserveBufferSpace, vixl::CodeBufferCheckScope::kNoAssert);
Literal l_BranchHost{LinkerAddress};
@@ -410,7 +469,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
record[0] = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
}
@@ -418,12 +477,13 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
return HostCode;
}
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
void Arm64JITCore::Op_NoOp(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx, 0)
, HostSupportsSVE{ctx->HostFeatures.SupportsAVX}
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
@@ -470,10 +530,10 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
// Common
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::ThreadRemoveCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -483,7 +543,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Arm64JITCore_ExitFunctionLink);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::Context::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
@@ -491,7 +551,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::Intern
// Platform Specific
auto &AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
AArch64.LUREM = reinterpret_cast<uint64_t>(LUREM);
@@ -508,6 +568,7 @@ void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
return Thread->CTX->Dispatcher->HandleSIGILL(Thread, Signal, info, ucontext);
}, true);
#ifdef _M_ARM_64
CTX->SignalDelegation->RegisterHostSignalHandler(SIGBUS, [](FEXCore::Core::InternalThreadState *Thread, int Signal, void *info, void *ucontext) -> bool {
if (!Thread->CPUBackend->IsAddressInCodeBuffer(ArchHelpers::Context::GetPc(ucontext))) {
// Wasn't a sigbus in JIT code
@@ -516,6 +577,7 @@ void Arm64JITCore::InitializeSignalHandlers(FEXCore::Context::Context *CTX) {
return FEXCore::ArchHelpers::Arm64::HandleSIGBUS(Thread->CTX->Config.ParanoidTSO(), Signal, info, ucontext);
}, true);
#endif
}
void Arm64JITCore::EmitDetectionString() {
@@ -527,7 +589,7 @@ void Arm64JITCore::EmitDetectionString() {
void Arm64JITCore::ClearCache() {
// Get the backing code buffer
auto CodeBuffer = GetEmptyCodeBuffer();
*GetBuffer() = vixl::CodeBuffer(CodeBuffer->Ptr, CodeBuffer->Size);
EmitDetectionString();
@@ -549,12 +611,12 @@ template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_32>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::GPRFixedClass.Val) {
return SRA64[Reg.Reg].W();
} else if (Reg.Class == IR::GPRClass.Val) {
return RA64[Reg.Reg].W();
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected Class: {}", Reg.Class);
}
FEX_UNREACHABLE;
@@ -564,12 +626,12 @@ template<>
aarch64::Register Arm64JITCore::GetReg<Arm64JITCore::RA_64>(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::GPRFixedClass.Val) {
return SRA64[Reg.Reg];
} else if (Reg.Class == IR::GPRClass.Val) {
return RA64[Reg.Reg];
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected Class: {}", Reg.Class);
}
FEX_UNREACHABLE;
@@ -590,12 +652,12 @@ std::pair<aarch64::Register, aarch64::Register> Arm64JITCore::GetSrcPair<Arm64JI
aarch64::VRegister Arm64JITCore::GetSrc(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::FPRFixedClass.Val) {
return SRAFPR[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
return RAFPR[Reg.Reg];
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected Class: {}", Reg.Class);
}
FEX_UNREACHABLE;
@@ -604,12 +666,12 @@ aarch64::VRegister Arm64JITCore::GetSrc(IR::NodeID Node) const {
aarch64::VRegister Arm64JITCore::GetDst(IR::NodeID Node) const {
auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
if (Reg.Class == IR::FPRFixedClass.Val) {
return SRAFPR[Reg.Reg];
} else if (Reg.Class == IR::FPRClass.Val) {
return RAFPR[Reg.Reg];
} else {
LOGMAN_THROW_A_FMT(false, "Unexpected Class: {}", Reg.Class);
}
FEX_UNREACHABLE;
@@ -665,7 +727,13 @@ bool Arm64JITCore::IsGPR(IR::NodeID Node) const {
return Class == IR::GPRClass || Class == IR::GPRFixedClass;
}
void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
void *Arm64JITCore::CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData,
bool GDBEnabled) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
using namespace aarch64;
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
@@ -718,10 +786,12 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
SpillSlots = RAData->SpillSlots();
if (SpillSlots) {
if (IsImmAddSub(SpillSlots * 16)) {
sub(sp, sp, SpillSlots * 16);
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (IsImmAddSub(TotalSpillSlotsSize)) {
sub(sp, sp, TotalSpillSlotsSize);
} else {
LoadConstant(x0, SpillSlots * 16);
LoadConstant(x0, TotalSpillSlotsSize);
sub(sp, sp, x0);
}
}
@@ -732,7 +802,7 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
using namespace FEXCore::IR;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = GetCursorAddress<uint8_t *>();
@@ -789,14 +859,17 @@ void *Arm64JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IR
}
void Arm64JITCore::ResetStack() {
if (SpillSlots == 0)
if (SpillSlots == 0) {
return;
}
if (IsImmAddSub(SpillSlots * 16)) {
add(sp, sp, SpillSlots * 16);
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (IsImmAddSub(TotalSpillSlotsSize)) {
add(sp, sp, TotalSpillSlotsSize);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(x0, SpillSlots * 16);
LoadConstant(x0, TotalSpillSlotsSize);
add(sp, sp, x0);
}
}
+15 -23
View File
@@ -23,16 +23,6 @@ $end_info$
#include <utility>
#include <vector>
#define STATE x28
#define TMP1 x0
#define TMP2 x1
#define TMP3 x2
#define TMP4 x3
#define VTMP1 v1
#define VTMP2 v2
#define VTMP3 v3
namespace FEXCore::Core {
struct InternalThreadState;
}
@@ -66,6 +56,7 @@ public:
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
const bool HostSupportsSVE{};
Label *PendingTargetLabel;
FEXCore::Context::Context *CTX;
@@ -127,6 +118,17 @@ private:
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
// NOTE: Will use TMP1 as a way to encode immediates that happen to fall outside
// the limits of the scalar plus immediate variant of SVE load/stores.
//
// TMP1 is safe to use again once this memory operand is used with its
// equivalent loads or stores that this was called for.
[[nodiscard]] SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize,
aarch64::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
[[nodiscard]] bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
[[nodiscard]] bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
@@ -211,7 +213,7 @@ private:
*/
uint8_t *GuestEntry{};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -223,7 +225,7 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -309,7 +311,7 @@ private:
DEF_OP(InlineSyscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -343,8 +345,6 @@ private:
DEF_OP(StoreMemTSO);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -364,13 +364,10 @@ private:
///< Move ops
DEF_OP(ExtractElementPair);
DEF_OP(CreateElementPair);
DEF_OP(Mov);
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(SplatVector2);
DEF_OP(SplatVector4);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
@@ -429,18 +426,13 @@ private:
DEF_OP(VUShrS);
DEF_OP(VSShrS);
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
DEF_OP(VUShrNI);
DEF_OP(VUShrNI2);
DEF_OP(VBitcast);
DEF_OP(VSXTL);
DEF_OP(VSXTL2);
DEF_OP(VUXTL);
File diff suppressed because it is too large. Load diff
+40 -38
View File
@@ -11,7 +11,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
@@ -37,43 +37,41 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
hlt(4);
break;
case FEXCore::IR::Break_Overflow: // overflow
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)));
br(TMP1);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)));
add(sp, TMP1, 0);
// Now we need to jump to the thread stop handler
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_Interrupt3: { // INT3
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)));
br(TMP1);
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
ResetStack();
// First we must reset the stack
ResetStack();
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)));
br(TMP1);
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Op->Reason.Signal,
.TrapNo = Op->Reason.TrapNumber,
.si_code = Op->Reason.si_code,
.err_code = Op->Reason.ErrorRegister,
};
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
uint64_t Constant{};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(x1, Constant);
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData)));
switch (Op->Reason.Signal) {
case SIGILL:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL)));
br(TMP1);
break;
case SIGTRAP:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)));
br(TMP1);
break;
case SIGSEGV:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV)));
br(TMP1);
break;
default:
ldr(TMP1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)));
br(TMP1);
break;
}
}
@@ -124,8 +122,11 @@ DEF_OP(SetRoundingMode) {
mrs(TMP1, FPCR);
// vixl simulator doesn't support anything beyond ties-to-even rounding
#ifndef VIXL_SIMULATOR
// Insert the rounding flags
bfi(TMP1, TMP2, 22, 2);
#endif
// Insert the FTZ flag
lsr(TMP2, Src, 2);
@@ -139,6 +140,7 @@ DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushDynamicRegsAndLR();
SpillStaticRegs();
if (IsGPR(Op->Value.ID())) {
mov(x0, GetReg<RA_64>(Op->Value.ID()));
@@ -150,10 +152,10 @@ DEF_OP(Print) {
fmov(x1, GetSrc(Op->Value.ID()).V1D(), 1);
ldr(x3, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)));
}
SpillStaticRegs();
blr(x3);
FillStaticRegs();
blr(x3);
FillStaticRegs();
PopDynamicRegsAndLR();
}
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
@@ -68,17 +68,11 @@ DEF_OP(CreateElementPair) {
}
}
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
mov(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()));
}
#undef DEF_OP
void Arm64JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &Arm64JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
#undef REGISTER_OP
}
}
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -33,7 +33,7 @@ namespace FEXCore::CPU {
DEF_OP(SignalReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
add(rsp, SpillSlots * MaxSpillSlotSize); // + 8 to consume return address
}
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.SignalReturnHandler)]);
@@ -42,7 +42,7 @@ DEF_OP(SignalReturn) {
DEF_OP(CallbackReturn) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * 16); // + 8 to consume return address
add(rsp, SpillSlots * MaxSpillSlotSize); // + 8 to consume return address
}
// Make sure to adjust the refcounter so we don't clear the cache now
@@ -71,7 +71,7 @@ DEF_OP(ExitFunction) {
if (SpillSlots) {
add(rsp, SpillSlots * 16);
add(rsp, SpillSlots * MaxSpillSlotSize);
}
uint64_t NewRIP;
@@ -117,9 +117,9 @@ DEF_OP(ExitFunction) {
DEF_OP(Jump) {
const auto Op = IROp->C<IR::IROp_Jump>();
const auto ArgID = Op->Args(0).ID();
const auto Target = Op->TargetBlock.ID();
PendingTargetLabel = &JumpTargets.try_emplace(ArgID).first->second;
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
#define GRCMP(Node) (Op->CompareSize == 4 ? GetSrc<RA_32>(Node) : GetSrc<RA_64>(Node))
@@ -209,7 +209,7 @@ DEF_OP(Thunk) {
if (NumPush & 1)
sub(rsp, 8); // Align
mov(rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rdi, GetSrc<RA_64>(Op->ArgPtr.ID()));
auto thunkFn = ThreadState->CTX->ThunkHandler->LookupThunk(Op->ThunkNameHash);
@@ -253,7 +253,7 @@ DEF_OP(ValidateCode) {
}
}
DEF_OP(RemoveThreadCodeEntry) {
DEF_OP(ThreadRemoveCodeEntry) {
auto NumPush = RA64.size();
for (auto &Reg : RA64)
@@ -266,7 +266,7 @@ DEF_OP(RemoveThreadCodeEntry) {
mov(rax, Entry); // imm64 move
mov(rsi, rax);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.RemoveThreadCodeEntryFromJIT)]);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT)]);
if (NumPush & 1)
add(rsp, 8); // Align
@@ -288,8 +288,8 @@ DEF_OP(CPUID) {
// Result: RAX, RDX. 4xi32
// rsi can be in the source registers, so copy argument to edx first
mov (edx, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov (esi, GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov (edx, GetSrc<RA_32>(Op->Leaf.ID()));
mov (esi, GetSrc<RA_32>(Op->Function.ID()));
mov (rdi, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj)]);
auto NumPush = RA64.size();
@@ -322,7 +322,7 @@ void X86JITCore::RegisterBranchHandlers() {
REGISTER_OP(SYSCALL, Syscall);
REGISTER_OP(THUNK, Thunk);
REGISTER_OP(VALIDATECODE, ValidateCode);
REGISTER_OP(REMOVETHREADCODEENTRY, RemoveThreadCodeEntry);
REGISTER_OP(THREADREMOVECODEENTRY, ThreadRemoveCodeEntry);
REGISTER_OP(CPUID, CPUID);
#undef REGISTER_OP
}
@@ -17,27 +17,76 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
auto Op = IROp->C<IR::IROp_VInsGPR>();
movapd(GetDst(Node), GetSrc(Op->DestVector.ID()));
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
switch (Op->Header.ElementSize) {
case 1: {
pinsrb(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
const auto Dst = GetDst(Node);
const auto DestVector = GetSrc(Op->DestVector.ID());
const auto DestIdx = Op->DestIdx;
const auto ElementSize = Op->Header.ElementSize;
const auto ElementSizeBits = ElementSize * 8;
const auto Offset = ElementSizeBits * DestIdx;
constexpr auto SSEBitSize = Core::CPUState::XMM_SSE_REG_SIZE * 8;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto InUpperLane = Offset >= SSEBitSize;
if (InUpperLane && !Is256Bit) {
LOGMAN_MSG_A_FMT("Attempt to access upper 128-bit lane in 128-bit operation! Offset={}",
Offset);
return;
}
if (Is256Bit) {
vmovapd(ToYMM(Dst), ToYMM(DestVector));
} else {
vmovapd(Dst, DestVector);
}
const auto Insert = [&](const Xbyak::Xmm& reg, int index) {
switch (ElementSize) {
case 1: {
if (InUpperLane) {
index -= 16;
}
pinsrb(reg, GetSrc<RA_32>(Op->Src.ID()), index);
break;
}
case 2: {
if (InUpperLane) {
index -= 8;
}
pinsrw(reg, GetSrc<RA_32>(Op->Src.ID()), index);
break;
}
case 4: {
if (InUpperLane) {
index -= 4;
}
pinsrd(reg, GetSrc<RA_32>(Op->Src.ID()), index);
break;
}
case 8: {
if (InUpperLane) {
index -= 2;
}
pinsrq(reg, GetSrc<RA_64>(Op->Src.ID()), index);
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
case 2: {
pinsrw(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 4: {
pinsrd(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()), Op->DestIdx);
break;
}
case 8: {
pinsrq(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()), Op->DestIdx);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
};
if (InUpperLane) {
vextracti128(xmm15, ToYMM(Dst), 1);
Insert(xmm15, DestIdx);
vinserti128(ToYMM(Dst), ToYMM(Dst), xmm15, 1);
} else {
Insert(Dst, DestIdx);
}
}
@@ -45,56 +94,64 @@ DEF_OP(VCastFromGPR) {
auto Op = IROp->C<IR::IROp_VCastFromGPR>();
switch (Op->Header.ElementSize) {
case 1:
movzx(rax, GetSrc<RA_8>(Op->Header.Args[0].ID()));
movzx(rax, GetSrc<RA_8>(Op->Src.ID()));
vmovq(GetDst(Node), rax);
break;
case 2:
movzx(rax, GetSrc<RA_16>(Op->Header.Args[0].ID()));
movzx(rax, GetSrc<RA_16>(Op->Src.ID()));
vmovq(GetDst(Node), rax);
break;
case 4:
vmovd(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()).cvt32());
vmovd(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()).cvt32());
break;
case 8:
vmovq(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()).cvt64());
vmovq(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()).cvt64());
break;
default: LOGMAN_MSG_A_FMT("Unknown VCastFromGPR element size: {}", Op->Header.ElementSize);
}
}
DEF_OP(Float_FromGPR_S) {
auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const auto Op = IROp->C<IR::IROp_Float_FromGPR_S>();
const uint16_t ElementSize = Op->Header.ElementSize;
const uint16_t Conv = (ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0404: { // Float <- int32_t
cvtsi2ss(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
cvtsi2ss(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()));
break;
}
case 0x0408: { // Float <- int64_t
cvtsi2ss(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
cvtsi2ss(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()));
break;
}
case 0x0804: { // Double <- int32_t
cvtsi2sd(GetDst(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
cvtsi2sd(GetDst(Node), GetSrc<RA_32>(Op->Src.ID()));
break;
}
case 0x0808: { // Double <- int64_t
cvtsi2sd(GetDst(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
cvtsi2sd(GetDst(Node), GetSrc<RA_64>(Op->Src.ID()));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled conversion mask: Mask=0x{:04x}, ElementSize={}, SrcElementSize={}",
Conv, ElementSize, Op->SrcElementSize);
break;
}
}
DEF_OP(Float_FToF) {
auto Op = IROp->C<IR::IROp_Float_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
switch (Conv) {
case 0x0804: { // Double <- Float
cvtss2sd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
cvtss2sd(GetDst(Node), GetSrc(Op->Scalar.ID()));
break;
}
case 0x0408: { // Float <- Double
cvtsd2ss(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
cvtsd2ss(GetDst(Node), GetSrc(Op->Scalar.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Float_FToF sizes: 0x{:x}", Conv);
@@ -102,99 +159,194 @@ DEF_OP(Float_FToF) {
}
DEF_OP(Vector_SToF) {
auto Op = IROp->C<IR::IROp_Vector_SToF>();
switch (Op->Header.ElementSize) {
const auto Op = IROp->C<IR::IROp_Vector_SToF>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
cvtdq2ps(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
if (Is256Bit) {
vcvtdq2ps(ToYMM(Dst), ToYMM(Vector));
} else {
vcvtdq2ps(Dst, Vector);
}
break;
case 8:
// This operation is a bit disgusting in x86
// There is no vector form of this instruction until AVX512VL + AVX512DQ (vcvtqq2pd)
// 1) First extract the top 64bits
// 2) Do a scalar conversion on each
// 3) Make sure to merge them together at the end
pextrq(rax, GetSrc(Op->Header.Args[0].ID()), 1);
pextrq(rcx, GetSrc(Op->Header.Args[0].ID()), 0);
cvtsi2sd(GetDst(Node), rcx);
pextrq(rax, Vector, 1);
pextrq(rcx, Vector, 0);
cvtsi2sd(Dst, rcx);
cvtsi2sd(xmm15, rax);
movlhps(GetDst(Node), xmm15);
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", Op->Header.ElementSize);
vmovlhps(Dst, Dst, xmm15);
if (Is256Bit) {
vextracti128(xmm15, ToYMM(Vector), 1);
pextrq(rax, xmm15, 1);
pextrq(rcx, xmm15, 0);
cvtsi2sd(xmm15, rcx);
cvtsi2sd(xmm14, rax);
movlhps(xmm15, xmm14);
vinserti128(ToYMM(Dst), ToYMM(Dst), xmm15, 1);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_SToF element size: {}", ElementSize);
break;
}
}
DEF_OP(Vector_FToZS) {
auto Op = IROp->C<IR::IROp_Vector_FToZS>();
switch (Op->Header.ElementSize) {
const auto Op = IROp->C<IR::IROp_Vector_FToZS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
cvttps2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
if (Is256Bit) {
vcvttps2dq(ToYMM(Dst), ToYMM(Vector));
} else {
vcvttps2dq(Dst, Vector);
}
break;
case 8:
cvttpd2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", Op->Header.ElementSize);
if (Is256Bit) {
vcvttpd2dq(ToYMM(Dst), ToYMM(Vector));
} else {
vcvttpd2dq(Dst, Vector);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToZS element size: {}", ElementSize);
break;
}
}
DEF_OP(Vector_FToS) {
auto Op = IROp->C<IR::IROp_Vector_FToS>();
switch (Op->Header.ElementSize) {
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
cvtps2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
if (Is256Bit) {
vcvtps2dq(ToYMM(Dst), ToYMM(Vector));
} else {
vcvtps2dq(Dst, Vector);
}
break;
case 8:
cvtpd2dq(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", Op->Header.ElementSize);
if (Is256Bit) {
vcvtpd2dq(ToYMM(Dst), ToYMM(Vector));
} else {
vcvtpd2dq(Dst, Vector);
}
break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToS element size: {}", ElementSize);
break;
}
}
DEF_OP(Vector_FToF) {
auto Op = IROp->C<IR::IROp_Vector_FToF>();
uint16_t Conv = (Op->Header.ElementSize << 8) | Op->SrcElementSize;
const auto Op = IROp->C<IR::IROp_Vector_FToF>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
switch (Conv) {
case 0x0804: { // Double <- Float
cvtps2pd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
if (Is256Bit) {
vcvtps2pd(ToYMM(Dst), Vector);
} else {
vcvtps2pd(Dst, Vector);
}
break;
}
case 0x0408: { // Float <- Double
cvtpd2ps(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
if (Is256Bit) {
vcvtpd2ps(Dst, ToYMM(Vector));
} else {
vcvtpd2ps(Dst, Vector);
}
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF conversion type : 0x{:04x}", Conv); break;
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToF conversion type : 0x{:04x}", Conv);
break;
}
}
DEF_OP(Vector_FToI) {
auto Op = IROp->C<IR::IROp_Vector_FToI>();
uint8_t RoundMode{};
const auto Op = IROp->C<IR::IROp_Vector_FToI>();
const auto OpSize = IROp->Size;
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
RoundMode = 0b0000'0'0'00;
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
RoundMode = 0b0000'0'0'01;
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
RoundMode = 0b0000'0'0'10;
break;
case FEXCore::IR::Round_Towards_Zero.Val:
RoundMode = 0b0000'0'0'11;
break;
case FEXCore::IR::Round_Host.Val:
RoundMode = 0b0000'0'1'00;
break;
}
const uint8_t RoundMode = [Op] {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
return 0b0000'0'0'00;
case FEXCore::IR::Round_Negative_Infinity.Val:
return 0b0000'0'0'01;
case FEXCore::IR::Round_Positive_Infinity.Val:
return 0b0000'0'0'10;
case FEXCore::IR::Round_Towards_Zero.Val:
return 0b0000'0'0'11;
case FEXCore::IR::Round_Host.Val:
return 0b0000'0'1'00;
default:
LOGMAN_MSG_A_FMT("Unhandled rounding mode");
return 0;
}
}();
switch (Op->Header.ElementSize) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
switch (ElementSize) {
case 4:
roundps(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), RoundMode);
break;
if (Is256Bit) {
vroundps(ToYMM(Dst), ToYMM(Vector), RoundMode);
} else {
vroundps(Dst, Vector, RoundMode);
}
break;
case 8:
roundpd(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), RoundMode);
break;
if (Is256Bit) {
vroundpd(ToYMM(Dst), ToYMM(Vector), RoundMode);
} else {
vroundpd(Dst, Vector, RoundMode);
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled element size: {}", ElementSize);
break;
}
}
@@ -17,44 +17,44 @@ namespace FEXCore::CPU {
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
vaesimc(GetDst(Node), GetSrc(Op->Header.Args[0].ID()));
vaesimc(GetDst(Node), GetSrc(Op->Vector.ID()));
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
vaesenc(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
vaesenc(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
vaesenclast(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
vaesenclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
vaesdec(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
vaesdec(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
vaesdeclast(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), GetSrc(Op->Header.Args[1].ID()));
vaesdeclast(GetDst(Node), GetSrc(Op->State.ID()), GetSrc(Op->Key.ID()));
}
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
vaeskeygenassist(GetDst(Node), GetSrc(Op->Header.Args[0].ID()), Op->RCON);
vaeskeygenassist(GetDst(Node), GetSrc(Op->Src.ID()), Op->RCON);
}
DEF_OP(CRC32) {
auto Op = IROp->C<IR::IROp_CRC32>();
switch (IROp->Size) {
case 4:
mov(TMP1, GetSrc<RA_32>(Op->Header.Args[1].ID()));
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Header.Args[0].ID()));
mov(TMP1, GetSrc<RA_32>(Op->Src2.ID()));
mov(GetDst<RA_32>(Node), GetSrc<RA_32>(Op->Src1.ID()));
break;
case 8:
mov(TMP1, GetSrc<RA_64>(Op->Header.Args[1].ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(TMP1, GetSrc<RA_64>(Op->Src2.ID()));
mov(GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Src1.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", IROp->Size);
}
@@ -18,7 +18,7 @@ namespace FEXCore::CPU {
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
mov(rax, GetSrc<RA_64>(Op->Header.Args[0].ID()));
mov(rax, GetSrc<RA_64>(Op->Value.ID()));
shr(rax, Op->Flag);
and_(rax, 1);
mov(GetDst<RA_64>(Node), rax);
+39 -24
View File
@@ -25,7 +25,9 @@ $end_info$
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <algorithm>
#include <array>
@@ -59,32 +61,42 @@ static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
namespace FEXCore::CPU {
void X86JITCore::PushRegs() {
sub(rsp, 16 * RAXMM_x.size());
const auto AVXRegSize = Core::CPUState::XMM_AVX_REG_SIZE;
sub(rsp, AVXRegSize * RAXMM_x.size());
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
movaps(ptr[rsp + i * 16], RAXMM_x[i]);
vmovups(ptr[rsp + i * AVXRegSize], ToYMM(RAXMM_x[i]));
}
for (auto &Reg : RA64)
for (const auto &Reg : RA64) {
push(Reg);
}
auto NumPush = RA64.size();
if (NumPush & 1)
sub(rsp, 8); // Align
const auto NumPush = RA64.size();
if ((NumPush & 1) != 0) {
// Align
sub(rsp, 8);
}
}
void X86JITCore::PopRegs() {
auto NumPush = RA64.size();
const auto AVXRegSize = Core::CPUState::XMM_AVX_REG_SIZE;
const auto NumPush = RA64.size();
if (NumPush & 1)
add(rsp, 8); // Align
for (uint32_t i = RA64.size(); i > 0; --i)
pop(RA64[i - 1]);
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
movaps(RAXMM_x[i], ptr[rsp + i * 16]);
if ((NumPush & 1) != 0) {
// Align
add(rsp, 8);
}
add(rsp, 16 * RAXMM_x.size());
for (uint32_t i = RA64.size(); i > 0; --i) {
pop(RA64[i - 1]);
}
for (size_t i = 0; i < RAXMM_x.size(); ++i) {
vmovups(ToYMM(RAXMM_x[i]), ptr[rsp + i * AVXRegSize]);
}
add(rsp, AVXRegSize * RAXMM_x.size());
}
void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
@@ -293,7 +305,8 @@ void X86JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}", FEXCore::IR::GetName(IROp->Op), Info.ABI);
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}",
IR::GetName(IROp->Op), ToUnderlying(Info.ABI));
#endif
break;
}
@@ -312,7 +325,7 @@ static uint64_t X86JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame,
}
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
Context::Context::ThreadAddBlockLink(Thread, GuestRip, (uintptr_t)record, [record, LinkerAddress]{
// undo the link
record[0] = LinkerAddress;
});
@@ -358,10 +371,10 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
{
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
Common.RemoveThreadCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::RemoveThreadCodeEntryFromJit);
Common.ThreadRemoveCodeEntryFromJIT = reinterpret_cast<uintptr_t>(&Context::Context::ThreadRemoveCodeEntryFromJit);
Common.CPUIDObj = reinterpret_cast<uint64_t>(&CTX->CPUID);
{
@@ -371,7 +384,7 @@ X86JITCore::X86JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalTh
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&X86JITCore_ExitFunctionLink);
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::Context::ThreadExitFunctionLink<X86JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
@@ -407,7 +420,7 @@ void X86JITCore::ClearCache() {
IR::PhysicalRegister X86JITCore::GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
LOGMAN_THROW_AA_FMT(PhyReg.Raw != 255, "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
return PhyReg;
}
@@ -570,6 +583,8 @@ std::tuple<X86JITCore::SetCC, X86JITCore::CMovCC, X86JITCore::JCC> X86JITCore::G
}
void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRListView const *IR, [[maybe_unused]] FEXCore::Core::DebugData *DebugData, FEXCore::IR::RegisterAllocationData *RAData, bool GDBEnabled) {
FEXCORE_PROFILE_SCOPED("x86::CompileCode");
JumpTargets.clear();
uint32_t SSACount = IR->GetSSACount();
@@ -592,12 +607,12 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
setSize(getSize() + GDBSize);
}
LOGMAN_THROW_A_FMT(RAData != nullptr, "Needs RA");
LOGMAN_THROW_AA_FMT(RAData != nullptr, "Needs RA");
SpillSlots = RAData->SpillSlots();
if (SpillSlots) {
sub(rsp, SpillSlots * 16);
sub(rsp, SpillSlots * MaxSpillSlotSize);
}
#ifdef BLOCKSTATS
@@ -649,7 +664,7 @@ void *X86JITCore::CompileCode(uint64_t Entry, [[maybe_unused]] FEXCore::IR::IRLi
using namespace FEXCore::IR;
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
auto BlockIROp = BlockHeader->CW<IROp_CodeBlock>();
LOGMAN_THROW_A_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = getCurr<uint8_t *>();
@@ -181,6 +181,10 @@ private:
[[nodiscard]] Xbyak::Xmm GetSrc(IR::NodeID Node) const;
[[nodiscard]] Xbyak::Xmm GetDst(IR::NodeID Node) const;
[[nodiscard]] static Xbyak::Ymm ToYMM(const Xbyak::Xmm& xmm) {
return Xbyak::Ymm{xmm.getIdx()};
}
[[nodiscard]] Xbyak::RegExp GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale) const;
@@ -205,7 +209,7 @@ private:
* @brief Current guest RIP entrypoint
*/
uint8_t *GuestEntry{};
using SetCC = void (X86JITCore::*)(const Operand& op);
using CMovCC = void (X86JITCore::*)(const Reg& reg, const Operand& op);
using JCC = void (X86JITCore::*)(const Label& label, LabelType type);
@@ -312,7 +316,7 @@ private:
DEF_OP(Syscall);
DEF_OP(Thunk);
DEF_OP(ValidateCode);
DEF_OP(RemoveThreadCodeEntry);
DEF_OP(ThreadRemoveCodeEntry);
DEF_OP(CPUID);
///< Conversion ops
@@ -333,6 +337,8 @@ private:
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
DEF_OP(LoadContextIndexed);
DEF_OP(StoreContextIndexed);
DEF_OP(SpillRegister);
@@ -341,8 +347,6 @@ private:
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -362,12 +366,10 @@ private:
///< Move ops
DEF_OP(ExtractElementPair);
DEF_OP(CreateElementPair);
DEF_OP(Mov);
///< Vector ops
DEF_OP(VectorZero);
DEF_OP(VectorImm);
DEF_OP(SplatVector);
DEF_OP(VMov);
DEF_OP(VAnd);
DEF_OP(VBic);
@@ -426,18 +428,13 @@ private:
DEF_OP(VUShrS);
DEF_OP(VSShrS);
DEF_OP(VInsElement);
DEF_OP(VInsScalarElement);
DEF_OP(VExtractElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
DEF_OP(VUShrNI);
DEF_OP(VUShrNI2);
DEF_OP(VBitcast);
DEF_OP(VSXTL);
DEF_OP(VSXTL2);
DEF_OP(VUXTL);
+365 -189
View File
@@ -21,130 +21,266 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void X86JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
auto Op = IROp->C<IR::IROp_LoadContext>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_LoadContext>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
movzx(GetDst<RA_32>(Node), byte [STATE + Op->Offset]);
break;
}
break;
case 2: {
movzx(GetDst<RA_32>(Node), word [STATE + Op->Offset]);
break;
}
break;
case 4: {
mov(GetDst<RA_32>(Node), dword [STATE + Op->Offset]);
break;
}
break;
case 8: {
mov(GetDst<RA_64>(Node), qword [STATE + Op->Offset]);
break;
}
break;
case 16: {
LOGMAN_MSG_A_FMT("Invalid GPR load of size 16");
break;
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
break;
}
}
else {
const auto Dst = GetDst(Node);
switch (OpSize) {
case 1: {
movzx(rax, byte [STATE + Op->Offset]);
vmovq(GetDst(Node), rax);
vmovq(Dst, rax);
break;
}
break;
case 2: {
movzx(rax, word [STATE + Op->Offset]);
vmovq(GetDst(Node), rax);
vmovq(Dst, rax);
break;
}
break;
case 4: {
vmovd(GetDst(Node), dword [STATE + Op->Offset]);
vmovd(Dst, dword [STATE + Op->Offset]);
break;
}
break;
case 8: {
vmovq(GetDst(Node), qword [STATE + Op->Offset]);
vmovq(Dst, qword [STATE + Op->Offset]);
break;
}
break;
case 16: {
if (Op->Offset % 16 == 0)
movaps(GetDst(Node), xword [STATE + Op->Offset]);
else
movups(GetDst(Node), xword [STATE + Op->Offset]);
if (Op->Offset % 16 == 0) {
vmovaps(Dst, xword [STATE + Op->Offset]);
} else {
vmovups(Dst, xword [STATE + Op->Offset]);
}
break;
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
case 32: {
vmovups(ToYMM(Dst), yword [STATE + Op->Offset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
break;
}
}
}
DEF_OP(StoreContext) {
auto Op = IROp->C<IR::IROp_StoreContext>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_StoreContext>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
mov(byte [STATE + Op->Offset], GetSrc<RA_8>(Op->Value.ID()));
break;
}
break;
case 2: {
mov(word [STATE + Op->Offset], GetSrc<RA_16>(Op->Value.ID()));
break;
}
break;
case 4: {
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Value.ID()));
break;
}
break;
case 8: {
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Value.ID()));
break;
}
break;
case 16:
LogMan::Msg::DFmt("Invalid store size of 16");
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
case 16: {
LOGMAN_MSG_A_FMT("Invalid store size of 16");
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
break;
}
}
else {
const auto Value = GetSrc(Op->Value.ID());
switch (OpSize) {
case 1: {
pextrb(byte [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
pextrb(byte [STATE + Op->Offset], Value, 0);
break;
}
break;
case 2: {
pextrw(word [STATE + Op->Offset], GetSrc(Op->Value.ID()), 0);
pextrw(word [STATE + Op->Offset], Value, 0);
break;
}
break;
case 4: {
vmovd(dword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
vmovd(dword [STATE + Op->Offset], Value);
break;
}
break;
case 8: {
vmovq(qword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
vmovq(qword [STATE + Op->Offset], Value);
break;
}
break;
case 16: {
if (Op->Offset % 16 == 0)
movaps(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
else
movups(xword [STATE + Op->Offset], GetSrc(Op->Value.ID()));
if (Op->Offset % 16 == 0) {
vmovaps(xword [STATE + Op->Offset], Value);
} else {
vmovups(xword [STATE + Op->Offset], Value);
}
break;
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
case 32: {
vmovups(yword [STATE + Op->Offset], ToYMM(Value));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
break;
}
}
}
DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
movzx(GetSrc<RA_32>(Node), byte [STATE + Op->Offset]);
break;
}
case 2: {
movzx(GetSrc<RA_32>(Node), word [STATE + Op->Offset]);
break;
}
case 4: {
mov(GetSrc<RA_32>(Node), dword [STATE + Op->Offset]);
break;
}
case 8: {
mov(GetSrc<RA_64>(Node), qword [STATE + Op->Offset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister size: {}", OpSize);
break;
}
}
else {
const auto Dst = GetSrc(Node);
switch (OpSize) {
case 1: {
movzx(rax, byte [STATE + Op->Offset]);
vmovq(Dst, rax);
break;
}
case 2: {
movzx(rax, word [STATE + Op->Offset]);
vmovq(Dst, rax);
break;
}
case 4: {
vmovd(Dst, dword [STATE + Op->Offset]);
break;
}
case 8: {
vmovq(Dst, qword [STATE + Op->Offset]);
break;
}
case 16: {
if (Op->Offset % 16 == 0) {
vmovaps(Dst, xword [STATE + Op->Offset]);
} else {
vmovups(Dst, xword [STATE + Op->Offset]);
}
break;
}
case 32: {
vmovups(ToYMM(Dst), yword [STATE + Op->Offset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister size: {}", OpSize);
break;
}
}
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
const auto regOffs = Op->Offset & 7;
switch (OpSize) {
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Value.ID()));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto Value = GetSrc(Op->Value.ID());
switch (OpSize) {
case 16: {
if (Op->Offset % 16 == 0) {
vmovaps(xword [STATE + Op->Offset], Value);
} else {
vmovups(xword [STATE + Op->Offset], Value);
}
break;
}
case 32: {
vmovups(yword [STATE + Op->Offset], ToYMM(Value));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
break;
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(LoadContextIndexed) {
auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
size_t size = IROp->Size;
Reg index = GetSrc<RA_64>(Op->Index.ID());
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
const Reg Index = GetSrc<RA_64>(Op->Index.ID());
if (Op->Class == IR::GPRClass) {
switch (Op->Stride) {
@@ -153,21 +289,21 @@ DEF_OP(LoadContextIndexed) {
case 4:
case 8: {
lea(rax, dword [STATE + Op->BaseOffset]);
switch (size) {
switch (OpSize) {
case 1:
movzx(GetDst<RA_32>(Node), byte [rax + index * Op->Stride]);
movzx(GetDst<RA_32>(Node), byte [rax + Index * Op->Stride]);
break;
case 2:
movzx(GetDst<RA_32>(Node), word [rax + index * Op->Stride]);
movzx(GetDst<RA_32>(Node), word [rax + Index * Op->Stride]);
break;
case 4:
mov(GetDst<RA_32>(Node), dword [rax + index * Op->Stride]);
mov(GetDst<RA_32>(Node), dword [rax + Index * Op->Stride]);
break;
case 8:
mov(GetDst<RA_64>(Node), qword [rax + index * Op->Stride]);
mov(GetDst<RA_64>(Node), qword [rax + Index * Op->Stride]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", OpSize);
break;
}
break;
@@ -186,53 +322,63 @@ DEF_OP(LoadContextIndexed) {
case 2:
case 4:
case 8: {
const auto Dst = GetDst(Node);
lea(rax, dword [STATE + Op->BaseOffset]);
switch (size) {
switch (OpSize) {
case 1:
movzx(eax, byte [rax + index * Op->Stride]);
vmovd(GetDst(Node), eax);
movzx(eax, byte [rax + Index * Op->Stride]);
vmovd(Dst, eax);
break;
case 2:
movzx(eax, word [rax + index * Op->Stride]);
vmovd(GetDst(Node), eax);
movzx(eax, word [rax + Index * Op->Stride]);
vmovd(Dst, eax);
break;
case 4:
vmovd(GetDst(Node), dword [rax + index * Op->Stride]);
vmovd(Dst, dword [rax + Index * Op->Stride]);
break;
case 8:
vmovq(GetDst(Node), qword [rax + index * Op->Stride]);
vmovq(Dst, qword [rax + Index * Op->Stride]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", OpSize);
break;
}
break;
}
case 16: {
mov(rax, index);
shl(rax, 4);
case 16:
case 32: {
const auto Dst = GetDst(Node);
const auto Shift = Op->Stride == 16 ? 4 : 5;
mov(rax, Index);
shl(rax, Shift);
lea(rax, dword [rax + Op->BaseOffset]);
switch (size) {
switch (OpSize) {
case 1:
pinsrb(GetDst(Node), byte [STATE + rax], 0);
pinsrb(Dst, byte [STATE + rax], 0);
break;
case 2:
pinsrw(GetDst(Node), word [STATE + rax], 0);
pinsrw(Dst, word [STATE + rax], 0);
break;
case 4:
vmovd(GetDst(Node), dword [STATE + rax]);
vmovd(Dst, dword [STATE + rax]);
break;
case 8:
vmovq(GetDst(Node), qword [STATE + rax]);
vmovq(Dst, qword [STATE + rax]);
break;
case 16:
if (Op->BaseOffset % 16 == 0)
movaps(GetDst(Node), xword [STATE + rax]);
else
movups(GetDst(Node), xword [STATE + rax]);
if (Op->BaseOffset % 16 == 0) {
vmovaps(Dst, xword [STATE + rax]);
} else {
vmovups(Dst, xword [STATE + rax]);
}
break;
case 32:
vmovups(ToYMM(Dst), yword [STATE + rax]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", IROp->Size);
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", OpSize);
break;
}
break;
@@ -245,12 +391,13 @@ DEF_OP(LoadContextIndexed) {
}
DEF_OP(StoreContextIndexed) {
auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
Reg index = GetSrc<RA_64>(Op->Index.ID());
size_t size = IROp->Size;
const auto Op = IROp->C<IR::IROp_StoreContextIndexed>();
const auto OpSize = IROp->Size;
const Reg Index = GetSrc<RA_64>(Op->Index.ID());
if (Op->Class == IR::GPRClass) {
auto value = GetSrc<RA_64>(Op->Value.ID());
const auto Value = GetSrc<RA_64>(Op->Value.ID());
lea(rax, dword [STATE + Op->BaseOffset]);
switch (Op->Stride) {
@@ -258,10 +405,10 @@ DEF_OP(StoreContextIndexed) {
case 2:
case 4:
case 8: {
if (!(size == 1 || size == 2 || size == 4 || size == 8)) {
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", IROp->Size);
if (!(OpSize == 1 || OpSize == 2 || OpSize == 4 || OpSize == 8)) {
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", OpSize);
}
mov(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
mov(AddressFrame(OpSize * 8) [rax + Index * Op->Stride], Value);
break;
}
default:
@@ -270,57 +417,64 @@ DEF_OP(StoreContextIndexed) {
}
}
else {
auto value = GetSrc(Op->Value.ID());
const auto Value = GetSrc(Op->Value.ID());
switch (Op->Stride) {
case 1:
case 2:
case 4:
case 8: {
lea(rax, dword [STATE + Op->BaseOffset]);
switch (size) {
switch (OpSize) {
case 1:
pextrb(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value, 0);
pextrb(AddressFrame(OpSize * 8) [rax + Index * Op->Stride], Value, 0);
break;
case 2:
pextrw(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value, 0);
pextrw(AddressFrame(OpSize * 8) [rax + Index * Op->Stride], Value, 0);
break;
case 4:
vmovd(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
vmovd(AddressFrame(OpSize * 8) [rax + Index * Op->Stride], Value);
break;
case 8:
vmovq(AddressFrame(IROp->Size * 8) [rax + index * Op->Stride], value);
vmovq(AddressFrame(OpSize * 8) [rax + Index * Op->Stride], Value);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", size);
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", OpSize);
break;
}
break;
}
case 16: {
mov(rax, index);
shl(rax, 4);
case 16:
case 32: {
const auto Shift = Op->Stride == 16 ? 4 : 5;
mov(rax, Index);
shl(rax, Shift);
lea(rax, dword [rax + Op->BaseOffset]);
switch (size) {
switch (OpSize) {
case 1:
pextrb(AddressFrame(IROp->Size * 8) [STATE + rax], value, 0);
pextrb(AddressFrame(OpSize * 8) [STATE + rax], Value, 0);
break;
case 2:
pextrw(AddressFrame(IROp->Size * 8) [STATE + rax], value, 0);
pextrw(AddressFrame(OpSize * 8) [STATE + rax], Value, 0);
break;
case 4:
vmovd(AddressFrame(IROp->Size * 8) [STATE + rax], value);
vmovd(AddressFrame(OpSize * 8) [STATE + rax], Value);
break;
case 8:
vmovq(AddressFrame(IROp->Size * 8) [STATE + rax], value);
vmovq(AddressFrame(OpSize * 8) [STATE + rax], Value);
break;
case 16:
if (Op->BaseOffset % 16 == 0)
movaps(xword [STATE + rax], value);
else
movups(xword [STATE + rax], value);
if (Op->BaseOffset % 16 == 0) {
vmovaps(xword [STATE + rax], Value);
} else {
vmovups(xword [STATE + rax], Value);
}
break;
case 32:
vmovups(yword [STATE + rax], ToYMM(Value));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", size);
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", OpSize);
break;
}
break;
@@ -333,10 +487,10 @@ DEF_OP(StoreContextIndexed) {
}
DEF_OP(SpillRegister) {
auto Op = IROp->C<IR::IROp_SpillRegister>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_SpillRegister>();
const uint8_t OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
uint32_t SlotOffset = Op->Slot * 16;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
@@ -355,36 +509,44 @@ DEF_OP(SpillRegister) {
mov(qword [rsp + SlotOffset], GetSrc<RA_64>(Op->Value.ID()));
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
default:
LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
const auto Src = GetSrc(Op->Value.ID());
switch (OpSize) {
case 4: {
movss(dword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
movss(dword [rsp + SlotOffset], Src);
break;
}
case 8: {
movsd(qword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
movsd(qword [rsp + SlotOffset], Src);
break;
}
case 16: {
movaps(xword [rsp + SlotOffset], GetSrc(Op->Value.ID()));
movaps(xword [rsp + SlotOffset], Src);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
case 32: {
vmovups(yword [rsp + SlotOffset], ToYMM(Src));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled SpillRegister size: {}", OpSize);
break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled SpillRegister class: {}", Op->Class.Val);
}
}
DEF_OP(FillRegister) {
auto Op = IROp->C<IR::IROp_FillRegister>();
uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_FillRegister>();
const uint8_t OpSize = IROp->Size;
const uint32_t SlotOffset = Op->Slot * MaxSpillSlotSize;
uint32_t SlotOffset = Op->Slot * 16;
if (Op->Class == FEXCore::IR::GPRClass) {
switch (OpSize) {
case 1: {
@@ -403,23 +565,33 @@ DEF_OP(FillRegister) {
mov(GetDst<RA_64>(Node), qword [rsp + SlotOffset]);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize);
default:
LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize);
break;
}
} else if (Op->Class == FEXCore::IR::FPRClass) {
const auto Dst = GetDst(Node);
switch (OpSize) {
case 4: {
movss(GetDst(Node), dword [rsp + SlotOffset]);
vmovss(Dst, dword [rsp + SlotOffset]);
break;
}
case 8: {
movsd(GetDst(Node), qword [rsp + SlotOffset]);
vmovsd(Dst, qword [rsp + SlotOffset]);
break;
}
case 16: {
movaps(GetDst(Node), xword [rsp + SlotOffset]);
vmovaps(Dst, xword [rsp + SlotOffset]);
break;
}
default: LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize);
case 32: {
vmovups(ToYMM(Dst), yword [rsp + SlotOffset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled FillRegister size: {}", OpSize);
break;
}
} else {
LOGMAN_MSG_A_FMT("Unhandled FillRegister class: {}", Op->Class.Val);
@@ -464,130 +636,136 @@ Xbyak::RegExp X86JITCore::GenerateModRM(Xbyak::Reg Base, IR::OrderedNodeWrapper
}
DEF_OP(LoadMem) {
auto Op = IROp->C<IR::IROp_LoadMem>();
const auto Op = IROp->C<IR::IROp_LoadMem>();
const auto OpSize = IROp->Size;
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
const Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == IR::GPRClass) {
auto Dst = GetDst<RA_64>(Node);
const auto Dst = GetDst<RA_64>(Node);
switch (IROp->Size) {
switch (OpSize) {
case 1: {
movzx (Dst, byte [MemPtr]);
movzx(Dst, byte [MemPtr]);
break;
}
break;
case 2: {
movzx (Dst, word [MemPtr]);
movzx(Dst, word [MemPtr]);
break;
}
break;
case 4: {
mov(Dst.cvt32(), dword [MemPtr]);
break;
}
break;
case 8: {
mov(Dst, qword [MemPtr]);
break;
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
default:
LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", OpSize);
break;
}
}
else
{
auto Dst = GetDst(Node);
const auto Dst = GetDst(Node);
switch (IROp->Size) {
switch (OpSize) {
case 1: {
movzx(eax, byte [MemPtr]);
vmovd(Dst, eax);
break;
}
break;
case 2: {
movzx(eax, word [MemPtr]);
vmovd(Dst, eax);
break;
}
break;
case 4: {
vmovd(Dst, dword [MemPtr]);
break;
}
break;
case 8: {
vmovq(Dst, qword [MemPtr]);
break;
}
break;
case 16: {
if (IROp->Size == Op->Align)
movups(GetDst(Node), xword [MemPtr]);
else
movups(GetDst(Node), xword [MemPtr]);
if (MemoryDebug) {
movq(rcx, GetDst(Node));
}
}
break;
default: LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", IROp->Size);
vmovups(Dst, xword [MemPtr]);
if (MemoryDebug) {
movq(rcx, Dst);
}
break;
}
case 32: {
vmovups(ToYMM(Dst), yword [MemPtr]);
if (MemoryDebug) {
movq(rcx, Dst);
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadMem size: {}", OpSize);
break;
}
}
}
DEF_OP(StoreMem) {
auto Op = IROp->C<IR::IROp_StoreMem>();
const auto Op = IROp->C<IR::IROp_StoreMem>();
const auto OpSize = IROp->Size;
Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
const Xbyak::Reg MemReg = GetSrc<RA_64>(Op->Addr.ID());
const auto MemPtr = GenerateModRM(MemReg, Op->Offset, Op->OffsetType, Op->OffsetScale);
if (Op->Class == IR::GPRClass) {
switch (IROp->Size) {
switch (OpSize) {
case 1:
mov(byte [MemPtr], GetSrc<RA_8>(Op->Value.ID()));
break;
break;
case 2:
mov(word [MemPtr], GetSrc<RA_16>(Op->Value.ID()));
break;
break;
case 4:
mov(dword [MemPtr], GetSrc<RA_32>(Op->Value.ID()));
break;
break;
case 8:
mov(qword [MemPtr], GetSrc<RA_64>(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", OpSize);
break;
}
}
else {
switch (IROp->Size) {
const auto Value = GetSrc(Op->Value.ID());
switch (OpSize) {
case 1:
pextrb(byte [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
pextrb(byte [MemPtr], Value, 0);
break;
case 2:
pextrw(word [MemPtr], GetSrc(Op->Value.ID()), 0);
break;
pextrw(word [MemPtr], Value, 0);
break;
case 4:
vmovd(dword [MemPtr], GetSrc(Op->Value.ID()));
break;
vmovd(dword [MemPtr], Value);
break;
case 8:
vmovq(qword [MemPtr], GetSrc(Op->Value.ID()));
break;
vmovq(qword [MemPtr], Value);
break;
case 16:
if (IROp->Size == Op->Align)
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
else
movups(xword [MemPtr], GetSrc(Op->Value.ID()));
break;
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", IROp->Size);
vmovups(xword [MemPtr], Value);
break;
case 32:
vmovups(yword [MemPtr], ToYMM(Value));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreMem size: {}", OpSize);
break;
}
}
}
DEF_OP(VLoadMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(VStoreMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -618,8 +796,8 @@ void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, Unhandled); // SRA specific, not supported on this backend
REGISTER_OP(STOREREGISTER, Unhandled);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
@@ -630,8 +808,6 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
+33 -51
View File
@@ -45,56 +45,38 @@ DEF_OP(Fence) {
DEF_OP(Break) {
auto Op = IROp->C<IR::IROp_Break>();
switch (Op->Reason) {
case FEXCore::IR::Break_Unimplemented: // Hard fault
case FEXCore::IR::Break_Interrupt: // Guest ud2
ud2();
break;
case FEXCore::IR::Break_Overflow: // overflow
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.OverflowExceptionHandler)]);
break;
case FEXCore::IR::Break_Halt: { // HLT
// Time to quit
// Set our stack to the starting stack location
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
break;
}
case FEXCore::IR::Break_Interrupt3: // INT3
{
if (CTX->GetGdbServerStatus()) {
// Adjust the stack first for a regular return
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
if (SpillSlots) {
add(rsp, SpillSlots * MaxSpillSlotSize);
}
// This jump target needs to be a constant offset here
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadPauseHandlerSpillSRA)]);
}
else {
// If we don't have a gdb server attached then....crash?
// Treat this case like HLT
mov(rsp, qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, ReturningStackLocation)]);
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Op->Reason.Signal,
.TrapNo = Op->Reason.TrapNumber,
.si_code = Op->Reason.si_code,
.err_code = Op->Reason.ErrorRegister,
};
// Now we need to jump to the thread stop handler
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadStopHandlerSpillSRA)]);
}
uint64_t Constant{};
memcpy(&Constant, &State, sizeof(State));
mov(TMP1, Constant);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData)], TMP1);
switch (Op->Reason.Signal) {
case SIGILL:
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGILL)]);
break;
case SIGTRAP:
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)]);
break;
case SIGSEGV:
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGSEGV)]);
break;
default:
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.GuestSignal_SIGTRAP)]);
break;
}
case FEXCore::IR::Break_InvalidInstruction:
{
if (SpillSlots) {
add(rsp, SpillSlots * 16);
}
// Need to be outside of JIT cache space to ensure cache clearing correctness
jmp(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.UnimplementedInstructionHandler)]);
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Break reason: {}", Op->Reason);
}
}
@@ -110,7 +92,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetSrc<RA_32>(Op->Header.Args[0].ID());
auto Src = GetSrc<RA_32>(Op->RoundMode.ID());
// Load old mxcsr
// Only stores to memory
@@ -135,13 +117,13 @@ DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
PushRegs();
if (IsGPR(Op->Header.Args[0].ID())) {
mov (rdi, GetSrc<RA_64>(Op->Header.Args[0].ID()));
if (IsGPR(Op->Value.ID())) {
mov (rdi, GetSrc<RA_64>(Op->Value.ID()));
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue)]);
}
else {
pextrq(rdi, GetSrc(Op->Header.Args[0].ID()), 0);
pextrq(rsi, GetSrc(Op->Header.Args[0].ID()), 1);
pextrq(rdi, GetSrc(Op->Value.ID()), 0);
pextrq(rsi, GetSrc(Op->Value.ID()), 1);
call(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue)]);
}
@@ -20,13 +20,13 @@ DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
case 4: {
auto Src = GetSrcPair<RA_32>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_32>(Op->Pair.ID());
std::array<Xbyak::Reg, 2> Regs = {Src.first, Src.second};
mov (GetDst<RA_32>(Node), Regs[Op->Element]);
break;
}
case 8: {
auto Src = GetSrcPair<RA_64>(Op->Header.Args[0].ID());
auto Src = GetSrcPair<RA_64>(Op->Pair.ID());
std::array<Xbyak::Reg, 2> Regs = {Src.first, Src.second};
mov (GetDst<RA_64>(Node), Regs[Op->Element]);
break;
@@ -45,15 +45,15 @@ DEF_OP(CreateElementPair) {
switch (IROp->ElementSize) {
case 4: {
Dst = GetSrcPair<RA_32>(Node);
RegFirst = GetSrc<RA_32>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_32>(Op->Header.Args[1].ID());
RegFirst = GetSrc<RA_32>(Op->Lower.ID());
RegSecond = GetSrc<RA_32>(Op->Upper.ID());
RegTmp = eax;
break;
}
case 8: {
Dst = GetSrcPair<RA_64>(Node);
RegFirst = GetSrc<RA_64>(Op->Header.Args[0].ID());
RegSecond = GetSrc<RA_64>(Op->Header.Args[1].ID());
RegFirst = GetSrc<RA_64>(Op->Lower.ID());
RegSecond = GetSrc<RA_64>(Op->Upper.ID());
RegTmp = rax;
break;
}
@@ -73,17 +73,11 @@ DEF_OP(CreateElementPair) {
}
}
DEF_OP(Mov) {
auto Op = IROp->C<IR::IROp_Mov>();
mov (GetDst<RA_64>(Node), GetSrc<RA_64>(Op->Header.Args[0].ID()));
}
#undef DEF_OP
void X86JITCore::RegisterMoveHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(EXTRACTELEMENTPAIR, ExtractElementPair);
REGISTER_OP(CREATEELEMENTPAIR, CreateElementPair);
REGISTER_OP(MOV, Mov);
#undef REGISTER_OP
}
}
File diff suppressed because it is too large. Load diff
@@ -91,7 +91,7 @@ bool X86JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint6
size_t DataIndex{};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_A_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
LOGMAN_THROW_AA_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
+2 -2
View File
@@ -37,11 +37,11 @@ LookupCache::LookupCache(FEXCore::Context::Context *CTX)
// We currently limit to 128MB of real memory for caching for the total cache size.
// Can end up being inefficient if we compile a small number of blocks per page
PageMemory = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, CODE_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
LOGMAN_THROW_AA_FMT(PageMemory != -1ULL, "Failed to allocate page memory");
// L1 Cache
L1Pointer = reinterpret_cast<uintptr_t>(FEXCore::Allocator::mmap(nullptr, L1_SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0));
LOGMAN_THROW_A_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
LOGMAN_THROW_AA_FMT(L1Pointer != -1ULL, "Failed to allocate L1Pointer");
VirtualMemSize = ctx->Config.VirtualMemSize;
}
+10 -9
View File
@@ -8,6 +8,7 @@
#include <utility>
#include <vector>
#include <mutex>
#include <tsl/robin_map.h>
namespace FEXCore {
namespace Context {
@@ -17,7 +18,7 @@ namespace Context {
class LookupCache {
public:
struct LookupCacheEntry {
struct LookupCacheEntry {
uintptr_t HostCode;
uintptr_t GuestCode;
};
@@ -54,7 +55,7 @@ public:
return L1Entry.HostCode;
}
}
// Try L3
auto HostCode = BlockList.find(Address);
@@ -62,7 +63,7 @@ public:
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
}
// Failed to find
return 0;
}
@@ -73,7 +74,7 @@ public:
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length -1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
@@ -88,9 +89,9 @@ public:
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void *HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_A_FMT(Inserted, "Duplicate block mapping added");
LOGMAN_THROW_AA_FMT(Inserted, "Duplicate block mapping added");
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
@@ -158,7 +159,7 @@ public:
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::DebugStore,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
@@ -167,7 +168,7 @@ public:
std::recursive_mutex WriteLock;
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Do L1
@@ -239,7 +240,7 @@ private:
std::map<BlockLinkTag, std::function<void()>> BlockLinks;
std::map<uint64_t, uint64_t> BlockList;
tsl::robin_map<uint64_t, uint64_t> BlockList;
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(LookupCacheEntry);
File diff suppressed because it is too large. Load diff
+38 -4
View File
@@ -76,10 +76,10 @@ public:
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
// Used during new op bringup
bool ShouldDump {false};
struct JumpTargetInfo {
OrderedNode* BlockEntry;
bool HaveEmitted;
@@ -278,6 +278,12 @@ public:
void NOTOp(OpcodeArgs);
void XADDOp(OpcodeArgs);
void PopcountOp(OpcodeArgs);
void DAAOp(OpcodeArgs);
void DASOp(OpcodeArgs);
void AAAOp(OpcodeArgs);
void AASOp(OpcodeArgs);
void AAMOp(OpcodeArgs);
void AADOp(OpcodeArgs);
void XLATOp(OpcodeArgs);
template<bool Reseed>
void RDRANDOp(OpcodeArgs);
@@ -292,6 +298,8 @@ public:
void WriteSegmentReg(OpcodeArgs);
void EnterOp(OpcodeArgs);
void SGDTOp(OpcodeArgs);
// SSE
void MOVAPSOp(OpcodeArgs);
void MOVUPSOp(OpcodeArgs);
@@ -398,6 +406,26 @@ public:
// ADX Ops
void ADXOp(OpcodeArgs);
// AVX Ops
template <IROps IROp, size_t ElementSize>
void AVXVectorALUOp(OpcodeArgs);
void VANDNOp(OpcodeArgs);
void VMOVAPS_VMOVAPD_Op(OpcodeArgs);
void VMOVUPS_VMOVUPD_Op(OpcodeArgs);
void VMOVHPOp(OpcodeArgs);
void VMOVLPOp(OpcodeArgs);
void VMOVDDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVVectorNTOp(OpcodeArgs);
void VZEROOp(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -515,7 +543,7 @@ public:
void X87FRSTORF64(OpcodeArgs);
void X87FXAMF64(OpcodeArgs);
void X87LDENVF64(OpcodeArgs);
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMIF64(OpcodeArgs);
@@ -646,6 +674,7 @@ private:
OrderedNode *Current_HeaderNode{};
OrderedNode *AppendSegmentOffset(OrderedNode *Value, uint32_t Flags, uint32_t DefaultPrefix = 0, bool Override = false);
void UpdatePrefixFromSegment(OrderedNode *Segment, uint32_t SegmentReg);
enum class MemoryAccessType {
// Choose TSO or Non-TSO depending on access type
@@ -657,6 +686,11 @@ private:
// Non-temporal streaming
ACCESS_STREAM,
};
OrderedNode *LoadGPRRegister(uint32_t GPR, int8_t Size = -1, uint8_t Offset = 0);
OrderedNode *LoadXMMRegister(uint32_t XMM);
void StoreGPRRegister(uint32_t GPR, OrderedNode *const Src, int8_t Size = -1, uint8_t Offset = 0);
void StoreXMMRegister(uint32_t XMM, OrderedNode *const Src);
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
@@ -665,7 +699,7 @@ private:
void StoreResult(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp Op, OrderedNode *const Src, int8_t Align, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
[[nodiscard]] static uint32_t GPROffset(X86State::X86Reg reg) {
LOGMAN_THROW_A_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
LOGMAN_THROW_AA_FMT(reg <= X86State::X86Reg::REG_R15, "Invalid reg used");
return static_cast<uint32_t>(offsetof(Core::CPUState, gregs[static_cast<size_t>(reg)]));
}
@@ -221,7 +221,8 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *XMM0 = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
@@ -769,7 +769,11 @@ void OpDispatchBuilder::CalculcateFlags_ShiftLeftImmediate(uint8_t SrcSize, Orde
// CF
{
// Extract the last bit shifted in to CF
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, SrcSize * 8 - Shift, Src1));
auto OpSize = SrcSize * 8;
if (OpSize < Shift) {
Shift &= (OpSize - 1);
}
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, OpSize - Shift, Src1));
}
// PF
@@ -934,6 +938,7 @@ void OpDispatchBuilder::CalculcateFlags_RotateRight(uint8_t SrcSize, OrderedNode
auto OldOF = GetRFLAG(FEXCore::X86State::RFLAG_OF_LOC);
// OF is set to the XOR of the new CF bit and the most significant bit of the result
// OF is architecturally only defined for 1-bit rotate, which is why this only happens when the shift is one.
auto NewOF = _Xor(_Bfe(1, OpSize - 2, Res), NewCF);
// If shift == 0, don't update flags
@@ -963,7 +968,9 @@ void OpDispatchBuilder::CalculcateFlags_RotateLeft(uint8_t SrcSize, OrderedNode
// OF
{
auto OldOF = GetRFLAG(FEXCore::X86State::RFLAG_OF_LOC);
// OF is set to the XOR of the new CF bit and the most significant bit of the result
// OF is the LSB and MSB XOR'd together.
// OF is set to the XOR of the new CF bit and the most significant bit of the result.
// OF is architecturally only defined for 1-bit rotate, which is why this only happens when the shift is one.
auto NewOF = _Xor(_Bfe(1, OpSize - 1, Res), NewCF);
auto OF = _Select(FEXCore::IR::COND_EQ, Src2, _Constant(0), OldOF, NewOF);
@@ -977,8 +984,7 @@ void OpDispatchBuilder::CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, Or
if (Shift == 0) return;
auto OpSize = SrcSize * 8;
auto NewCF = _Bfe(1, OpSize - Shift, Src1);
auto NewCF = _Bfe(1, OpSize - 1, Res);
// CF
{
@@ -989,8 +995,10 @@ void OpDispatchBuilder::CalculcateFlags_RotateRightImmediate(uint8_t SrcSize, Or
// OF
{
if (Shift == 1) {
// OF is set to the XOR of the new CF bit and the most significant bit of the result
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(_Bfe(1, OpSize - 1, Res), NewCF));
// OF is the top two MSBs XOR'd together
// OF is architecturally only defined for 1-bit rotate, which is why this only happens when the shift is one.
auto NewOF = _Xor(_Bfe(1, OpSize - 2, Res), NewCF);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(NewOF);
}
}
}
@@ -1000,17 +1008,22 @@ void OpDispatchBuilder::CalculcateFlags_RotateLeftImmediate(uint8_t SrcSize, Ord
auto OpSize = SrcSize * 8;
auto NewCF = _Bfe(1, 0, Res);
// CF
{
// Extract the last bit shifted in to CF
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(_Bfe(1, Shift, Src1));
SetRFLAG<FEXCore::X86State::RFLAG_CF_LOC>(NewCF);
}
// OF
{
if (Shift == 1) {
// OF is the top two MSBs XOR'd together
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(_Xor(_Bfe(1, OpSize - 1, Src1), _Bfe(1, OpSize - 2, Src1)));
// OF is the LSB and MSB XOR'd together.
// OF is set to the XOR of the new CF bit and the most significant bit of the result.
// OF is architecturally only defined for 1-bit rotate, which is why this only happens when the shift is one.
auto NewOF = _Xor(_Bfe(1, OpSize - 1, Res), NewCF);
SetRFLAG<FEXCore::X86State::RFLAG_OF_LOC>(NewOF);
}
}
}
@@ -33,11 +33,46 @@ void OpDispatchBuilder::MOVVectorNTOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::VMOVVectorNTOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1, true, false, MemoryAccessType::ACCESS_STREAM);
// TODO: When stores and loads gain the ability to explicitly express
// whether a vector extension or an insert is desirable, ensure
// the 128-bit case here is a zero extend on store if the destination
// is a register.
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::MOVAPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
StoreResult(FPRClass, Op, Src, -1);
}
void OpDispatchBuilder::VMOVAPS_VMOVAPD_Op(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Is128BitDest = GetDstSize(Op) == Core::CPUState::XMM_SSE_REG_SIZE;
if (Op->Dest.IsGPR() && Is128BitDest) {
// Perform 32 byte store to clear the upper lane.
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 32, -1);
} else {
StoreResult(FPRClass, Op, Src, -1);
}
}
void OpDispatchBuilder::VMOVUPS_VMOVUPD_Op(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1);
const auto Is128BitDest = GetDstSize(Op) == Core::CPUState::XMM_SSE_REG_SIZE;
if (Op->Dest.IsGPR() && Is128BitDest) {
// Perform 32 byte store to clear the upper lane.
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 32, 1);
} else {
StoreResult(FPRClass, Op, Src, 1);
}
}
void OpDispatchBuilder::MOVUPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1);
StoreResult(FPRClass, Op, Src, 1);
@@ -68,20 +103,33 @@ void OpDispatchBuilder::MOVHPDOp(OpcodeArgs) {
}
}
void OpDispatchBuilder::VMOVHPOp(OpcodeArgs) {
if (Op->Dest.IsGPR()) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 8, 1, 0, Src1, Src2);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
} else {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Result = _VInsElement(16, 8, 0, 1, Src, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 8, 8);
}
}
void OpDispatchBuilder::MOVLPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
if (Op->Dest.IsGPR()) {
// xmm, xmm is movhlps special case
if (Op->Src[0].IsGPR()) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, 8, 16);
Src = _VExtractElement(16, 8, Src, 1);
auto Result = _VInsScalarElement(16, 8, 0, Dest, Src);
auto Result = _VInsElement(16, 8, 0, 1, Dest, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 16, 16);
}
else {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, 8, 16);
auto Result = _VInsScalarElement(16, 8, 0, Dest, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 8, 16);
auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
auto Result = _VInsElement(16, 8, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
}
else {
@@ -89,30 +137,68 @@ void OpDispatchBuilder::MOVLPOp(OpcodeArgs) {
}
}
void OpDispatchBuilder::VMOVLPOp(OpcodeArgs) {
if (Op->Dest.IsGPR()) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 8, 0, 0, Src1, Src2);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
} else {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 8, 8);
}
}
void OpDispatchBuilder::MOVSHDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 4, 3, 3, Src, Src);
Result = _VInsElement(16, 4, 2, 3, Result, Src);
Result = _VInsElement(16, 4, 1, 1, Result, Src);
OrderedNode *Result = _VInsElement(16, 4, 2, 3, Src, Src);
Result = _VInsElement(16, 4, 0, 1, Result, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VMOVSHDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto SrcSize = GetSrcSize(Op);
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
OrderedNode *Result = _VInsElement(SrcSize, 4, 2, 3, Src, Src);
Result = _VInsElement(SrcSize, 4, 0, 1, Result, Src);
if (Is256Bit) {
Result = _VInsElement(SrcSize, 4, 4, 5, Result, Src);
Result = _VInsElement(SrcSize, 4, 6, 7, Result, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
}
void OpDispatchBuilder::MOVSLDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 4, 3, 2, Src, Src);
Result = _VInsElement(16, 4, 2, 2, Result, Src);
Result = _VInsElement(16, 4, 1, 0, Result, Src);
Result = _VInsElement(16, 4, 0, 0, Result, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VMOVSLDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto SrcSize = GetSrcSize(Op);
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
OrderedNode *Result = _VInsElement(SrcSize, 4, 3, 2, Src, Src);
Result = _VInsElement(SrcSize, 4, 1, 0, Result, Src);
if (Is256Bit) {
Result = _VInsElement(SrcSize, 4, 5, 4, Result, Src);
Result = _VInsElement(SrcSize, 4, 7, 6, Result, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
}
void OpDispatchBuilder::MOVSSOp(OpcodeArgs) {
if (Op->Dest.IsGPR() && Op->Src[0].IsGPR()) {
// MOVSS xmm1, xmm2
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, 16, Op->Flags, -1);
OrderedNode *Src = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 4, Op->Flags, -1);
auto Result = _VInsScalarElement(16, 4, 0, Dest, Src);
auto Result = _VInsElement(16, 4, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
else if (Op->Dest.IsGPR()) {
@@ -133,7 +219,7 @@ void OpDispatchBuilder::MOVSDOp(OpcodeArgs) {
// xmm1[63:0] <- xmm2[63:0]
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
auto Result = _VInsScalarElement(16, 8, 0, Dest, Src);
auto Result = _VInsElement(16, 8, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
else if (Op->Dest.IsGPR()) {
@@ -302,6 +388,48 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 2>(OpcodeArgs);
template <IROps IROp, size_t ElementSize>
void OpDispatchBuilder::AVXVectorALUOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
const auto Is128Bit = Size == Core::CPUState::XMM_SSE_REG_SIZE;
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
auto ALUOp = _VAdd(Size, ElementSize, Src1, Src2);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
OrderedNode* Result = ALUOp;
if (Is128Bit) {
// 128-bit variants need to zero the upper lane.
Result = _VMov(Size, ALUOp);
}
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 1>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 2>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 4>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 8>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VFADD, 4>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VFADD, 8>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VAND, 16>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VOR, 16>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VXOR, 16>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorALUROp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
@@ -322,8 +450,9 @@ void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// If OpSize == ElementSize then it only does the lower scalar op
@@ -333,9 +462,9 @@ void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
OrderedNode* Result = ALUOp;
if (Size != ElementSize) {
if (DstSize != ElementSize) {
// Insert the lower bits
Result = _VInsScalarElement(Size, ElementSize, 0, Dest, Result);
Result = _VInsElement(DstSize, ElementSize, 0, 0, Dest, ALUOp);
}
StoreResult(FPRClass, Op, Result, -1);
@@ -369,11 +498,12 @@ void OpDispatchBuilder::VectorScalarALUOp<IR::OP_VFMAX, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize, bool Scalar>
void OpDispatchBuilder::VectorUnaryOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
auto DstSize = GetDstSize(Op);
if constexpr (Scalar) {
Size = ElementSize;
}
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
auto ALUOp = _VFSqrt(Size, ElementSize, Src);
// Overwrite our IR's op type
@@ -381,7 +511,7 @@ void OpDispatchBuilder::VectorUnaryOp(OpcodeArgs) {
if constexpr (Scalar) {
// Insert the lower bits
auto Result = _VInsScalarElement(GetSrcSize(Op), ElementSize, 0, Dest, ALUOp);
auto Result = _VInsElement(DstSize, ElementSize, 0, 0, Dest, ALUOp);
StoreResult(FPRClass, Op, Result, -1);
}
else {
@@ -440,10 +570,10 @@ void OpDispatchBuilder::MOVQOp(OpcodeArgs) {
// This instruction is a bit special that if the destination is a register then it'll ZEXT the 64bit source to 128bit
if (Op->Dest.IsGPR()) {
const auto gpr = Op->Dest.Data.GPR.GPR;
const auto gprIndex = gpr - X86State::REG_XMM_0;
_StoreContext(8, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][0]));
auto Const = _Constant(0);
_StoreContext(8, GPRClass, Const, offsetof(FEXCore::Core::CPUState, xmm[gpr - FEXCore::X86State::REG_XMM_0][1]));
auto Reg = _VMov(16, Src);
StoreXMMRegister(gprIndex, Reg);
}
else {
// This is simple, just store the result
@@ -562,7 +692,7 @@ void OpDispatchBuilder::PSHUFBOp(OpcodeArgs) {
template<size_t ElementSize, bool HalfSize, bool Low>
void OpDispatchBuilder::PSHUFDOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
LOGMAN_THROW_AA_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
uint8_t Shuffle = Op->Src[1].Data.Literal.Value;
@@ -599,7 +729,7 @@ void OpDispatchBuilder::PSHUFDOp<4, false, true>(OpcodeArgs);
template<size_t ElementSize>
void OpDispatchBuilder::SHUFOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(ElementSize != 0, "What. No element size?");
LOGMAN_THROW_AA_FMT(ElementSize != 0, "What. No element size?");
const auto Size = GetSrcSize(Op);
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -654,6 +784,22 @@ void OpDispatchBuilder::ANDNOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Dest, -1);
}
void OpDispatchBuilder::VANDNOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
const auto Is128Bit = Size == Core::CPUState::XMM_SSE_REG_SIZE;
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
Src1 = _VNot(Size, Size, Src1);
OrderedNode *Dest = _VAnd(Size, Size, Src1, Src2);
if (Is128Bit) {
Dest = _VMov(16, Dest);
}
StoreResult(FPRClass, Op, Dest, -1);
}
template<size_t ElementSize>
void OpDispatchBuilder::PINSROp(OpcodeArgs) {
auto Size = GetDstSize(Op);
@@ -921,7 +1067,11 @@ void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
auto Size = GetDstSize(Op);
auto Result = _VSRI(Size, 16, Dest, Shift);
OrderedNode *Result = _VectorZero(Size);
if (Shift < Size) {
Result = _VExtr(Size, 1, Result, Dest, Shift);
}
StoreResult(FPRClass, Op, Result, -1);
}
@@ -933,7 +1083,10 @@ void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
auto Size = GetDstSize(Op);
auto Result = _VSLI(Size, 16, Dest, Shift);
OrderedNode *Result = _VectorZero(Size);
if (Shift < Size) {
Result = _VExtr(Size, 1, Dest, Result, Size - Shift);
}
StoreResult(FPRClass, Op, Result, -1);
}
@@ -972,10 +1125,28 @@ void OpDispatchBuilder::PAVGOp<2>(OpcodeArgs);
void OpDispatchBuilder::MOVDDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Res = _SplatVector2(Src);
OrderedNode *Res = _VDupElement(16, GetSrcSize(Op), Src, 0);
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::VMOVDDUPOp(OpcodeArgs) {
const auto SrcSize = GetSrcSize(Op);
const auto IsSrcGPR = Op->Src[0].IsGPR();
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto MemSize = Is256Bit ? 32 : 8;
OrderedNode *Src = IsSrcGPR ? LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags, -1)
: LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], MemSize, Op->Flags, -1);
OrderedNode *Res = _VInsElement(SrcSize, 8, 1, 0, Src, Src);
if (Is256Bit) {
Res = _VInsElement(SrcSize, 8, 3, 2, Res, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Res, 32, -1);
}
template<size_t DstElementSize>
void OpDispatchBuilder::CVTGPR_To_FPR(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -986,7 +1157,7 @@ void OpDispatchBuilder::CVTGPR_To_FPR(OpcodeArgs) {
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, 16, Op->Flags, -1);
Src = _VInsScalarElement(16, DstElementSize, 0, Dest, Src);
Src = _VInsElement(16, DstElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Src, -1);
}
@@ -1081,13 +1252,16 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, false>(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void OpDispatchBuilder::Scalar_CVT_Float_To_Float(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
const auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
Src = _Float_FToF(DstElementSize, SrcElementSize, Src);
Src = _VInsScalarElement(16, DstElementSize, 0, Dest, Src);
Src = _VInsElement(16, DstElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Src, -1);
auto Result = _VInsElement(DstSize, DstElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
template
@@ -1097,8 +1271,9 @@ void OpDispatchBuilder::Scalar_CVT_Float_To_Float<8, 4>(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void OpDispatchBuilder::Vector_CVT_Float_To_Float(OpcodeArgs) {
const auto Size = GetDstSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
if constexpr (DstElementSize > SrcElementSize) {
Src = _Vector_FToF(Size, SrcElementSize << 1, Src, SrcElementSize);
@@ -1175,13 +1350,12 @@ void OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, true>(OpcodeArgs);
void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
// Until we get correct PHI nodes this is required to be a loop unroll
const auto GPRSize = CTX->GetGPRSize();
const auto Size = uint32_t{GetSrcSize(Op)} * 8;
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *MemDest = _LoadContext(GPRSize, GPRClass, GPROffset(X86State::REG_RDI));
auto MemDest = LoadGPRRegister(X86State::REG_RDI);
const size_t NumElements = Size / 64;
for (size_t Element = 0; Element < NumElements; ++Element) {
@@ -1275,7 +1449,7 @@ void OpDispatchBuilder::VFCMPOp(OpcodeArgs) {
if constexpr (Scalar) {
// Insert the lower bits
Result = _VInsScalarElement(GetDstSize(Op), ElementSize, 0, Dest, Result);
Result = _VInsElement(GetDstSize(Op), ElementSize, 0, 0, Dest, Result);
}
StoreResult(FPRClass, Op, Result, -1);
@@ -1390,10 +1564,11 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
_StoreMem(FPRClass, 16, MemLocation, MMReg, 16);
}
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *XMMReg = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[i]));
OrderedNode *XMMReg = LoadXMMRegister(i);
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
_StoreMem(FPRClass, 16, MemLocation, XMMReg, 16);
@@ -1439,12 +1614,13 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
auto MMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(16, FPRClass, MMReg, offsetof(FEXCore::Core::CPUState, mm[i]));
}
unsigned NumRegs = CTX->Config.Is64BitMode ? 16 : 8;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
auto XMMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(16, FPRClass, XMMReg, offsetof(FEXCore::Core::CPUState, xmm[i]));
StoreXMMRegister(i, XMMReg);
}
}
@@ -1541,6 +1717,9 @@ void OpDispatchBuilder::PACKSSOp<4>(OpcodeArgs);
template<size_t ElementSize, bool Signed>
void OpDispatchBuilder::PMULLOp(OpcodeArgs) {
static_assert(ElementSize == sizeof(uint32_t),
"Currently only handles 32-bit -> 64-bit");
auto Size = GetSrcSize(Op);
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
@@ -1557,17 +1736,8 @@ void OpDispatchBuilder::PMULLOp(OpcodeArgs) {
}
}
else {
OrderedNode* Srcs1[2]{};
OrderedNode* Srcs2[2]{};
Srcs1[0] = _VExtr(Size, ElementSize, Src1, Src1, 0);
Srcs1[1] = _VExtr(Size, ElementSize, Src1, Src1, 2);
Srcs2[0] = _VExtr(Size, ElementSize, Src2, Src2, 0);
Srcs2[1] = _VExtr(Size, ElementSize, Src2, Src2, 2);
Src1 = _VInsElement(Size, ElementSize, 1, 0, Srcs1[0], Srcs1[1]);
Src2 = _VInsElement(Size, ElementSize, 1, 0, Srcs2[0], Srcs2[1]);
Src1 = _VInsElement(Size, ElementSize, 1, 2, Src1, Src1);
Src2 = _VInsElement(Size, ElementSize, 1, 2, Src2, Src2);
if constexpr (Signed) {
Res = _VSMull(Size, ElementSize, Src1, Src2);
@@ -1590,8 +1760,10 @@ void OpDispatchBuilder::MOVQ2DQ(OpcodeArgs) {
// This instruction is a bit special in that if the source is MMX then it zexts to 128bit
if constexpr (ToXMM) {
const auto Index = Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0;
Src = _VMov(16, Src);
_StoreContext(16, FPRClass, Src, offsetof(FEXCore::Core::CPUState, xmm[Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0][0]));
StoreXMMRegister(Index, Src);
}
else {
// This is simple, just store the result
@@ -1688,8 +1860,8 @@ void OpDispatchBuilder::PFNACCOp(OpcodeArgs) {
OrderedNode *ResSubSrc{};
OrderedNode *ResSubDest{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
auto UpperSubSrc = _VExtractElement(Size, 4, Src, 1);
auto UpperSubDest = _VDupElement(Size, 4, Dest, 1);
auto UpperSubSrc = _VDupElement(Size, 4, Src, 1);
ResSubDest = _VFSub(4, 4, Dest, UpperSubDest);
ResSubSrc = _VFSub(4, 4, Src, UpperSubSrc);
@@ -1707,7 +1879,7 @@ void OpDispatchBuilder::PFPNACCOp(OpcodeArgs) {
OrderedNode *ResAdd{};
OrderedNode *ResSub{};
auto UpperSubDest = _VExtractElement(Size, 4, Dest, 1);
auto UpperSubDest = _VDupElement(Size, 4, Dest, 1);
ResSub = _VFSub(4, 4, Dest, UpperSubDest);
ResAdd = _VFAddP(Size, 4, Src, Src);
@@ -1840,8 +2012,6 @@ void OpDispatchBuilder::PMADDWD(OpcodeArgs) {
if (Size == 8) {
Size <<= 1;
Src1 = _VBitcast(Size, 2, Src1);
Src2 = _VBitcast(Size, 2, Src2);
}
auto Src1_L = _VSXTL(Size, 2, Src1); // [15:0 ], [31:16], [32:47 ], [63:48 ]
@@ -1928,9 +2098,6 @@ void OpDispatchBuilder::PMULHW(OpcodeArgs) {
OrderedNode *Res{};
if (Size == 8) {
Dest = _VBitcast(Size * 2, 2, Dest);
Src = _VBitcast(Size * 2, 2, Src);
// Implementation is more efficient for 8byte registers
if (Signed)
Res = _VSMull(Size * 2, 2, Dest, Src);
@@ -2307,7 +2474,7 @@ void OpDispatchBuilder::VectorRound(OpcodeArgs) {
if constexpr (Scalar) {
// Insert the lower bits
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
auto Result = _VInsScalarElement(GetDstSize(Op), ElementSize, 0, Dest, Src);
auto Result = _VInsElement(GetDstSize(Op), ElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
else {
@@ -2356,7 +2523,8 @@ void OpDispatchBuilder::VectorVariableBlend(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// The mask is hardcoded to be xmm0 in this instruction
OrderedNode *Mask = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm[0]));
auto Mask = LoadXMMRegister(0);
// Each element is selected by the high bit of that element size
// Dest[ElementIdx] = Xmm0[ElementIndex][HighBit] ? Src : Dest;
//
@@ -2591,4 +2759,28 @@ void OpDispatchBuilder::MPSADBWOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VZEROOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto IsVZEROALL = DstSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
if (IsVZEROALL) {
// NOTE: Despite the name being VZEROALL, this will still only ever
// zero out up to the first 16 registers (even on AVX-512, where we have 32 registers)
OrderedNode* ZeroVector = _VectorZero(DstSize);
for (uint32_t i = 0; i < NumRegs; i++) {
StoreXMMRegister(i, ZeroVector);
}
} else {
// Likewise, VZEROUPPER will only ever zero only up to the first 16 registers
for (uint32_t i = 0; i < NumRegs; i++) {
OrderedNode* Reg = LoadXMMRegister(i);
OrderedNode* Dst = _VMov(16, Reg);
StoreXMMRegister(i, Dst);
}
}
}
}
@@ -782,6 +782,12 @@ void OpDispatchBuilder::X87UnaryOp(OpcodeArgs) {
// Overwrite the op
result.first->Header.Op = IROp;
if constexpr (IROp == IR::OP_F80SIN ||
IROp == IR::OP_F80COS) {
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
// Write to ST[TOP]
_StoreContextIndexed(result, top, 16, MMBaseOffset(), 16, FPRClass);
}
@@ -809,7 +815,8 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
// Overwrite the op
result.first->Header.Op = IROp;
if constexpr (IROp == IR::OP_F80FPREM) {
if constexpr (IROp == IR::OP_F80FPREM ||
IROp == IR::OP_F80FPREM1) {
//TODO: Set C0 to Q2, C3 to Q1, C1 to Q0
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
@@ -854,6 +861,9 @@ void OpDispatchBuilder::X87SinCos(OpcodeArgs) {
auto sin = _F80SIN(a);
auto cos = _F80COS(a);
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
// Write to ST[TOP]
_StoreContextIndexed(sin, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(cos, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -900,6 +910,9 @@ void OpDispatchBuilder::X87TAN(OpcodeArgs) {
OrderedNode *data = _VCastFromGPR(16, 8, low);
data = _VInsGPR(16, 8, 1, data, high);
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
// Write to ST[TOP]
_StoreContextIndexed(result, orig_top, 16, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
@@ -1185,7 +1198,7 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
// upper 16 bits [79:64]
_StoreMem(FPRClass, 8, ST0Location, data, 1);
ST0Location = _Add(ST0Location, _Constant(8));
auto topBytes = _VExtractElement(16, 2, data, 4);
auto topBytes = _VDupElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
// reset to default
@@ -778,6 +778,12 @@ void OpDispatchBuilder::X87UnaryOpF64(OpcodeArgs) {
// Overwrite the op
result.first->Header.Op = IROp;
if constexpr (IROp == IR::OP_F64SIN ||
IROp == IR::OP_F64COS) {
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
// Write to ST[TOP]
_StoreContextIndexed(result, top, 8, MMBaseOffset(), 16, FPRClass);
}
@@ -804,7 +810,8 @@ void OpDispatchBuilder::X87BinaryOpF64(OpcodeArgs) {
// Overwrite the op
result.first->Header.Op = IROp;
if constexpr (IROp == IR::OP_F64FPREM) {
if constexpr (IROp == IR::OP_F80FPREM ||
IROp == IR::OP_F80FPREM1) {
//TODO: Set C0 to Q2, C3 to Q1, C1 to Q0
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
@@ -831,6 +838,9 @@ void OpDispatchBuilder::X87SinCosF64(OpcodeArgs) {
auto sin = _F64SIN(a);
auto cos = _F64COS(a);
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
// Write to ST[TOP]
_StoreContextIndexed(sin, orig_top, 8, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(cos, top, 8, MMBaseOffset(), 16, FPRClass);
@@ -871,6 +881,9 @@ void OpDispatchBuilder::X87TANF64(OpcodeArgs) {
auto one = _VCastFromGPR(8, 8, _Constant(0x3FF0000000000000));
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
// Write to ST[TOP]
_StoreContextIndexed(result, orig_top, 8, MMBaseOffset(), 16, FPRClass);
_StoreContextIndexed(one, top, 8, MMBaseOffset(), 16, FPRClass);
@@ -996,7 +1009,7 @@ void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
// upper 16 bits [79:64]
_StoreMem(FPRClass, 8, ST0Location, data, 1);
ST0Location = _Add(ST0Location, _Constant(8));
auto topBytes = _VExtractElement(16, 2, data, 4);
auto topBytes = _VDupElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
// reset to default
@@ -266,10 +266,10 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0x17, 1, X86InstInfo{"POP SS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1E, 1, X86InstInfo{"PUSH DS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x1F, 1, X86InstInfo{"POP DS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS, 0, nullptr}},
{0x27, 1, X86InstInfo{"DAA", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x2F, 1, X86InstInfo{"DAS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x37, 1, X86InstInfo{"AAA", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x3F, 1, X86InstInfo{"AAS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x27, 1, X86InstInfo{"DAA", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x2F, 1, X86InstInfo{"DAS", TYPE_INST, GenFlagsDstSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x37, 1, X86InstInfo{"AAA", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x3F, 1, X86InstInfo{"AAS", TYPE_INST, GenFlagsDstSize(SIZE_16BIT) | FLAGS_SF_DST_RAX, 0, nullptr}},
{0x40, 8, X86InstInfo{"INC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
{0x48, 8, X86InstInfo{"DEC", TYPE_INST, FLAGS_SF_REX_IN_BYTE, 0, nullptr}},
@@ -283,8 +283,8 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xA1, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_DST_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xA3, 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_MEM_OFFSET, 4, nullptr}},
{0xCE, 1, X86InstInfo{"INTO", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xD4, 1, X86InstInfo{"AAM", TYPE_INST, FLAGS_NONE, 1, nullptr}},
{0xD5, 1, X86InstInfo{"AAD", TYPE_INST, FLAGS_NONE, 1, nullptr}},
{0xD4, 1, X86InstInfo{"AAM", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xD5, 1, X86InstInfo{"AAD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_SF_DST_RAX, 1, nullptr}},
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
@@ -67,7 +67,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_6, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// GROUP 7
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -76,7 +76,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -85,7 +85,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -94,7 +94,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -295,7 +295,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x20, 4, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x24, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SS", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2B, 1, X86InstInfo{"MOVNTSS", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2D, 1, X86InstInfo{"CVTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
@@ -305,18 +305,18 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x40, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x50, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 1, X86InstInfo{"RSQRTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x53, 1, X86InstInfo{"RCPSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 1, X86InstInfo{"RSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x53, 1, X86InstInfo{"RCPSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x54, 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSS2SD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSS2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5B, 1, X86InstInfo{"CVTTPS2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x60, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x68, 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -383,16 +383,16 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x40, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x50, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSD2SS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5B, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x60, 16, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -17,23 +17,23 @@ void InitializeVEXTables() {
static constexpr U16U8InfoStruct VEXTable[] = {
// Map 0 (Reserved)
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMODUPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMODUPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x14), 1, X86InstInfo{"VUNPCKLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x14), 1, X86InstInfo{"VUNPCKLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -41,12 +41,12 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x15), 1, X86InstInfo{"VUNPCKHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x15), 1, X86InstInfo{"VUNPCKHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x16), 1, X86InstInfo{"VMOVSHDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b10, 0x16), 1, X86InstInfo{"VMOVSHDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x50), 1, X86InstInfo{"VMOVMSKPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x50), 1, X86InstInfo{"VMOVMSKPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -62,17 +62,17 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x53), 1, X86InstInfo{"VRCPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x55), 1, X86InstInfo{"VANDNPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x55), 1, X86InstInfo{"VANDNPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x55), 1, X86InstInfo{"VANDNPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x55), 1, X86InstInfo{"VANDNPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x56), 1, X86InstInfo{"VORPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x56), 1, X86InstInfo{"VORPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x56), 1, X86InstInfo{"VORPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x56), 1, X86InstInfo{"VORPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x57), 1, X86InstInfo{"VXORPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x57), 1, X86InstInfo{"VDORPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x57), 1, X86InstInfo{"VXORPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x57), 1, X86InstInfo{"VXORPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x60), 1, X86InstInfo{"VPUNPCKLBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x61), 1, X86InstInfo{"VPUNPCKLWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -95,7 +95,7 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x75), 1, X86InstInfo{"VPCMPEQW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x76), 1, X86InstInfo{"VPCMPEQD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x77), 1, X86InstInfo{"VZERO*", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x77), 1, X86InstInfo{"VZERO*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT), 0, nullptr}},
{OPD(1, 0b00, 0xC2), 1, X86InstInfo{"VCMPccPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC2), 1, X86InstInfo{"VCMPccPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -112,17 +112,17 @@ void InitializeVEXTables() {
// This table doesn't state which VEX.pp is for which instruction
// XXX: Confirm all the above encoding opcodes
{OPD(1, 0b00, 0x28), 1, X86InstInfo{"VMOVAPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x28), 1, X86InstInfo{"VMOVAPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x28), 1, X86InstInfo{"VMOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x28), 1, X86InstInfo{"VMOVAPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x29), 1, X86InstInfo{"VMOVAPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x29), 1, X86InstInfo{"VMOVAPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x29), 1, X86InstInfo{"VMOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x29), 1, X86InstInfo{"VMOVAPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2A), 1, X86InstInfo{"VCVTSI2SS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2A), 1, X86InstInfo{"VCVTSI2SD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -136,8 +136,8 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VUCOMISS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VUCOMISD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -179,8 +179,8 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x6D), 1, X86InstInfo{"VPUNPCKHQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6E), 1, X86InstInfo{"VMOV*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7C), 1, X86InstInfo{"VHADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7C), 1, X86InstInfo{"VHADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -188,11 +188,11 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x7D), 1, X86InstInfo{"VHSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7D), 1, X86InstInfo{"VHSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
{OPD(1, 0b01, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
@@ -205,19 +205,19 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0xD1), 1, X86InstInfo{"VPSRLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD2), 1, X86InstInfo{"VPSRLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD3), 1, X86InstInfo{"VPSRLQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD5), 1, X86InstInfo{"VPMULLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD7), 1, X86InstInfo{"VPMOVMSKB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(1, 0b01, 0xD8), 1, X86InstInfo{"VPSUBUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD9), 1, X86InstInfo{"VPSUBUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDA), 1, X86InstInfo{"VPMINUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDB), 1, X86InstInfo{"VPAND", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDB), 1, X86InstInfo{"VPAND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDC), 1, X86InstInfo{"VPADDUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDD), 1, X86InstInfo{"VPADDUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDE), 1, X86InstInfo{"VPMAXUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDF), 1, X86InstInfo{"VPANDN", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDF), 1, X86InstInfo{"VPANDN", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE0), 1, X86InstInfo{"VPAVGB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE1), 1, X86InstInfo{"VPSRAW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -230,16 +230,16 @@ void InitializeVEXTables() {
{OPD(1, 0b10, 0xE6), 1, X86InstInfo{"VCVTDQ2PD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xE6), 1, X86InstInfo{"VCVTPD2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE7), 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE7), 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE8), 1, X86InstInfo{"VPSUBSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE9), 1, X86InstInfo{"VPSUBSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEB), 1, X86InstInfo{"VPOR", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEB), 1, X86InstInfo{"VPOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEC), 1, X86InstInfo{"VPADDSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xED), 1, X86InstInfo{"VPADDSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -255,9 +255,9 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0xF9), 1, X86InstInfo{"VPSUBW", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFA), 1, X86InstInfo{"VPSUBD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFB), 1, X86InstInfo{"VPSUBQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFC), 1, X86InstInfo{"VPADDB", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFD), 1, X86InstInfo{"VPADDW", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFE), 1, X86InstInfo{"VPADDD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFC), 1, X86InstInfo{"VPADDB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFD), 1, X86InstInfo{"VPADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFE), 1, X86InstInfo{"VPADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
// VEX Map 2
{OPD(2, 0b01, 0x00), 1, X86InstInfo{"VPSHUFB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -298,7 +298,7 @@ void InitializeVEXTables() {
{OPD(2, 0b01, 0x28), 1, X86InstInfo{"VPMULDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x29), 1, X86InstInfo{"VPCMPEQQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2B), 1, X86InstInfo{"VPACKUSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2C), 1, X86InstInfo{"VMASKMOVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2D), 1, X86InstInfo{"VMASKMOVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -33,7 +33,7 @@ static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInfoStruct<Op
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
FinalTable[OpNum + i] = Info;
#ifndef NDEBUG
++Total;
@@ -51,7 +51,7 @@ static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoS
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if (Info.Type == TYPE_COPY_OTHER) {
FinalTable[OpNum + i] = OtherLocal[OpNum + i];
}
@@ -74,7 +74,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_A_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if ((OpNum & 0b11'000'000) == 0b11'000'000) {
// If the mod field is 0b11 then it is a regular op
FinalTable[OpNum + i] = Info;
@@ -82,7 +82,7 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct
else {
// If the mod field is !0b11 then this instruction is duplicated through the whole mod [0b00, 0b10] range
// and the modrm.rm space because that is used part of the instruction encoding
LOGMAN_THROW_A_FMT((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
LOGMAN_THROW_AA_FMT((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
for (uint16_t mod = 0b00'000'000; mod < 0b11'000'000; mod += 0b01'000'000) {
for (uint16_t rm = 0b000; rm < 0b1'000; ++rm) {
FinalTable[(OpNum | mod | rm) + i] = Info;
+173 -104
View File
@@ -28,6 +28,10 @@ $end_info$
#include <string>
#include <utility>
#ifdef ENABLE_JEMALLOC
#include "jemalloc/jemalloc.h"
#endif
struct LoadlibArgs {
const char *Name;
};
@@ -66,12 +70,23 @@ namespace FEXCore {
struct ExportEntry { uint8_t *sha256; ThunkedFunction* Fn; };
struct TrampolineInstanceInfo {
uintptr_t HostPacker;
void* HostPacker;
uintptr_t CallCallback;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
};
// Opaque type pointing to an instance of HostToGuestTrampolineTemplate and its
// embedded TrampolineInstanceInfo
struct HostToGuestTrampolinePtr;
const auto HostToGuestTrampolineSize = __stop_HostToGuestTrampolineTemplate - __start_HostToGuestTrampolineTemplate;
static TrampolineInstanceInfo& GetInstanceInfo(HostToGuestTrampolinePtr* Trampoline) {
const auto Length = __stop_HostToGuestTrampolineTemplate - __start_HostToGuestTrampolineTemplate;
const auto InstanceInfoOffset = Length - sizeof(TrampolineInstanceInfo);
return *reinterpret_cast<TrampolineInstanceInfo*>(reinterpret_cast<char*>(Trampoline) + InstanceInfoOffset);
}
struct GuestcallInfo {
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
@@ -94,7 +109,9 @@ namespace FEXCore {
}
};
class ThunkHandler_impl final: public ThunkHandler {
HostToGuestTrampolinePtr* MakeHostTrampolineForGuestFunction(void* HostPacker, uintptr_t GuestTarget, uintptr_t GuestUnpacker);
struct ThunkHandler_impl final: public ThunkHandler {
std::shared_mutex ThunksMutex;
std::unordered_map<IR::SHA256Sum, ThunkedFunction*, TruncatingSHA256Hash> Thunks = {
@@ -108,23 +125,28 @@ namespace FEXCore {
{ 0xee, 0x57, 0xba, 0x0c, 0x5f, 0x6e, 0xef, 0x2a, 0x8c, 0xb5, 0x19, 0x81, 0xc9, 0x23, 0xe6, 0x51, 0xae, 0x65, 0x02, 0x8f, 0x2b, 0x5d, 0x59, 0x90, 0x6a, 0x7e, 0xe2, 0xe7, 0x1c, 0x33, 0x8a, 0xff },
&IsLibLoaded
},
{
// sha256(fex:is_host_heap_allocation)
{ 0xf5, 0x77, 0x68, 0x43, 0xbb, 0x6b, 0x28, 0x18, 0x40, 0xb0, 0xdb, 0x8a, 0x66, 0xfb, 0x0e, 0x2d, 0x98, 0xc2, 0xad, 0xe2, 0x5a, 0x18, 0x5a, 0x37, 0x2e, 0x13, 0xc9, 0xe7, 0xb9, 0x8c, 0xa9, 0x3e },
&IsHostHeapAllocation
},
{
// sha256(fex:link_address_to_function)
{ 0xe6, 0xa8, 0xec, 0x1c, 0x7b, 0x74, 0x35, 0x27, 0xe9, 0x4f, 0x5b, 0x6e, 0x2d, 0xc9, 0xa0, 0x27, 0xd6, 0x1f, 0x2b, 0x87, 0x8f, 0x2d, 0x35, 0x50, 0xea, 0x16, 0xb8, 0xc4, 0x5e, 0x42, 0xfd, 0x77 },
&LinkAddressToGuestFunction
},
{
// sha256(fex:make_host_trampoline_for_guest_function)
{ 0x1e, 0x51, 0x6b, 0x07, 0x39, 0xeb, 0x50, 0x59, 0xb3, 0xf3, 0x4f, 0xca, 0xdd, 0x58, 0x37, 0xe9, 0xf0, 0x30, 0xe5, 0x89, 0x81, 0xc7, 0x14, 0xfb, 0x24, 0xf9, 0xba, 0xe7, 0x0e, 0x00, 0x1e, 0x86 },
&MakeHostTrampolineForGuestFunction
}
// sha256(fex:allocate_host_trampoline_for_guest_function)
{ 0x9b, 0xb2, 0xf4, 0xb4, 0x83, 0x7d, 0x28, 0x93, 0x40, 0xcb, 0xf4, 0x7a, 0x0b, 0x47, 0x85, 0x87, 0xf9, 0xbc, 0xb5, 0x27, 0xca, 0xa6, 0x93, 0xa5, 0xc0, 0x73, 0x27, 0x24, 0xae, 0xc8, 0xb8, 0x5a },
&AllocateHostTrampolineForGuestFunction
},
};
// Can't be a string_view. We need to keep a copy of the library name in-case string_view pointer goes away.
// Ideally we track when a library has been unloaded and remove it from this set before the memory backing goes away.
std::set<std::string> Libs;
std::unordered_map<GuestcallInfo, uintptr_t, GuestcallInfoHash> GuestcallToHostTrampoline;
std::unordered_map<GuestcallInfo, HostToGuestTrampolinePtr*, GuestcallInfoHash> GuestcallToHostTrampoline;
uint8_t *HostTrampolineInstanceDataPtr;
size_t HostTrampolineInstanceDataAvailable = 0;
@@ -157,11 +179,11 @@ namespace FEXCore {
auto args = reinterpret_cast<args_t*>(argsv);
auto CTX = Thread->CTX;
LOGMAN_THROW_A_FMT(args->original_callee, "Tried to link null pointer address to guest function");
LOGMAN_THROW_A_FMT(args->target_addr, "Tried to link address to null pointer guest function");
LOGMAN_THROW_AA_FMT(args->original_callee, "Tried to link null pointer address to guest function");
LOGMAN_THROW_AA_FMT(args->target_addr, "Tried to link address to null pointer guest function");
if (!CTX->Config.Is64BitMode) {
LOGMAN_THROW_A_FMT((args->original_callee >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_A_FMT((args->target_addr >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_AA_FMT((args->original_callee >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
LOGMAN_THROW_AA_FMT((args->target_addr >> 32) == 0, "Tried to link 64-bit address in 32-bit mode");
}
LogMan::Msg::DFmt("Thunks: Adding guest trampoline from address {:#x} to guest function {:#x}",
@@ -177,7 +199,7 @@ namespace FEXCore {
const uint8_t GPRSize = CTX->GetGPRSize();
emit->_StoreContext(GPRSize, IR::GPRClass, emit->_Constant(Entrypoint), offsetof(Core::CPUState, gregs[X86State::REG_R11]));
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, gregs[X86State::REG_R11]), IR::GPRClass, IR::GPRFixedClass, GPRSize);
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
}, CTX->ThunkHandler.get(), (void*)args->target_addr);
@@ -194,98 +216,44 @@ namespace FEXCore {
}
/**
* Generates a host-callable trampoline to call guest functions via the host ABI.
* Guest-side helper to initiate creation of a host trampoline for
* calling guest functions. This must be followed by a host-side call
* to FinalizeHostTrampolineForGuestFunction to make the trampoline
* usable.
*
* This trampoline uses the same calling convention as the given HostPacker. Trampolines
* are cached, so it's safe to call this function repeatedly on the same arguments without
* leaking memory.
*
* Invoking the returned trampoline has the effect of:
* - packing the arguments (using the HostPacker identified by its SHA256)
* - performing a host->guest transition
* - unpacking the arguments via GuestUnpacker
* - calling the function at GuestTarget
*
* The primary use case of this is ensuring that guest function pointers ("callbacks")
* passed to thunked APIs can safely be called by the native host library.
* This two-step initialization is equivalent to a host-side call to
* MakeHostTrampolineForGuestFunction. The split is needed if the
* host doesn't have all information needed to create the trampoline
* on its own.
*/
static void MakeHostTrampolineForGuestFunction(void* ArgsRV) {
struct ArgsRV_t {
IR::SHA256Sum *HostPackerSha256;
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
uintptr_t rv; // Pointer to host trampoline + TrampolineInstanceInfo
} *args = reinterpret_cast<ArgsRV_t*>(ArgsRV);
static void AllocateHostTrampolineForGuestFunction(void* ArgsRV) {
struct ArgsRV_t {
uintptr_t GuestUnpacker;
uintptr_t GuestTarget;
uintptr_t rv; // Pointer to host trampoline + TrampolineInstanceInfo
} *args = reinterpret_cast<ArgsRV_t*>(ArgsRV);
LOGMAN_THROW_A_FMT(args->GuestTarget, "Tried to create host-trampoline to null pointer guest function");
args->rv = (uintptr_t)MakeHostTrampolineForGuestFunction(nullptr, args->GuestTarget, args->GuestUnpacker);
}
const auto CTX = Thread->CTX;
const auto ThunkHandler = reinterpret_cast<ThunkHandler_impl *>(CTX->ThunkHandler.get());
/**
* Checks if the given pointer is allocated on the host heap.
*
* This is useful for thunking APIs that need to work with both guest
* and host heap pointers.
*/
static void IsHostHeapAllocation(void* ArgsRV) {
#ifdef ENABLE_JEMALLOC
struct ArgsRV_t {
void* ptr;
bool rv;
} *args = reinterpret_cast<ArgsRV_t*>(ArgsRV);
const GuestcallInfo gci = { args->GuestUnpacker, args->GuestTarget };
// Try first with shared_lock
{
std::shared_lock lk(ThunkHandler->ThunksMutex);
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
std::lock_guard lk(ThunkHandler->ThunksMutex);
// Retry lookup with full lock before making a new trampoline to avoid double trampolines
{
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
args->rv = found->second;
return;
}
}
// No entry found => create new trampoline
auto HostPackerEntry = ThunkHandler->Thunks.find(*args->HostPackerSha256);
if (HostPackerEntry == ThunkHandler->Thunks.end()) {
ERROR_AND_DIE_FMT("Unknown host packing function for callback");
}
LogMan::Msg::DFmt("Thunks: Adding host trampoline for guest function {:#x}",
args->GuestTarget);
const auto Length = __stop_HostToGuestTrampolineTemplate - __start_HostToGuestTrampolineTemplate;
const auto InstanceInfoOffset = Length - sizeof(TrampolineInstanceInfo);
if (ThunkHandler->HostTrampolineInstanceDataAvailable < Length) {
const auto allocation_step = 16 * 1024;
ThunkHandler->HostTrampolineInstanceDataAvailable = allocation_step;
ThunkHandler->HostTrampolineInstanceDataPtr = (uint8_t *)mmap(
0, ThunkHandler->HostTrampolineInstanceDataAvailable,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
LOGMAN_THROW_A_FMT(ThunkHandler->HostTrampolineInstanceDataPtr != MAP_FAILED, "Failed to mmap HostTrampolineInstanceDataPtr");
}
const TrampolineInstanceInfo NewTrampolineInfo {
.HostPacker = reinterpret_cast<uintptr_t>(HostPackerEntry->second),
.CallCallback = (uintptr_t)&CallCallback,
.GuestUnpacker = args->GuestUnpacker,
.GuestTarget = args->GuestTarget
};
uint8_t* const HostTrampoline = ThunkHandler->HostTrampolineInstanceDataPtr;
ThunkHandler->HostTrampolineInstanceDataAvailable -= Length;
ThunkHandler->HostTrampolineInstanceDataPtr += Length;
memcpy(HostTrampoline, (void*)&HostToGuestTrampolineTemplate, Length);
memcpy(HostTrampoline + InstanceInfoOffset, &NewTrampolineInfo, sizeof(NewTrampolineInfo));
args->rv = reinterpret_cast<uintptr_t>(HostTrampoline);
ThunkHandler->GuestcallToHostTrampoline[gci] = args->rv;
args->rv = je_is_known_allocation(args->ptr);
#else
// Thunks usage without jemalloc isn't supported
ERROR_AND_DIE_FMT("Unsupported: Thunks querying for host heap allocation information");
#endif
}
static void LoadLib(void *ArgsV) {
@@ -352,9 +320,7 @@ namespace FEXCore {
}
}
public:
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) {
ThunkedFunction* LookupThunk(const IR::SHA256Sum &sha256) override {
std::shared_lock lk(ThunksMutex);
@@ -367,12 +333,115 @@ namespace FEXCore {
}
}
void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) {
void RegisterTLSState(FEXCore::Core::InternalThreadState *Thread) override {
::Thread = Thread;
}
void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) override {
for (auto & Definition : Definitions) {
Thunks.emplace(Definition.Sum, Definition.ThunkFunction);
}
}
};
ThunkHandler* ThunkHandler::Create() {
return new ThunkHandler_impl();
return new ThunkHandler_impl();
}
/**
* Generates a host-callable trampoline to call guest functions via the host ABI.
*
* This trampoline uses the same calling convention as the given HostPacker. Trampolines
* are cached, so it's safe to call this function repeatedly on the same arguments without
* leaking memory.
*
* Invoking the returned trampoline has the effect of:
* - packing the arguments (using the HostPacker identified by its SHA256)
* - performing a host->guest transition
* - unpacking the arguments via GuestUnpacker
* - calling the function at GuestTarget
*
* The primary use case of this is ensuring that guest function pointers ("callbacks")
* passed to thunked APIs can safely be called by the native host library.
*
* Returns a pointer to the generated host trampoline and its TrampolineInstanceInfo.
*
* If HostPacker is zero, the trampoline will be partially initialized and needs to be
* finalized with a call to FinalizeHostTrampolineForGuestFunction. A typical use case
* is to allocate the trampoline for a given GuestTarget/GuestUnpacker on the guest-side,
* and provide the HostPacker host-side.
*/
FEX_DEFAULT_VISIBILITY
HostToGuestTrampolinePtr* MakeHostTrampolineForGuestFunction(void* HostPacker, uintptr_t GuestTarget, uintptr_t GuestUnpacker) {
LOGMAN_THROW_AA_FMT(GuestTarget, "Tried to create host-trampoline to null pointer guest function");
const auto CTX = Thread->CTX;
const auto ThunkHandler = reinterpret_cast<ThunkHandler_impl *>(CTX->ThunkHandler.get());
const GuestcallInfo gci = { GuestUnpacker, GuestTarget };
// Try first with shared_lock
{
std::shared_lock lk(ThunkHandler->ThunksMutex);
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
return found->second;
}
}
std::lock_guard lk(ThunkHandler->ThunksMutex);
// Retry lookup with full lock before making a new trampoline to avoid double trampolines
{
auto found = ThunkHandler->GuestcallToHostTrampoline.find(gci);
if (found != ThunkHandler->GuestcallToHostTrampoline.end()) {
return found->second;
}
}
LogMan::Msg::DFmt("Thunks: Adding host trampoline for guest function {:#x} via unpacker {:#x}",
GuestTarget, GuestUnpacker);
if (ThunkHandler->HostTrampolineInstanceDataAvailable < HostToGuestTrampolineSize) {
const auto allocation_step = 16 * 1024;
ThunkHandler->HostTrampolineInstanceDataAvailable = allocation_step;
ThunkHandler->HostTrampolineInstanceDataPtr = (uint8_t *)mmap(
0, ThunkHandler->HostTrampolineInstanceDataAvailable,
PROT_READ | PROT_WRITE | PROT_EXEC,
MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
LOGMAN_THROW_AA_FMT(ThunkHandler->HostTrampolineInstanceDataPtr != MAP_FAILED, "Failed to mmap HostTrampolineInstanceDataPtr");
}
auto HostTrampoline = reinterpret_cast<HostToGuestTrampolinePtr* const>(ThunkHandler->HostTrampolineInstanceDataPtr);
ThunkHandler->HostTrampolineInstanceDataAvailable -= HostToGuestTrampolineSize;
ThunkHandler->HostTrampolineInstanceDataPtr += HostToGuestTrampolineSize;
memcpy(HostTrampoline, (void*)&HostToGuestTrampolineTemplate, HostToGuestTrampolineSize);
GetInstanceInfo(HostTrampoline) = TrampolineInstanceInfo {
.HostPacker = HostPacker,
.CallCallback = (uintptr_t)&ThunkHandler_impl::CallCallback,
.GuestUnpacker = GuestUnpacker,
.GuestTarget = GuestTarget
};
ThunkHandler->GuestcallToHostTrampoline[gci] = HostTrampoline;
return HostTrampoline;
}
FEX_DEFAULT_VISIBILITY
void FinalizeHostTrampolineForGuestFunction(HostToGuestTrampolinePtr* TrampolineAddress, void* HostPacker) {
if (TrampolineAddress == nullptr) return;
auto& Trampoline = GetInstanceInfo(TrampolineAddress);
LOGMAN_THROW_A_FMT(Trampoline.CallCallback == (uintptr_t)&ThunkHandler_impl::CallCallback,
"Invalid trampoline at {} passed to {}", fmt::ptr(TrampolineAddress), __FUNCTION__);
if (!Trampoline.HostPacker) {
LogMan::Msg::DFmt("Thunks: Finalizing trampoline at {} with host packer {}", fmt::ptr(TrampolineAddress), fmt::ptr(HostPacker));
Trampoline.HostPacker = HostPacker;
}
}
}
+6
View File
@@ -6,6 +6,10 @@ $end_info$
#pragma once
#include <FEXCore/IR/IR.h>
#include <vector>
namespace FEXCore::Context {
struct Context;
}
@@ -28,5 +32,7 @@ namespace FEXCore {
virtual ~ThunkHandler() { }
static ThunkHandler* Create();
virtual void AppendThunkDefinitions(std::vector<FEXCore::IR::ThunkDefinition> const& Definitions) = 0;
};
};
+3 -3
View File
@@ -127,7 +127,7 @@ namespace FEXCore::IR {
auto Array = (AOTIRInlineIndex *)((char*)FilePtr + IndexOffset);
LOGMAN_THROW_A_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
LOGMAN_THROW_AA_FMT(Entry->Array == nullptr && Entry->FilePtr == nullptr, "Entry must not be initialized here");
Entry->Array = Array;
Entry->FilePtr = FilePtr;
Entry->Size = Size;
@@ -387,7 +387,7 @@ namespace FEXCore::IR {
auto Inserted = AOTIRCache.insert({fileid, AOTIRCacheEntry { .FileId = fileid, .Filename = filename }});
auto Entry = &(Inserted.first->second);
LOGMAN_THROW_A_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
LOGMAN_THROW_AA_FMT(Entry->Array == nullptr, "Duplicate LoadAOTIRCacheEntry");
if (CTX->Config.AOTIRLoad && AOTIRLoader) {
auto streamfd = AOTIRLoader(fileid);
@@ -403,7 +403,7 @@ namespace FEXCore::IR {
}
void AOTIRCaptureCache::UnloadAOTIRCacheEntry(AOTIRCacheEntry *Entry) {
LOGMAN_THROW_A_FMT(Entry != nullptr, "Removing not existing entry");
LOGMAN_THROW_AA_FMT(Entry != nullptr, "Removing not existing entry");
if (Entry->Array) {
FEXCore::Allocator::munmap(Entry->FilePtr, Entry->Size);
+48 -94
View File
@@ -123,12 +123,12 @@
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_UXTW {1}",
"constexpr FEXCore::IR::MemOffsetType MEM_OFFSET_SXTW {2}",
"constexpr FEXCore::IR::BreakReason Break_Unimplemented {0}",
"constexpr FEXCore::IR::BreakReason Break_Interrupt {1}",
"constexpr FEXCore::IR::BreakReason Break_Interrupt3 {2}",
"constexpr FEXCore::IR::BreakReason Break_Halt {3}",
"constexpr FEXCore::IR::BreakReason Break_Overflow {4}",
"constexpr FEXCore::IR::BreakReason Break_InvalidInstruction {5}"
"struct BreakDefinition {",
" uint16_t ErrorRegister;",
" uint8_t Signal;",
" uint8_t TrapNumber;",
" uint8_t si_code;",
"};"
],
"IRTypes" : {
"i1": "bool",
@@ -150,7 +150,7 @@
"SyscallFlags": "FEXCore::IR::SyscallFlags",
"SHA256Sum": "SHA256Sum",
"MemOffsetType": "MemOffsetType",
"BreakReason": "BreakReason",
"BreakDefinition": "BreakDefinition",
"RoundType": "RoundType"
},
"Ops": {
@@ -192,7 +192,7 @@
"DestSize": "8"
},
"RemoveThreadCodeEntry": {
"ThreadRemoveCodeEntry": {
"HasSideEffects": true
},
@@ -261,7 +261,7 @@
"HasSideEffects": true,
"DestSize": "GetOpSize(_NewRIP)"
},
"Break BreakReason:$Reason, u8:$Literal": {
"Break BreakDefinition:$Reason": {
"HasSideEffects": true
},
"SignalReturn": {
@@ -302,10 +302,6 @@
}
},
"Moves": {
"GPR = Mov GPR:$Value": {
"DestSize": "GetOpSize(_Value)"
},
"GPR = ExtractElementPair GPRPair:$Pair, u8:$Element": {
"Desc": ["Extracts a register for the register pair"],
"DestSize": "GetOpSize(_Pair) >> 1"
@@ -359,22 +355,26 @@
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContext to XMM\""
]
},
"StoreContext u8:#ByteSize, RegisterClass:$Class, SSA:$Value, u32:$Offset": {
"Desc": ["Stores a value to the context with offset",
"Ctx[Offset] = Value",
"Zero Extends if value's type is too small",
"Truncates if value's type is too large"
],
"Desc": ["Stores a value to the context with offset",
"Ctx[Offset] = Value",
"Zero Extends if value's type is too small",
"Truncates if value's type is too large"
],
"HasSideEffects": true,
"DestSize": "ByteSize",
"EmitValidation": [
"WalkFindRegClass($Value) == $Class",
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContext to XMM\""
]
},
@@ -385,7 +385,9 @@
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContextIndexed to XMM\""
]
},
"StoreContextIndexed SSA:$Value, GPR:$Index, u8:#ByteSize, u32:$BaseOffset, u32:$Stride, RegisterClass:$Class": {
@@ -397,7 +399,9 @@
"EmitValidation": [
"WalkFindRegClass($Value) == $Class",
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContextIndexed to XMM\""
]
},
@@ -475,22 +479,6 @@
]
},
"FPR = VLoadMemElement u8:#RegisterSize, u8:#ElementSize, FPR:$Value, GPR:$Addr, u8:$Index, u8:$Align{1}": {
"Desc": ["Loads an element of size #ElementSize in to $Value from $Addr at $Index"
],
"OpClass": "Memory",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"VStoreMemElement u8:#RegisterSize, u8:#ElementSize, FPR:$Value, GPR:$Addr, u8:$Index, u8:$Align": {
"Desc": ["Stores an element of size #ElementSize from $Value[$Index] to $Addr"
],
"HasSideEffects": true,
"DestSize": "ElementSize",
"NumElements": "RegisterSize / ElementSize"
},
"CacheLineClear GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline clear at the address specified",
"Only clears the data cachelines. Doesn't do any zeroing"
@@ -845,27 +833,25 @@
"DestSize": "std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(_TrueVal), GetOpSize(_FalseVal)))"
},
"GPR = Extr GPR:$Upper, GPR:$Lower, u8:$LSB": {
"Desc": ["Concats the two GPRs to create a value that is the size of the full two GPRs",
"It then extracts a bitfield width that size of a GPR from the LSB",
"Valid LSB range is 0-31 for 32bit and 0-63 for 64bit",
"<Size * 2> ConcatValue = $Upper:$Lower",
"Result = ConcatValue<LSB+Size - 1: LSB>"
]
"Desc": ["Concats the two GPRs to create a value that is the size of the full two GPRs",
"It then extracts a bitfield width that size of a GPR from the LSB",
"Valid LSB range is 0-31 for 32bit and 0-63 for 64bit",
"<Size * 2> ConcatValue = $Upper:$Lower",
"Result = ConcatValue<LSB+Size - 1: LSB>"
]
},
"GPR = PDep GPR:$Input, GPR:$Mask": {
"Desc": [
"Performs a parallel bit deposit.",
"Takes the contiguous low-order bits and deposits them into",
"the destination at the locations specified by the Mask."
]
"Desc": ["Performs a parallel bit deposit.",
"Takes the contiguous low-order bits and deposits them into",
"the destination at the locations specified by the Mask."
]
},
"GPR = PExt GPR:$Input, GPR:$Mask": {
"Desc": [
"Performs a parallel bit extract.",
"Each bit set in the mask will select the corresponding bit in the Input",
"and transfers them to the lower contiguous bits in the destination."
]
"Desc": ["Performs a parallel bit extract.",
"Each bit set in the mask will select the corresponding bit in the Input",
"and transfers them to the lower contiguous bits in the destination."
]
},
"GPR = LDiv GPR:$Lower, GPR:$Upper, GPR:$Divisor": {
@@ -924,15 +910,6 @@
}
},
"Vector": {
"FPR = SplatVector2 FPR:$Scalar": {
"NumElements": "2",
"DestSize": "GetOpSize(_Scalar) * 2"
},
"FPR = SplatVector4 FPR:$Scalar": {
"NumElements": "4",
"DestSize": "GetOpSize(_Scalar) * 4"
},
"FPR = VMov u8:#RegisterSize, FPR:$Source": {
"Desc" : ["Copy vector register",
"When Register size is smaller than Source register size,",
@@ -941,12 +918,6 @@
"DestSize": "RegisterSize"
},
"FPR = VBitcast u8:#RegisterSize, u8:#ElementSize, FPR:$Source": {
"Desc": ["Workaround for issue with LLVM breaking when loading scalar elements to vectors"],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VectorZero u8:#RegisterSize": {
"Desc": ["Generates a vector zero",
"Useful to generate a zero vector without any previous dependencies"
@@ -1038,24 +1009,11 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VExtractElement u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$Index": {
"DestSize": "ElementSize"
},
"FPR = VDupElement u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$Index": {
"Desc": ["Duplicates one element from the source register across the whole register"],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VSLI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$ByteShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VSRI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$ByteShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VShlI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
@@ -1089,7 +1047,7 @@
},
"FPR = VSXTL2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"Desc": ["Sign extends elements from the source element size to the next size up",
"Source elements come from the upper 64bits of the register"
"Source elements come from the upper half of the register"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
@@ -1101,7 +1059,7 @@
},
"FPR = VUXTL2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"Desc": ["Zero extends elements from the source element size to the next size up",
"Source elements come from the upper 64bits of the register"
"Source elements come from the upper half of the register"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
@@ -1124,9 +1082,9 @@
},
"FPR = VRev64 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"Desc" : ["Reverses elements in 64-bit halfwords",
"Available element size: 1byte, 2 byte, 4 byte"
],
"Desc" : ["Reverses elements in 64-bit halfwords",
"Available element size: 1byte, 2 byte, 4 byte"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
@@ -1320,10 +1278,6 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VInsScalarElement u8:#RegisterSize, u8:#ElementSize, u8:$DestIdx, FPR:$DestVector, FPR:$SrcScalar": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VInsGPR u8:#RegisterSize, u8:#ElementSize, u8:$DestIdx, FPR:$DestVector, GPR:$Src": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
@@ -1587,9 +1541,9 @@
"DestSize": "16"
},
"GPR = F80Cmp FPR:$X80Src1, FPR:$X80Src2, u32:$Flags": {
"Desc": ["Does a scalar unordered compare and stores the asked for flags in to a GPR",
"Ordering flag result is true if either float input is NaN"
],
"Desc": ["Does a scalar unordered compare and stores the asked for flags in to a GPR",
"Ordering flag result is true if either float input is NaN"
],
"DestSize": "4"
},
"FPR = F80BCDLoad FPR:$X80Src": {
+6 -1
View File
@@ -182,7 +182,12 @@ static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const*
}
}
static void PrintArg(std::stringstream *out, [[maybe_unused]] IRListView const* IR, FEXCore::IR::BreakDefinition Arg) {
*out << "{" << Arg.ErrorRegister << ".";
*out << static_cast<uint32_t>(Arg.Signal) << ".";
*out << static_cast<uint32_t>(Arg.TrapNumber) << ".";
*out << static_cast<uint32_t>(Arg.si_code) << "}";
}
void Dump(std::stringstream *out, IRListView const* IR, IR::RegisterAllocationData *RAData) {
auto HeaderOp = IR->GetHeader();
+4 -2
View File
@@ -8,6 +8,7 @@ $end_info$
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <array>
@@ -87,7 +88,8 @@ FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(OrderedNode *Node) {
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled op type: {} {} in argument class validation", IROp->Op, GetOpName(Node));
LOGMAN_MSG_A_FMT("Unhandled op type: {} {} in argument class validation",
ToUnderlying(IROp->Op), GetOpName(Node));
break;
}
return InvalidClass;
@@ -167,7 +169,7 @@ IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode
if (insertAfter) {
LinkCodeBlocks(insertAfter, CodeNode);
} else {
LOGMAN_THROW_A_FMT(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
LOGMAN_THROW_AA_FMT(CurrentCodeBlock != nullptr, "CurrentCodeBlock must not be null here");
// Find last block
auto LastBlock = CurrentCodeBlock;
Loaded 100 of 509 files, more files were not shown because too many files have changed in this diff. Show more