Compare commits

...
204 Commits
Author SHA1 Message Date
Ryan Houdek 2f5ebf1dd1 Docs: Update for release FEX-2212 2022-12-05 14:02:57 -08:00
Ryan Houdek 8f157e45bb Merge pull request #2198 from lioncash/add
OpcodeDispatcher: Handle VADDPD/VADDPS/VPADDB/VPADDW/VPADDD/VPADDQ
2022-12-05 13:19:30 -08:00
lioncash 6b259e2731 OpcodeDispatcher: Handle VPADDQ 2022-12-05 17:53:33 +00:00
lioncash 200660aba5 OpcodeDispatcher: Handle VPADDD 2022-12-05 17:44:12 +00:00
lioncash 065c12cfbb OpcodeDispatcher: Handle VPADDW 2022-12-05 17:33:52 +00:00
lioncash 318972620f OpcodeDispatcher: Handle VPADDB 2022-12-05 17:23:39 +00:00
lioncash e7f54d1592 OpcodeDispatcher: Handle VADDPD 2022-12-05 16:59:04 +00:00
lioncash a8571282b2 OpcodeDispatcher: Handle VADDPS 2022-12-05 16:43:26 +00:00
Ryan Houdek fc28062052 Merge pull request #2193 from Sonicadvance1/support_radeon_ioctl_emu
IoctlEmu: Support radeon
2022-12-05 06:36:44 -08:00
Mai cc6306aa32 Merge pull request #2194 from Sonicadvance1/fix_confusing_error
FEXServerClient: Disable confusing connection log
2022-12-05 13:39:50 +00:00
Ryan Houdek f0caa81253 FEXServerClient: Disable confusing connection log
On first FEXInterpreter execution, it is expected that `ConnectToServer`
will fail with `ECONNREFUSED` because FEXServer won't be running.

Skip printing this first messaage to stderr if configured.
If it is some other error message then ensure it is still printed.
2022-12-05 05:18:41 -08:00
Ryan Houdek cd98871f8f IoctlEmu: Support radeon 2022-12-05 05:12:04 -08:00
Mai 66e0d46d89 Merge pull request #2196 from Sonicadvance1/optimize_symlink_following
Linux: Improve performance of hot paths in path searching
2022-12-05 13:05:22 +00:00
Mai 1dd5642e46 Merge pull request #2195 from Sonicadvance1/improve_interpreter_check
FEXLoader: Make `IsInterpreterInstalled` check less horrible.
2022-12-05 13:04:11 +00:00
Mai 9ca34ca306 Merge pull request #2178 from Sonicadvance1/const_jit
Arm64: Const on unmodified argument
2022-12-05 13:03:21 +00:00
Ryan Houdek 69f39a0bc3 Arm64: Const on unmodified argument
Just noticed this when tinkering around the JIT.
These arguments can safely be const.
2022-12-05 03:54:17 -08:00
Ryan Houdek 50cf74db0a Linux: Improve performance of hot paths in path searching
`GetEmulatedPath` and `OpenAt` are called a /lot/ in applications.
std::filesystem::path handling here is quite heavy and costly for what
we are trying to achieve.

Remove this usage and instead use lstat, access, and readlink directly
which is a heck of a lot faster.

In particular this helps out pressure-vessel, shaving off launch times
by 1-2 seconds.
Going from ~22 seconds down to ~20 seconds.
2022-12-05 03:51:50 -08:00
Ryan Houdek 6886d8ff65 FEXLoader: Make IsInterpreterInstalled check less horrible.
`std::filesystem::exists` is particularly gnarly in how it checks to see
if the file exists.
It allocates memory, it creates lists, it splits things, then eventually
checking the status

Remove all this overhead to help out minorly for applications that
execve a lot.
2022-12-05 02:58:15 -08:00
Mai c1d118c1d4 Merge pull request #2191 from Sonicadvance1/minor_aeskeygenassist_optimization
Arm64: Minor optimization in AESKEYGENASSIST
2022-12-04 06:59:31 +00:00
Ryan Houdek 5e46d63c42 Arm64: Minor optimization in AESKEYGENASSIST
The less number of FPR<->GPR movement instructions the better.
This removes one instance of `ins` and replaces the other with a 64-bit
`dup` instead.
The LoadConstant still turns in to a single `movz` instruction with the
shift.
2022-12-03 03:59:27 -08:00
xianwei zheng 863a59a8e2 Thunk: Crash on XSetErrorHandler(NULL) (#2190)
* BUGFIX:Adding the nullptr check to FinalizeHostTrampolineForGuestFunction. It will lead crash on interpreter code set function nullptr callback. eg. XSetErrorHandler(NULL)

* fix #2189
2022-11-30 21:57:09 -08:00
Ryan Houdek c37fcf136a Merge pull request #2187 from lioncash/and
OpcodeDispatcher: Handle VANDPD/VANDPS/VPAND/VANDNPD/VANDNPS/VPANDN
2022-11-30 16:22:42 -08:00
lioncash 02a2292115 OpcodeDispatcher: Handle VPANDN 2022-11-30 15:51:10 +00:00
lioncash a483bc9837 OpcodeDispatcher: Handle VANDNPD 2022-11-30 15:51:10 +00:00
lioncash 120a6b85f4 OpcodeDispatcher: Handle VANDNPS 2022-11-30 15:51:05 +00:00
Ryan Houdek 9fea774a93 Merge pull request #2188 from lioncash/pclmul
unittests: Expand VPCLMULQDQ unit test
2022-11-29 17:15:34 -08:00
lioncash bf1e619ead unittests: Expand vpclmulqdq unit test
Now that we have some AVX instructions in place, we can make the test
use them and also enforce correctness behavior in the upper lane.
2022-11-29 22:07:38 +00:00
lioncash 0f8fcfc43e OpcodeDispatcher: Handle VPAND 2022-11-29 19:08:58 +00:00
lioncash 698b7fda06 OpcodeDispatcher: Handle VANDPD 2022-11-29 19:06:25 +00:00
lioncash 23caa6e20f OpcodeDispatcher: Handle VANDPS 2022-11-29 19:04:29 +00:00
Ryan Houdek 34e39c996e Merge pull request #2186 from lioncash/or
OpcodeDispatcher: Handle VORPD/VORPS/VPOR
2022-11-29 10:57:17 -08:00
lioncash 16ed20cfae OpcodeDispatcher: Handle VPOR 2022-11-29 18:43:38 +00:00
lioncash ef368ceafa OpcodeDispatcher: Handle VORPD 2022-11-29 18:40:04 +00:00
lioncash 45480ef32c OpcodeDispatcher: Handle VORPS 2022-11-29 18:38:13 +00:00
Ryan Houdek 4de69029e5 Merge pull request #2185 from lioncash/xor
OpcodeDispatcher: Handle VPXOR/VXORPD/VXORPS
2022-11-29 10:28:02 -08:00
lioncash c065770f48 OpcodeDispatcher: Handle VPXOR 2022-11-29 18:12:04 +00:00
lioncash 27957ea051 OpcodeDispatcher: Handle VXORPD 2022-11-29 18:05:28 +00:00
lioncash 94e9d1ab3b OpcodeDispatcher: Handle VXORPS 2022-11-29 18:04:29 +00:00
Ryan Houdek a374a9af35 Merge pull request #2184 from lioncash/vzero
OpcodeDispatcher: Handle VZEROUPPER/VZEROALL
2022-11-29 08:33:07 -08:00
lioncash 3e80416eb6 OpcodeDispatcher: Handle VZEROUPPER/VZEROALL 2022-11-29 16:15:32 +00:00
Ryan Houdek b35c6c6d22 Merge pull request #2183 from lioncash/vmovq
OpcodeDispatcher: Handle VMOVQ
2022-11-28 18:33:26 -08:00
lioncash 69d26cfee6 OpcodeDispatcher: Handle combined VMOVQ/VMOVD 2022-11-29 01:49:58 +00:00
lioncash e3be1540f1 OpcodeDispatcher: Handle VMOVQ
Fairly trivial, we can reuse the existing implementation for MOVQ.
2022-11-29 01:49:15 +00:00
Ryan Houdek 8e2b0d10e5 Merge pull request #2181 from Sonicadvance1/defer_cpuinfo
EmulatedFiles: Defer cpuinfo file initialization to first access
2022-11-28 17:48:42 -08:00
Ryan Houdek 57c5761920 Merge pull request #2179 from Sonicadvance1/tsl_maps
Core: Replace a couple maps with tsl robin_map
2022-11-28 17:48:27 -08:00
Ryan Houdek 0841ff5feb Merge pull request #2182 from lioncash/ntdq
OpcodeDispatcher: Handle VMOVNTDQ/VMOVNTDQA/VMOVNTPD/VMOVNTPS
2022-11-28 17:47:27 -08:00
lioncash 45115384b5 OpcodeDispatcher: Handle VMOVNTPD 2022-11-28 16:58:37 +00:00
lioncash dae2563850 OpcodeDispatcher: Handle VMOVNTPS 2022-11-28 16:56:00 +00:00
lioncash fb4df5a0b7 OpcodeDispatcher: Handle VMOVNTDQA 2022-11-28 16:51:06 +00:00
lioncash b2f0303d1e OpcodeDispatcher: Handle VMOVNTDQ 2022-11-28 16:49:01 +00:00
Ryan Houdek f8b2a0b4d8 Merge pull request #2180 from Sonicadvance1/disable_aot_stores
FEXLoader: Disables some AOT shutdown overhead when not enabled
2022-11-28 01:08:39 -08:00
Ryan Houdek e1fcb78ce3 FEXLoader: Disables some AOT shutdown overhead when not enabled
When AOT wasn't enabled it was still doing some accesses to the
filesystem on shutdown.

Slightly improves shutdown time.
2022-11-27 16:02:58 -08:00
Ryan Houdek d6309088c0 EmulatedFiles: Defer cpuinfo file initialization to first access
This improves startup time by a couple of milliseconds.

Most applications don't query cpuinfo so deferring improves most
application's startup times.
2022-11-27 15:34:10 -08:00
Ryan Houdek 1fb3a2e28f Core: Replace a couple maps with tsl robin_map
Improves the shutdown time a small amount and performance of these maps.
2022-11-27 04:35:27 -08:00
Mai df25d4e03e Merge pull request #2177 from Sonicadvance1/disable_multiblock_default
Config: Disable multiblock by default
2022-11-26 06:23:24 +00:00
Ryan Houdek 9c2f0287e0 Config: Disable multiblock by default
This causes users pain currently since the JIT isn't doing any caching
and our RA being a hack mess means it is quite slow.

Disable by default to improve JIT time performance.
2022-11-25 18:08:25 -08:00
Mai 9912d41714 Merge pull request #2174 from Sonicadvance1/emulated_sgdt
OpcodeDispatcher: Implement SGDT
2022-11-25 05:23:56 +00:00
Ryan Houdek 8b2cd87d9e unittests: Disable SGDT tests on host
The Zen+ CI runner doesn't support the UMIP hardware feature, so it
doesn't hit the kernel emulated path.

Instead the instruction returns real data on this hardware. Still in
kernel space, so it is unmapped as expected.
2022-11-24 18:29:05 -08:00
Ryan Houdek 3e6d23ae7e unittests: SGDT tests 2022-11-24 17:47:31 -08:00
Ryan Houdek c96c39d5b1 OpcodeDispatcher: Implement SGDT
Ran in to this when running Team Sonic Racing.
This game uses Denuvo Anti-Tamper which some versions rely on SGDT.

The Linux kernel catches and emulates this instruction.
It will always return a limit of 0 and a base of `0xFFFFFFFFFFFE0000ULL`
which is guaranteed to be in kernel space.

This gets the game slightly farther but still not running entirely.
2022-11-24 17:43:52 -08:00
Mai a4556e90cd Merge pull request #2170 from Sonicadvance1/optimize_sra_step_one
OpcodeDispatcher: Moves all GPR and XMM accesses to direct register accesses
2022-11-24 06:17:55 +00:00
Mai 4d2c4b4423 Merge pull request #2172 from Sonicadvance1/minor_vector_initialization_improvement
Syscalls: Minor optimization with initialization of syscall definition vector
2022-11-23 22:16:36 +00:00
Ryan Houdek 530de3f031 Syscalls: Minor optimization with initialization of syscall definition vector
Shaves a couple milliseconds off initialization time.

Noticed this while poking around.
2022-11-23 12:59:08 -08:00
Ryan Houdek c9622f6fd4 unittests/IR: Update tests for new IR semantics 2022-11-22 23:06:18 -08:00
Ryan Houdek 854628d959 IREmitter relies on CoreState.h now 2022-11-22 23:06:18 -08:00
Ryan Houdek 35dbf6c44b OpcodeDispatcher: Moves all GPR and XMM accesses to direct register accesses
This is a bit of a large commit since it is an all or nothing sort of
change.

Instead of doing LoadContext and StoreContext for GPRs and FPRs, any
that are statically allocated (GPR and XMM) should use
{Load,Store}Register directly.

This is now enforced that {Load,Store}Context{,Indexed} will assert if
trying to access these ranges.

This is the first step towards accelerating our JIT compile times in
that it removes the need for the SRA pass to convert all the IR over.

This ensures that from the Dispatcher directly we are handling SRA
accesses as "registers", ensuring that future changes don't need to run
in to this problem.

Even though it wasn't a target of this change, in a simple test
application this removes ~12% of the total JIT compile time.
2022-11-22 23:06:18 -08:00
Ryan Houdek b9fec7436f X86Tables: Fixes some incorrectly defined instruction sizes.
These were working because of partial loadstore handling with context
loadstores
2022-11-22 21:06:58 -08:00
Ryan Houdek 808d19c374 IR: Disallow loading and storing registers through {Load,Store}Context{,Indexed}
All SRA allocated registers must be handled explicitly through {Load,Store}Register
2022-11-22 21:06:58 -08:00
Ryan Houdek 441d7205ed Passes: Remove StaticRegisterAllocation pass 2022-11-22 21:06:58 -08:00
Ryan Houdek 8b3d3b68c6 Jit64: Implement {Load,Store}Register 2022-11-22 21:06:58 -08:00
Ryan Houdek aa0e038ef7 Interpreter: Implement {Load,Store}Register 2022-11-22 21:06:58 -08:00
Ryan Houdek 5e5e5a35d9 Merge pull request #2157 from Sonicadvance1/more_systemd_stuff
FEXServer: More Systemd fixes
2022-11-22 03:17:17 -08:00
Ryan Houdek 10e35a55ea FEXServerClient: Cleanup AF_UNIX abstract socket string math
Makes it a bit easier to reason about.
2022-11-22 03:05:33 -08:00
Ryan Houdek 83cebea780 FEXServerClient: Clean up comments in server mount folder. 2022-11-22 02:56:31 -08:00
Ryan Houdek 91bbb92c50 FEXServer: More Systemd fixes
Two changes here.

- Make the mount path follow server temp folder requirements.
  - Will be mounted in `/tmp/` or `$XDG_RUNTIME_DIR/` now
- Switch the FEXServer socket to an "abstract" AF_UNIX socket.
  - If the socket is in `/tmp/` then systemd will put the service in a
    private `/tmp` folder that only exists for the service.
  - If the socket is in `$XDG_RUNTIME_DIR` then pressure-vessel can't
    chroot anymore since they make their own runtime directory.
  - If it is in `$HOME/.fex-emu/` then it breaks usage where the
    filesystem is a mount that doesn't support AF_UNIX like sshfs.

The only reasonable thing to do is to switch over to `abstract` sockets
which will work in all cases.
Tested with pressure-vessel and systemd and now it works in all
situations.
2022-11-22 02:55:08 -08:00
Ryan Houdek c7dd6ff28a Merge pull request #2167 from Sonicadvance1/optimize_break_codegen
Arm64: Optimize Break IR op codegen
2022-11-21 23:37:39 -08:00
Ryan Houdek df761a99ce Merge pull request #2171 from lioncash/dqa
OpcodeDispatcher: Handle VMOVDQA/VMOVDQU
2022-11-21 23:29:21 -08:00
Ryan Houdek 359416e2b6 Arm64: Optimize Break IR op codegen
This isn't really a performance issue, more just something that is gross
looking at while looking at unit tests.

Break /usually/ isn't abused heavily by games (Except Denuvo) so not
really a perf concern regardless.
2022-11-21 23:20:15 -08:00
lioncash e140c0d60c OpcodeDispatcher: Handle VMOVDQU 2022-11-22 06:47:55 +00:00
Ryan Houdek 0030971f6f Merge pull request #2168 from Sonicadvance1/cmake_typo
CMake: Fix typo in clang thunks option.
2022-11-21 22:46:52 -08:00
lioncash 3a90aaf1e6 OpcodeDispatcher: Implement VMOVDQA 2022-11-22 06:40:07 +00:00
Ryan Houdek f8a199af49 Merge pull request #2169 from lioncash/movddup
OpcodeDispatcher: Handle VMOVDDUP
2022-11-21 22:21:35 -08:00
lioncash 3a5de8e10c OpcodeDispatcher: Handle VMOVDDUP 2022-11-22 05:33:53 +00:00
Ryan Houdek 3c881809f4 Merge pull request #2166 from Sonicadvance1/wrapnode_helper
IntrusiveIRList: Add a utility helper for getting an OrderedNodeWrapper
2022-11-21 21:33:30 -08:00
Ryan Houdek 8b6e9e08c0 Merge pull request #2165 from Sonicadvance1/remove_migrate_log
Core: Removes log about migrating to shared memory mode
2022-11-21 21:33:11 -08:00
Ryan Houdek 3c8da3e3b4 Merge pull request #2164 from Sonicadvance1/debug_logs_on_bad_socket
FEXServerClient: Add some debug logs for when FEX can't connect to se…
2022-11-21 21:33:00 -08:00
Ryan Houdek 0a39d909b2 CMake: Fix typo in clang thunks option. 2022-11-21 21:11:01 -08:00
Ryan Houdek d35d1092a4 IntrusiveIRList: Add a utility helper for getting an OrderedNodeWrapper
This is a nice helper that was otherwise missing.
2022-11-21 21:02:39 -08:00
Ryan Houdek 7ae655c56a Core: Removes log about migrating to shared memory mode
This hasn't ever caused problems and instead just adds a message that
appears in logs for nearly every application invocation.

Remove it because it isn't necessary to track.
2022-11-21 21:00:10 -08:00
Ryan Houdek 58f35ba413 Merge pull request #2163 from lioncash/movshdup
OpcodeDispatcher: Handle VMOVSHDUP/VMOVSLDUP
2022-11-21 20:58:49 -08:00
Ryan Houdek e61eb24ec2 FEXServerClient: Add some debug logs for when FEX can't connect to server
Sometimes when the socket fails to connect we have no debug information
at all as to why.

This at least gives us a little bit more.
2022-11-21 20:58:36 -08:00
lioncash 2e93d2ce51 OpcodeDispatcher: Simplify SSE MOVSLDUP
Like with MOVSHDUP, we only need to duplicate two values rather than
four.
2022-11-22 04:38:39 +00:00
lioncash 9d21e1efd5 OpcodeDispatcher: Handle VMOVSLDUP 2022-11-22 04:37:48 +00:00
lioncash e0e6b3ad6b OpcodeDispatcher: Simplify SSE MOVSHDUP
We only need to insert two values, rather than four.
2022-11-22 04:16:55 +00:00
lioncash 815cdc5b3c OpcodeDispatcher: Handle VMOVSHDUP 2022-11-22 04:16:34 +00:00
Ryan Houdek bc20f1e684 Merge pull request #2162 from lioncash/vmovhps
OpcodeDispatcher: Handle VMOVHPD/VMOVHPS
2022-11-21 18:02:06 -08:00
lioncash 20e5f2bec6 OpcodeDispatcher: Handle VMOVHPD 2022-11-22 01:06:47 +00:00
lioncash c3b6fa55b6 OpcodeDispatcher: Handle VMOVHPS 2022-11-22 01:01:28 +00:00
Ryan Houdek 69045db3a9 Merge pull request #2161 from lioncash/vmovlps
OpcodeDispatcher: Handle VMOVLPD/VMOVLPS
2022-11-21 14:20:35 -08:00
lioncash 7b2240c80b OpcodeDispatcher: Handle VMOVLPD 2022-11-21 21:45:58 +00:00
lioncash dcfbd90dd7 OpcodeDispatcher: Handle VMOVLPS 2022-11-21 21:45:36 +00:00
Ryan Houdek 70e6ab5782 Merge pull request #2160 from lioncash/mov
Arm64/VectorOps: Simplify VMov IR op on SVE
2022-11-21 12:40:02 -08:00
lioncash 175879823f Arm64/VectorOps: Add future clarifying context comment
Just so this doesn't get overlooked in the distant future.
2022-11-21 20:19:03 +00:00
Ryan Houdek 2271a90adb Merge pull request #2159 from lioncash/vmovapd
OpcodeDispatcher: Handle VMOVAPD/VMOVUPD/VMOVUPS
2022-11-21 12:08:55 -08:00
lioncash 2bdde5845e Arm64/VectorOps: Simplify VMov IR op on SVE
Initially I put this in very conservatively to make sure we always clear
out the upper lanes, but since Adv. SIMD operations have zero-extending
behavior when storing results, we can just use a lot of operations as
is, without needing to unnecessarily do the same work twice.
2022-11-21 20:05:58 +00:00
lioncash 0c40497a01 OpcodeDispatcher: Handle VMOVUPD 2022-11-21 17:14:25 +00:00
lioncash 45dff0f550 OpcodeDecoder: Handle VMOVUPS 2022-11-21 17:06:31 +00:00
lioncash e9035ef6ee OpcodeDecoder: Handle VMOVAPD 2022-11-21 17:06:27 +00:00
Ryan Houdek 02ca94e6e6 Merge pull request #2158 from Sonicadvance1/steam_appid_configs
Config: Add support for steamid based configurations.
2022-11-18 14:40:27 -08:00
Ryan Houdek 432b7d2dc8 Config: Add support for steamid based configurations.
This will be useful for keying specific executables to steamids.
This is sadly required because a bunch of games end up naming themselves
"game.exe" so we can't safely enable thunks for all things shipping a
generic name.
2022-11-17 18:27:42 -08:00
Ryan Houdek 181d315d2c Merge pull request #2155 from lioncash/x86
x86_64/VectorOps: Separate 128-bit/256-bit paths
2022-11-16 19:23:53 -08:00
Mai d5f7e616eb Merge pull request #2156 from Sonicadvance1/fexserver_systemd_fixes
Systemd fixes
2022-11-17 02:20:48 +00:00
Ryan Houdek 9a8869e8a4 FEXServer: Send shutdown signal to image mount programs on exit. 2022-11-16 18:08:32 -08:00
Ryan Houdek 85d56ed76f FEXServerClient: Print a message when the server socket fails 2022-11-16 18:08:32 -08:00
Ryan Houdek 06c827b5c8 FEXServerClient: Support XDG_RUNTIME_DIR
In the case of a platform enabling PrivateTmp then the FEXServer and
FEXInterpreter won't have a tmp folder that shares the socket location.

The runtime directory is a perfect place to share these across
processes.
2022-11-16 18:08:02 -08:00
Ryan Houdek 41259ff361 FEXServer: Don't deparent when running as a systemd service.
Due to how systemd watches PIDs, it will terminate the entire process
tree if the first pid exits.
2022-11-16 18:08:00 -08:00
Ryan Houdek cccee1a668 FEXServer: Ensure fatal messages are printed 2022-11-16 18:08:00 -08:00
lioncash c46b35362b x86_64/VectorOps: Separate 128-bit VShlI path 2022-11-16 19:46:12 +00:00
lioncash 2f96a6d8bf x86_64/VectorOps: Separate 128-bit VSShrI path 2022-11-16 19:44:46 +00:00
lioncash c78a47a3e5 x86_64/VectorOps: Separate 128-bit VUShrI path 2022-11-16 19:40:50 +00:00
lioncash b049721683 x86_64/VectorOps: Separate 128-bit VUABDL path 2022-11-16 19:37:01 +00:00
lioncash 11c06fe9fe x86_64/VectorOps: Separate 128-bit VMul path 2022-11-16 19:32:04 +00:00
lioncash c5f6e53d0d x86_64/VectorOps: Separate 128-bit VSShrS path 2022-11-16 19:29:13 +00:00
lioncash 1184672bb1 x86_64/VectorOps: Separate 128-bit VUShrS path 2022-11-16 19:27:47 +00:00
lioncash 57d3a2ba35 x86_64/VectorOps: Separate 128-bit VUShlS path 2022-11-16 19:26:20 +00:00
lioncash 1d813b0183 x86_64/VectorOps: Separate 128-bit VFCMPUNO path 2022-11-16 19:22:48 +00:00
lioncash 9b24931518 x86_64/VectorOps: Separate 128-bit VFCMPORD path 2022-11-16 19:22:01 +00:00
lioncash 7e5b8b7bdf x86_64/VectorOps: Separate 128-bit VFCMPLE path 2022-11-16 19:20:56 +00:00
lioncash 7e36473aff x86_64/VectorOps: Separate 128-bit VFCMPGT path 2022-11-16 19:19:28 +00:00
lioncash adbb512306 x86_64/VectorOps: Separate 128-bit VFCMPLT path 2022-11-16 19:18:01 +00:00
lioncash b951f4ad4b x86_64/VectorOps: Separate 128-bit VFCMPNEQ path 2022-11-16 19:16:57 +00:00
lioncash 48900662ae x86_64/VectorOps: Separate 128-bit VFCMPEQ path 2022-11-16 19:15:48 +00:00
lioncash da31a66c07 x86_64/VectorOps: Separate 128-bit VCMPLTZ path 2022-11-16 19:13:00 +00:00
lioncash ca710b1cbb x86_64/VectorOps: Separate 128-bit VCMPGTZ path 2022-11-16 19:10:39 +00:00
lioncash a4d7eec145 x86_64/VectorOps: Separate 128-bit VCMPGT path 2022-11-16 19:08:30 +00:00
lioncash 232c2fe87f x86_64/VectorOps: Separate 128-bit VCMPEQZ path 2022-11-16 19:06:54 +00:00
lioncash 661112cfd4 x86_64/VectorOps: Separate 128-bit VCMPEQ path 2022-11-16 19:01:31 +00:00
lioncash d4a84eaa9a x86_64/VectorOps: Separate 128-bit VBSL path 2022-11-16 18:59:40 +00:00
lioncash 99f8af64d0 x86_64/VectorOps: Separate 128-bit VSMax path 2022-11-16 18:57:16 +00:00
lioncash 1541ea9ffc x86_64/VectorOps: Separate 128-bit VUMax path 2022-11-16 18:55:23 +00:00
lioncash ced86e693c x86_64/VectorOps: Separate 128-bit VSMin path 2022-11-16 18:53:46 +00:00
lioncash 53920d5bd3 x86_64/VectorOps: Separate 128-bit VUMin path 2022-11-16 18:51:56 +00:00
lioncash 15c5a9dac0 x86_64/VectorOps: Separate 128-bit VNot path 2022-11-16 18:49:13 +00:00
lioncash d63cfdbe7d x86_64/VectorOps: Separate 128-bit VFNeg path 2022-11-16 18:46:35 +00:00
lioncash 569461a01c x86_64/VectorOps: Separate 128-bit VNeg path 2022-11-16 18:41:16 +00:00
lioncash 37c7dee236 x86_64/VectorOps: Separate 128-bit VFRSqrt path 2022-11-16 18:35:43 +00:00
lioncash cc230091c8 x86_64/VectorOps: Separate 128-bit VFSqrt path 2022-11-16 18:33:14 +00:00
lioncash bac33cf246 x86_64/VectorOps: Separate 128-bit VFRecp path 2022-11-16 18:31:38 +00:00
lioncash 7a9c0506b4 x86_64/VectorOps: Separate 128-bit VFMax path 2022-11-16 18:28:31 +00:00
lioncash fbd7c15a4b x86_64/VectorOps: Separate 128-bit VFMin path 2022-11-16 18:27:00 +00:00
lioncash 64edf24bc7 x86_64/VectorOps: Separate 128-bit VFDiv path 2022-11-16 18:25:00 +00:00
lioncash 3f456d683f x86_64/VectorOps: Separate 128-bit VFMul path 2022-11-16 18:23:21 +00:00
lioncash aac0824fd3 x86_64/VectorOps: Separate 128-bit VFSub path 2022-11-16 18:21:29 +00:00
lioncash 1b32ca0b93 x86_64/VectorOps: Separate 128-bit VFAddP path 2022-11-16 18:19:13 +00:00
lioncash c29456aac9 x86_64/VectorOps: Separate 128-bit VFAdd path 2022-11-16 18:17:23 +00:00
lioncash 715f25d059 x86_64/VectorOps: Separate 128-bit VAbs path 2022-11-16 18:13:51 +00:00
lioncash c85e31ec7a x86_64/VectorOps: Separate 128-bit VURAvg path 2022-11-16 18:10:53 +00:00
lioncash 0d6bfc4fa4 x86_64/VectorOps: Separate 128-bit VSQSub path 2022-11-16 18:07:59 +00:00
lioncash 106917add2 x86_64/VectorOps: Separate 128-bit VSQAdd path 2022-11-16 18:06:17 +00:00
lioncash 1e10a2bac3 x86_64/VectorOps: Separate 128-bit VUQSub path 2022-11-16 18:02:09 +00:00
lioncash 63dd09cd6f x86_64/VectorOps: Separate 128-bit VUQAdd path 2022-11-16 17:56:37 +00:00
lioncash c7e6935f42 x86_64/VectorOps: Separate 128-bit VSub path 2022-11-16 17:54:57 +00:00
lioncash 1be5054e86 x86_64/VectorOps: Separate 128-bit VAdd path 2022-11-16 17:51:57 +00:00
lioncash f40755aca7 x86_64/VectorOps: Separate 128-bit VXor path 2022-11-16 17:49:25 +00:00
lioncash d49b78cf34 x86_64/VectorOps: Separate 128-bit VOr path 2022-11-16 17:46:49 +00:00
lioncash 10e80ae064 x86_64/VectorOps: Separate 128-bit VBic path 2022-11-16 17:44:10 +00:00
lioncash f3ebb214aa x86_64/VectorOps: Separate 128-bit VAnd path 2022-11-16 17:40:25 +00:00
lioncash dc41c3bd9f x86_64/VectorOps: Separate 128-bit VectorImm path 2022-11-16 17:40:14 +00:00
Ryan Houdek 56ff09f3ac Merge pull request #2154 from lioncash/decoder
OpcodeDispatcher: Handle VMOVAPS
2022-11-15 12:10:41 -08:00
lioncash 1707b27d14 OpcodeDecoder: Only install AVX ops if host supports it 2022-11-15 19:57:08 +00:00
lioncash 7c92963eca x86_64/MemoryOps: Use unaligned loads/stores for 256-bit cases
Alignment will always be a little finicky in this case, so let's just
use unaligned variants to make behavior always consistent.
2022-11-15 19:57:08 +00:00
lioncash ecf82c90ee OpcodeDispatcher: Handle VMOVAPS 2022-11-15 19:57:05 +00:00
lioncash 580f06fe00 Frontend: Handle 256-bit vectors 2022-11-15 03:11:34 +00:00
Ryan Houdek 5a403b7765 Merge pull request #2153 from lioncash/extr
IR: Handle 256-bit VExtr
2022-11-14 15:00:12 -08:00
lioncash f7367e56af IR: Handle 256-bit VExtr
Extends VExtr to handle 256-bit vectors.
2022-11-14 22:26:52 +00:00
Ryan Houdek f066abc151 Merge pull request #2152 from lioncash/vixl
Externals: Update vixl submodule
2022-11-14 09:49:36 -08:00
lioncash 2bee92f332 Externals: Update vixl submodule
Incorporates an upstreamed fix that allows movprfx to be used with
destructive SVE EXT.
2022-11-14 15:45:11 +00:00
Mai 7d9ed4e1bf Merge pull request #2151 from Sonicadvance1/remove_VSLI_VSRI
IR: Removes the only uses of VSLI and VSRI
2022-11-14 05:15:27 +00:00
Ryan Houdek 24696e6b98 IR: Removes the only uses of VSLI and VSRI
We use these two IR ops for PSRLQ and PSLLQ respectively, Which is
actually implemented in a quite inefficient way.

Instead switch the ops over to using VExtr which maps directly to one
instruction and can emulate both.
2022-11-13 19:51:31 -08:00
Mai 71f658b07d Merge pull request #2150 from Sonicadvance1/sort_named_rootfs
FEXConfig: Sort named rootfs vector
2022-11-13 23:14:38 +00:00
Ryan Houdek 920f56353f FEXConfig: Sort named rootfs vector
This was driving me nuts that this wasn't sorted by name.
2022-11-13 14:36:47 -08:00
Ryan Houdek e0fe9167ea Merge pull request #2149 from Sonicadvance1/calculate_minstack_size
ELFCodeLoader: Calculate AT_MINSIGSTKSZ
2022-11-13 14:26:59 -08:00
Ryan Houdek 45a2349a0d ELFCodeLoader: Calculate AT_MINSIGSTKSZ
This is the last remaining auxv value that we were missing.

This requires a little bit of setup to match what we are doing in
FEXCore's Dispatcher.

This only tracks how much space is required by the kernel to store its
required state.
2022-11-13 14:09:51 -08:00
Ryan Houdek d7b0e8469e Merge pull request #2148 from Sonicadvance1/fix_null_end
ELFCodeLoader: Fixes AT_PLATFORM null terminator
2022-11-13 13:58:47 -08:00
Mai 8afc3b8e23 Merge pull request #2147 from Sonicadvance1/auxv_secure
ELFCodeLoader: Pass through AT_SECURE
2022-11-13 21:49:27 +00:00
Ryan Houdek 9caa63d5b2 ELFCodeLoader: Fixes AT_PLATFORM null terminator
We were failing to null terminate the platform which is causing strcmp
to fail.
2022-11-13 13:46:36 -08:00
Mai 1d32df91ce Merge pull request #2146 from Sonicadvance1/fix_vsyscall_auxv
ELFCodeLoader: Ensure we set AT_SYSINFO for 32-bit
2022-11-13 20:52:46 +00:00
Ryan Houdek fa973f65bf ELFCodeLoader: Pass through AT_SECURE
When we are installed as a binfmt_misc handler the Linux kernel will
assign us AT_SECURE for setuid binaries.

Ensure we are passing this through for any application that will end up
needing it.
2022-11-13 12:50:58 -08:00
Ryan Houdek c027f02e5a ELFCodeLoader: Ensure we set AT_SYSINFO for 32-bit
AT_SYSINFO points to the vsyscall location that the kernel provides.
This AUXV value isn't used on x86-64.
If we have VDSO installed then we can use the one provided from there,
otherwise we need to provide a code page.

It seems like some behaviour changed with glibc provided in Ubuntu 22.10
that it now requires AT_SYSINFO.

Fixes wine 7.0 execution inside the Ubuntu 22.10 rootfs.
2022-11-13 12:39:07 -08:00
Ryan Houdek 9cee0126d7 Merge pull request #2145 from lioncash/memload
IR: Remove VLoadMemElement and VStoreMemElement
2022-11-11 11:44:54 -08:00
lioncash c713602d56 IR: Remove VLoadMemElement and VStoreMemElement
These are unused, so we can get rid of them.
2022-11-11 19:06:38 +00:00
Mai c8293cbbda Merge pull request #2137 from Sonicadvance1/pclmul_disable
OpcodeDispatcher: Disable PCLMUL if not supported on host
2022-11-10 22:23:48 +00:00
Ryan Houdek a9c51388cf Merge pull request #2144 from lioncash/reg
IR: Handle 256-bit LoadRegister/StoreRegister
2022-11-08 21:31:51 -08:00
lioncash ad1d65e91a IR: Handle 256-bit LoadRegister
Extends LoadRegister to handle 256-bit vectors.
2022-11-09 03:19:38 +00:00
lioncash 3a6c7803e8 IR: Handle 256-bit StoreRegister
Extends StoreRegister to handle 256-bit vectors.
2022-11-09 03:19:34 +00:00
Ryan Houdek aa837eddd5 Merge pull request #2143 from Sonicadvance1/kinetic_support
InstallFEX.py: Adds support for Kinetic
2022-11-08 14:38:52 -08:00
Ryan Houdek faef57838f InstallFEX.py: Adds support for Kinetic
Also removes Hirsute and Impish since Ubuntu's PPA system doesn't
support these anymore.

Fixes #2140
2022-11-08 14:26:59 -08:00
Ryan Houdek 04d4c5e017 Merge pull request #2141 from lioncash/addv
IR: Handle 256-bit VAddV
2022-11-08 12:14:47 -08:00
Ryan Houdek 9158877569 Merge pull request #2139 from neobrain/fix_thunks_guest_ide_integration
Thunks: Fix guest targets not being detected by IDEs
2022-11-08 12:09:44 -08:00
lioncash 1eae07f1b8 IR: Handle 256-bit VAddV
Extends VAddV to handle 256-bit vectors.
2022-11-08 18:21:30 +00:00
Tony Wasserka 14b22487f1 Thunks: Fix guest targets not being detected by IDEs
The IDE integration path didn't set up the BITNESS variable. This change
also unmarks that variable as a CMake option since it's not a boolean value.
2022-11-08 15:30:49 +01:00
Ryan Houdek 1ac7cd5835 OpcodeDispatcher: Disable PCLMUL if not supported on host
Might fix Steam on Pi4
2022-11-05 12:11:45 -07:00
Mai 5336f01725 Merge pull request #2135 from Sonicadvance1/release_process
Update release process to include AUR
2022-11-03 19:38:19 +00:00
Ryan Houdek b8b66b1829 Update release process to include AUR 2022-11-03 01:23:32 -07:00
131 changed files with 5371 additions and 1596 deletions

No files matched your search

+1 -1
View File
@@ -7,7 +7,7 @@ CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
option(BUILD_TESTS "Build unit tests to ensure sanity" TRUE)
option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(BUILD_CLANG_THUNKS "Build thunks with clang" FALSE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
-1
View File
@@ -135,7 +135,6 @@ set (SRCS
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/StaticRegisterAllocationPass.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/SyscallOptimization.cpp
Utils/Allocator.cpp
+10 -6
View File
@@ -211,10 +211,12 @@ namespace JSON {
static std::map<FEXCore::Config::LayerType, std::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static FEXCore::Config::Layer *Meta{};
constexpr std::array<FEXCore::Config::LayerType, 7> LoadOrder = {
constexpr std::array<FEXCore::Config::LayerType, 9> LoadOrder = {
FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN,
FEXCore::Config::LayerType::LAYER_MAIN,
FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_GLOBAL_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_APP,
FEXCore::Config::LayerType::LAYER_ARGUMENTS,
FEXCore::Config::LayerType::LAYER_ENVIRONMENT,
@@ -629,7 +631,7 @@ namespace JSON {
class AppLoader final : public FEXCore::Config::OptionMapper {
public:
explicit AppLoader(const std::string& Filename, bool Global);
explicit AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type);
void Load();
private:
@@ -681,8 +683,10 @@ namespace JSON {
});
}
AppLoader::AppLoader(const std::string& Filename, bool Global)
: FEXCore::Config::OptionMapper(Global ? FEXCore::Config::LayerType::LAYER_GLOBAL_APP : FEXCore::Config::LayerType::LAYER_LOCAL_APP) {
AppLoader::AppLoader(const std::string& Filename, FEXCore::Config::LayerType Type)
: FEXCore::Config::OptionMapper(Type) {
const bool Global = Type == FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP ||
Type == FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP;
Config = FEXCore::Config::GetApplicationConfig(Filename, Global);
// Immediately load so we can reload the meta layer
@@ -754,8 +758,8 @@ namespace JSON {
}
}
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, bool Global) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Global);
std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, FEXCore::Config::LayerType Type) {
return std::make_unique<FEXCore::Config::AppLoader>(Filename, Type);
}
std::unique_ptr<FEXCore::Config::Layer> CreateEnvironmentLayer(char *const _envp[]) {
+3 -2
View File
@@ -16,10 +16,11 @@
},
"Multiblock": {
"Type": "bool",
"Default": "true",
"Default": "false",
"ShortArg": "m",
"Desc": [
"Controls multiblock code compilation"
"Controls multiblock code compilation",
"Can cause long JIT compilation times and stutter"
]
},
"MaxInst": {
+1 -1
View File
@@ -421,7 +421,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) {
Res.ecx =
(1 << 0) | // SSE3
(1 << 1) | // PCLMULQDQ
(CTX->HostFeatures.SupportsPMULL_128Bit << 1) | // PCLMULQDQ
(1 << 2) | // DS area supports 64bit layout
(1 << 3) | // MWait
(0 << 4) | // DS-CPL
-2
View File
@@ -1212,8 +1212,6 @@ namespace FEXCore::Context {
IsMemoryShared = true;
if (Config.TSOAutoMigration) {
LogMan::Msg::IFmt("Migrating to shared memory mode");
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
LogMan::Throw::AFmt(Threads.size() == 1, "First MarkMemoryShared called must be before creating any threads");
+17 -3
View File
@@ -451,8 +451,13 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DestSize = 2;
}
else if (DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_128BIT);
DestSize = 16;
if (Options.L) {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_256BIT);
DestSize = 32;
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeDstSize(DecodeFlags::SIZE_128BIT);
DestSize = 16;
}
}
else if (HasNarrowingDisplacement &&
(DstSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_DEF ||
@@ -483,7 +488,14 @@ bool Decoder::NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op,
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_16BIT);
}
else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_128BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
if (Options.L) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
} else {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_128BIT);
}
}
else if (SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_256BIT) {
DecodeInst->Flags |= DecodeFlags::GenSizeSrcSize(DecodeFlags::SIZE_256BIT);
}
else if (HasNarrowingDisplacement &&
(SrcSizeFlag == FEXCore::X86Tables::InstFlags::SIZE_DEF ||
@@ -777,6 +789,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
if (Op == 0xC5) { // Two byte VEX
pp = Byte1 & 0b11;
options.vvvv = 15 - ((Byte1 & 0b01111000) >> 3);
options.L = (Byte1 & 0b100) != 0;
}
else { // 0xC4 = Three byte VEX
const uint8_t Byte2 = ReadByte();
@@ -784,6 +797,7 @@ bool Decoder::NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16
map_select = Byte1 & 0b11111;
options.vvvv = 15 - ((Byte2 & 0b01111000) >> 3);
options.w = (Byte2 & 0b10000000) != 0;
options.L = (Byte2 & 0b100) != 0;
if ((Byte1 & 0b01000000) == 0) {
LOGMAN_THROW_A_FMT(CTX->Config.Is64BitMode, "VEX.X shouldn't be 0 in 32-bit mode!");
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_X;
+1
View File
@@ -49,6 +49,7 @@ private:
struct DecodedHeader {
uint8_t vvvv; // Encoded operand in a VEX prefix.
bool w; // VEX.W bit.
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Context::Context *CTX;
@@ -65,6 +65,7 @@ HostFeatures::HostFeatures() {
SupportsFlushInputsToZero = Features.Has(vixl::CPUFeatures::Feature::kAFP);
SupportsRCPC = Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsTSOImm9 = Features.Has(vixl::CPUFeatures::Feature::kRCpcImm);
SupportsPMULL_128Bit = Features.Has(vixl::CPUFeatures::Feature::kPmull1Q);
Supports3DNow = true;
SupportsSSE4A = true;
@@ -127,6 +128,7 @@ HostFeatures::HostFeatures() {
SupportsSHA = Features.has(Xbyak::util::Cpu::tSHA);
SupportsBMI1 = Features.has(Xbyak::util::Cpu::tBMI1);
SupportsBMI2 = Features.has(Xbyak::util::Cpu::tBMI2);
SupportsPMULL_128Bit = Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
// xbyak doesn't know how to check for CLZero
uint32_t eax, ebx, ecx, edx;
@@ -146,7 +146,7 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info) {
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80LOADFCW: {
@@ -154,8 +154,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
@@ -245,8 +243,6 @@ constexpr OpHandlerArray InterpreterOpHandlers = [] {
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
@@ -49,7 +49,7 @@ namespace FEXCore::CPU {
public:
static void InterpretIR(FEXCore::Core::CpuStateFrame *Frame, FEXCore::IR::IRListView const *IR);
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header *IROp, FallbackInfo *Info);
static bool GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info);
struct IROpData {
FEXCore::Core::InternalThreadState *State{};
@@ -181,8 +181,6 @@ namespace FEXCore::CPU {
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -265,8 +263,6 @@ namespace FEXCore::CPU {
DEF_OP(VInsElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
@@ -69,11 +69,47 @@ DEF_OP(StoreContext) {
}
DEF_OP(LoadRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Src = ContextPtr + Op->Offset;
#define LOAD_CTX(x, y) \
case x: { \
y const *MemData = reinterpret_cast<y const*>(Src); \
GD = *MemData; \
break; \
}
switch (OpSize) {
LOAD_CTX(1, uint8_t)
LOAD_CTX(2, uint16_t)
LOAD_CTX(4, uint32_t)
LOAD_CTX(8, uint64_t)
case 16:
case 32: {
void const *MemData = reinterpret_cast<void const*>(Src);
memcpy(GDP, MemData, OpSize);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContext size: {}", OpSize);
break;
}
#undef LOAD_CTX
}
DEF_OP(StoreRegister) {
LOGMAN_MSG_A_FMT("Unimplemented");
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
const auto ContextPtr = reinterpret_cast<uintptr_t>(Data->State->CurrentFrame);
const auto Dst = ContextPtr + Op->Offset;
void *MemData = reinterpret_cast<void*>(Dst);
void *Src = GetSrc<void*>(Data->SSAData, Op->Value);
memcpy(MemData, Src, OpSize);
}
DEF_OP(LoadContextIndexed) {
@@ -236,36 +272,6 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(VLoadMemElement) {
auto Op = IROp->C<IR::IROp_VLoadMemElement>();
void const *MemData = *GetSrc<void const**>(Data->SSAData, Op->Value);
memcpy(GDP, GetSrc<void*>(Data->SSAData, Op->Addr), 16);
memcpy(reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(GDP) + (Op->Header.ElementSize * Op->Index)),
MemData, Op->Header.ElementSize);
}
DEF_OP(VStoreMemElement) {
#define STORE_DATA(x, y) \
case x: { \
y *MemData = *GetSrc<y**>(Data->SSAData, Op->Value); \
memcpy(MemData, &GetSrc<y*>(Data->SSAData, Op->Addr)[Op->Index], sizeof(y)); \
break; \
}
auto Op = IROp->C<IR::IROp_VStoreMemElement>();
uint8_t OpSize = IROp->Size;
switch (OpSize) {
STORE_DATA(1, uint8_t)
STORE_DATA(2, uint16_t)
STORE_DATA(4, uint32_t)
STORE_DATA(8, uint64_t)
default: LOGMAN_MSG_A_FMT("Unhandled StoreMem size"); break;
}
#undef STORE_DATA
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -11,6 +11,7 @@ $end_info$
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/BitUtils.h>
#include <array>
#include <bit>
#include <cstdint>
#include <limits>
@@ -358,23 +359,26 @@ DEF_OP(VAddP) {
}
DEF_OP(VAddV) {
auto Op = IROp->C<IR::IROp_VAddV>();
const uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VAddV>();
const auto OpSize = IROp->Size;
void *Src = GetSrc<void*>(Data->SSAData, Op->Vector);
uint8_t Tmp[16];
uint8_t Tmp[Core::CPUState::XMM_AVX_REG_SIZE];
const uint8_t Elements = OpSize / Op->Header.ElementSize;
const uint8_t ElementSize = Op->Header.ElementSize;
const uint8_t Elements = OpSize / ElementSize;
const auto Func = [](auto current, auto a) { return current + a; };
switch (Op->Header.ElementSize) {
switch (ElementSize) {
DO_VECTOR_REDUCE_1SRC_OP(1, int8_t, Func, 0)
DO_VECTOR_REDUCE_1SRC_OP(2, int16_t, Func, 0)
DO_VECTOR_REDUCE_1SRC_OP(4, int32_t, Func, 0)
DO_VECTOR_REDUCE_1SRC_OP(8, int64_t, Func, 0)
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
return;
}
memcpy(GDP, Tmp, Op->Header.ElementSize);
memcpy(GDP, Tmp, ElementSize);
}
DEF_OP(VUMinV) {
@@ -1619,41 +1623,52 @@ DEF_OP(VDupElement) {
}
DEF_OP(VExtr) {
auto Op = IROp->C<IR::IROp_VExtr>();
const uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VExtr>();
const auto OpSize = IROp->Size;
const auto OpSizeBits = OpSize * 8;
const auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->VectorLower);
const auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->VectorUpper);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = Op->Header.ElementSize;
const auto Index = Op->Index;
uint64_t Offset = Op->Index * Op->Header.ElementSize * 8;
__uint128_t Dst{};
if (Offset >= (OpSize * 8)) {
Offset -= OpSize * 8;
Dst = Src1 >> Offset;
if (Is256Bit) {
const auto ByteIndex = Index * ElementSize;
const auto IsUpperVectorZero = ByteIndex >= OpSize;
const auto SanitizedByteIndex = IsUpperVectorZero ? ByteIndex - OpSize
: ByteIndex;
const auto Vectors = IsUpperVectorZero
?
std::array<InterpVector256, 2>{
*GetSrc<InterpVector256*>(Data->SSAData, Op->VectorLower),
InterpVector256{},
}
:
std::array<InterpVector256, 2>{
*GetSrc<InterpVector256*>(Data->SSAData, Op->VectorUpper),
*GetSrc<InterpVector256*>(Data->SSAData, Op->VectorLower),
};
const auto* VectorsPtr = reinterpret_cast<const uint8_t*>(Vectors.data());
const auto* SrcPtr = VectorsPtr + SanitizedByteIndex;
memcpy(GDP, SrcPtr, OpSize);
} else {
uint64_t Offset = Index * ElementSize * 8;
const auto Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->VectorLower);
const auto Src2 = *GetSrc<__uint128_t*>(Data->SSAData, Op->VectorUpper);
__uint128_t Dst{};
if (Offset >= OpSizeBits) {
Offset -= OpSizeBits;
Dst = Src1 >> Offset;
} else {
Dst = (Src1 << (OpSizeBits - Offset)) | (Src2 >> Offset);
}
memcpy(GDP, &Dst, OpSize);
}
else {
Dst = (Src1 << (OpSize * 8 - Offset)) | (Src2 >> Offset);
}
memcpy(GDP, &Dst, OpSize);
}
DEF_OP(VSLI) {
auto Op = IROp->C<IR::IROp_VSLI>();
const __uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Vector);
const __uint128_t Src2 = Op->ByteShift * 8;
const __uint128_t Dst = Op->ByteShift >= sizeof(__uint128_t) ? 0 : Src1 << Src2;
memcpy(GDP, &Dst, 16);
}
DEF_OP(VSRI) {
auto Op = IROp->C<IR::IROp_VSRI>();
const __uint128_t Src1 = *GetSrc<__uint128_t*>(Data->SSAData, Op->Vector);
const __uint128_t Src2 = Op->ByteShift * 8;
const __uint128_t Dst = Op->ByteShift >= sizeof(__uint128_t) ? 0 : Src1 >> Src2;
memcpy(GDP, &Dst, 16);
}
DEF_OP(VUShrI) {
@@ -14,7 +14,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(TruncElementPair) {
auto Op = IROp->C<IR::IROp_TruncElementPair>();
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
// Size is the size of each pair element
@@ -19,7 +19,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(SignalReturn) {
// First we must reset the stack
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
@@ -10,7 +10,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(AESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -18,7 +18,7 @@ DEF_OP(AESImc) {
}
DEF_OP(AESEnc) {
auto Op = IROp->C<IR::IROp_VAESEnc>();
auto Op = IROp->C<IR::IROp_VAESEnc>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
@@ -27,7 +27,7 @@ DEF_OP(AESEnc) {
}
DEF_OP(AESEncLast) {
auto Op = IROp->C<IR::IROp_VAESEncLast>();
auto Op = IROp->C<IR::IROp_VAESEncLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aese(VTMP1.V16B(), VTMP2.V16B());
@@ -35,7 +35,7 @@ DEF_OP(AESEncLast) {
}
DEF_OP(AESDec) {
auto Op = IROp->C<IR::IROp_VAESDec>();
auto Op = IROp->C<IR::IROp_VAESDec>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
@@ -44,7 +44,7 @@ DEF_OP(AESDec) {
}
DEF_OP(AESDecLast) {
auto Op = IROp->C<IR::IROp_VAESDecLast>();
auto Op = IROp->C<IR::IROp_VAESDecLast>();
eor(VTMP2.V16B(), VTMP2.V16B(), VTMP2.V16B());
mov(VTMP1.V16B(), GetSrc(Op->State.ID()).V16B());
aesd(VTMP1.V16B(), VTMP2.V16B());
@@ -52,7 +52,7 @@ DEF_OP(AESDecLast) {
}
DEF_OP(AESKeyGenAssist) {
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
auto Op = IROp->C<IR::IROp_VAESKeyGenAssist>();
aarch64::Literal ConstantLiteral (0x0C030609'0306090CULL, 0x040B0E01'0B0E0104ULL);
aarch64::Label PastConstant;
@@ -69,9 +69,8 @@ DEF_OP(AESKeyGenAssist) {
if (Op->RCON) {
tbl(VTMP1.V16B(), VTMP1.V16B(), VTMP3.V16B());
LoadConstant(TMP1.W(), Op->RCON);
ins(VTMP2.V4S(), 1, TMP1.W());
ins(VTMP2.V4S(), 3, TMP1.W());
LoadConstant(TMP1, static_cast<uint64_t>(Op->RCON) << 32);
dup(VTMP2.V2D(), TMP1);
eor(GetDst(Node).V16B(), VTMP1.V16B(), VTMP2.V16B());
}
else {
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(GetReg<RA_64>(Node), GetReg<RA_64>(Op->Value.ID()), Op->Flag, 1);
+2 -2
View File
@@ -80,7 +80,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
void Arm64JITCore::Op_Unhandled(IR::IROp_Header *IROp, IR::NodeID Node) {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -477,7 +477,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
return HostCode;
}
void Arm64JITCore::Op_NoOp(IR::IROp_Header *IROp, IR::NodeID Node) {
void Arm64JITCore::Op_NoOp(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
Arm64JITCore::Arm64JITCore(FEXCore::Context::Context *ctx, FEXCore::Core::InternalThreadState *Thread)
@@ -213,7 +213,7 @@ private:
*/
uint8_t *GuestEntry{};
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header *IROp, IR::NodeID Node);
using OpHandler = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
std::array<OpHandler, IR::IROps::OP_LAST + 1> OpHandlers {};
void RegisterALUHandlers();
void RegisterAtomicHandlers();
@@ -225,7 +225,7 @@ private:
void RegisterMoveHandlers();
void RegisterVectorHandlers();
void RegisterEncryptionHandlers();
#define DEF_OP(x) void Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
///< Unhandled handler
DEF_OP(Unhandled);
@@ -345,8 +345,6 @@ private:
DEF_OP(StoreMemTSO);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -430,8 +428,6 @@ private:
DEF_OP(VInsElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
+297 -84
View File
@@ -13,7 +13,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(LoadContext) {
const auto Op = IROp->C<IR::IROp_LoadContext>();
@@ -127,17 +127,18 @@ DEF_OP(StoreContext) {
DEF_OP(LoadRegister) {
auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
auto regOffs = Op->Offset & 7;
const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.gregs[0])) / Core::CPUState::GPR_REG_SIZE;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
const auto reg = SRA64[regId];
switch(Op->Header.Size) {
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
ubfx(GetReg<RA_64>(Node), reg, regOffs * 8, 8);
@@ -156,57 +157,161 @@ DEF_OP(LoadRegister) {
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
if (GetReg<RA_64>(Node).GetCode() != reg.GetCode())
if (GetReg<RA_64>(Node).GetCode() != reg.GetCode()) {
mov(GetReg<RA_64>(Node), reg);
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = CTX->HostFeatures.SupportsAVX ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
const auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "out of range regId");
auto guest = SRAFPR[regId];
auto host = GetSrc(Node);
const auto guest = SRAFPR[regId];
const auto host = GetSrc(Node);
switch(Op->Header.Size) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
mov(host.B(), guest.B());
break;
if (HostSupportsSVE) {
const auto regOffs = Op->Offset & 31;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
fmov(host.H(), guest.H());
break;
aarch64::Label DataLocation;
const auto LoadPredicate = [this, &DataLocation] {
const auto Predicate = p0;
adr(TMP1, &DataLocation);
ldr(Predicate, SVEMemOperand(TMP1));
return Predicate.Merging();
};
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs");
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode())
fmov(host.S(), guest.S());
} else {
ins(host.V4S(), 0, guest.V4S(), regOffs/4);
using DataLiteral = aarch64::Literal<uint32_t>;
const auto EmitData = [this, &DataLocation](DataLiteral& Data) {
aarch64::Label PastConstant;
b(&PastConstant);
bind(&DataLocation);
place(&Data);
bind(&PastConstant);
};
switch (OpSize) {
case 1: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
mov(host.B(), guest.B());
break;
}
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs");
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode())
mov(host.D(), guest.D());
} else {
ins(host.V2D(), 0, guest.V2D(), regOffs/8);
case 2: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
fmov(host.H(), guest.H());
break;
}
break;
case 4: {
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode()) {
fmov(host.S(), guest.S());
}
} else {
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
if (host.GetCode() != guest.GetCode())
mov(host.Q(), guest.Q());
break;
dup(VTMP1.Z().VnS(), host.Z().VnS(), 0);
mov(guest.Z().VnS(), Predicate, VTMP1.Z().VnS());
EmitData(Data);
}
break;
}
case 8: {
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode()) {
mov(host.D(), guest.D());
}
} else {
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
dup(VTMP1.Z().VnD(), host.Z().VnD(), 0);
mov(guest.Z().VnD(), Predicate, VTMP1.Z().VnD());
EmitData(Data);
}
break;
}
case 16: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (host.GetCode() != guest.GetCode()) {
mov(host.Q(), guest.Q());
}
break;
}
case 32: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (host.GetCode() != guest.GetCode()) {
mov(host.Z().VnD(), PRED_TMP_32B.Merging(), guest.Z().VnD());
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister FPR size: {}", OpSize);
break;
}
} else {
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
mov(host.B(), guest.B());
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
fmov(host.H(), guest.H());
break;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode()) {
fmov(host.S(), guest.S());
}
} else {
ins(host.V4S(), 0, guest.V4S(), regOffs/4);
}
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
if (regOffs == 0) {
if (host.GetCode() != guest.GetCode()) {
mov(host.D(), guest.D());
}
} else {
ins(host.V2D(), 0, guest.V2D(), regOffs/8);
}
break;
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (host.GetCode() != guest.GetCode()) {
mov(host.Q(), guest.Q());
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister FPR size: {}", OpSize);
break;
}
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
@@ -214,17 +319,18 @@ DEF_OP(LoadRegister) {
}
DEF_OP(StoreRegister) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
auto regOffs = Op->Offset & 7;
const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < SRA64.size(), "out of range regId");
auto reg = SRA64[regId];
const auto reg = SRA64[regId];
switch(Op->Header.Size) {
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
bfi(reg, GetReg<RA_64>(Op->Value.ID()), regOffs * 8, 8);
@@ -242,46 +348,163 @@ DEF_OP(StoreRegister) {
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
if (GetReg<RA_64>(Op->Value.ID()).GetCode() != reg.GetCode())
if (GetReg<RA_64>(Op->Value.ID()).GetCode() != reg.GetCode()) {
mov(reg, GetReg<RA_64>(Op->Value.ID()));
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = CTX->HostFeatures.SupportsAVX ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto regSize = HostSupportsSVE ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
const auto regOffs = Op->Offset & 15;
LOGMAN_THROW_A_FMT(regId < SRAFPR.size(), "regId out of range");
auto guest = SRAFPR[regId];
auto host = GetSrc(Op->Value.ID());
const auto guest = SRAFPR[regId];
const auto host = GetSrc(Op->Value.ID());
switch(Op->Header.Size) {
case 1:
ins(guest.V16B(), regOffs, host.V16B(), 0);
break;
if (HostSupportsSVE) {
// 256-bit capable hardware allows us to expand the allowed
// offsets used, however we cannot use Adv. SIMD's INS instruction
// at all, since it will zero out the upper lanes of the 256-bit SVE
// vectors, so we'll need to set up a proper predicate for performing
// the insert.
case 2:
LOGMAN_THROW_AA_FMT((regOffs & 1) == 0, "unexpected regOffs");
ins(guest.V8H(), regOffs/2, host.V8H(), 0);
break;
const auto regOffs = Op->Offset & 31;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs");
ins(guest.V4S(), regOffs/4, host.V4S(), 0);
break;
// Compartmentalized setting up of the predicate for the cases that need it.
aarch64::Label DataLocation;
const auto LoadPredicate = [this, &DataLocation] {
const auto Predicate = p0;
adr(TMP1, &DataLocation);
ldr(Predicate, SVEMemOperand(TMP1));
return Predicate.Merging();
};
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs");
ins(guest.V2D(), regOffs / 8, host.V2D(), 0);
break;
// Emits the predicate data and provides the necessary jump to go around the
// emitted data instead of trying to execute it. Place at end of necessary code.
// It's helpful to treat LoadPredicate and EmitData as a prologue and epilogue
// respectfully.
using DataLiteral = aarch64::Literal<uint32_t>;
const auto EmitData = [this, &DataLocation](DataLiteral& Data) {
aarch64::Label PastConstant;
b(&PastConstant);
bind(&DataLocation);
place(&Data);
bind(&PastConstant);
};
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
if (guest.GetCode() != host.GetCode())
mov(guest.Q(), host.Q());
break;
switch (OpSize) {
case 1: {
LOGMAN_THROW_AA_FMT(regOffs <= 31, "unexpected reg index: {}", regOffs);
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
dup(VTMP1.Z().VnB(), host.Z().VnB(), 0);
mov(guest.Z().VnB(), Predicate, VTMP1.Z().VnB());
EmitData(Data);
break;
}
case 2: {
LOGMAN_THROW_AA_FMT((regOffs / 2) <= 15, "unexpected reg index: {}", regOffs / 2);
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
dup(VTMP1.Z().VnH(), host.Z().VnH(), 0);
mov(guest.Z().VnH(), Predicate, VTMP1.Z().VnH());
EmitData(Data);
break;
}
case 4: {
LOGMAN_THROW_AA_FMT((regOffs / 4) <= 7, "unexpected reg index: {}", regOffs / 4);
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
dup(VTMP1.Z().VnS(), host.Z().VnS(), 0);
mov(guest.Z().VnS(), Predicate, VTMP1.Z().VnS());
EmitData(Data);
break;
}
case 8: {
LOGMAN_THROW_AA_FMT((regOffs / 8) <= 3, "unexpected reg index: {}", regOffs / 8);
const auto Predicate = LoadPredicate();
DataLiteral Data{1U << regOffs};
dup(VTMP1.Z().VnD(), host.Z().VnD(), 0);
mov(guest.Z().VnD(), Predicate, VTMP1.Z().VnD());
EmitData(Data);
break;
}
case 16: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (guest.GetCode() != host.GetCode()) {
mov(guest.Q(), host.Q());
}
break;
}
case 32: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (guest.GetCode() != host.GetCode()) {
mov(guest.Z().VnD(), PRED_TMP_32B.Merging(), host.Z().VnD());
}
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister FPR size: {}", OpSize);
break;
}
} else {
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1:
ins(guest.V16B(), regOffs, host.V16B(), 0);
break;
case 2:
LOGMAN_THROW_AA_FMT((regOffs & 1) == 0, "unexpected regOffs: {}", regOffs);
ins(guest.V8H(), regOffs/2, host.V8H(), 0);
break;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
ins(guest.V4S(), regOffs/4, host.V4S(), 0);
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
ins(guest.V2D(), regOffs / 8, host.V2D(), 0);
break;
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
if (guest.GetCode() != host.GetCode()) {
mov(guest.Q(), host.Q());
}
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister FPR size: {}", OpSize);
break;
}
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
@@ -1167,14 +1390,6 @@ DEF_OP(ParanoidStoreMemTSO) {
}
}
DEF_OP(VLoadMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(VStoreMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -1235,8 +1450,6 @@ void Arm64JITCore::RegisterMemoryHandlers() {
REGISTER_OP(LOADMEMTSO, LoadMemTSO);
REGISTER_OP(STOREMEMTSO, StoreMemTSO);
}
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
+14 -11
View File
@@ -11,7 +11,7 @@ $end_info$
namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
@@ -41,16 +41,19 @@ DEF_OP(Break) {
// First we must reset the stack
ResetStack();
LoadConstant(w1, 1);
strb(w1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException)));
LoadConstant(w1, Op->Reason.Signal);
strb(w1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.Signal)));
LoadConstant(w1, Op->Reason.TrapNumber);
str(w1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.TrapNo)));
LoadConstant(w1, Op->Reason.si_code);
str(w1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.si_code)));
LoadConstant(x1, Op->Reason.ErrorRegister);
str(w1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.err_code)));
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Op->Reason.Signal,
.TrapNo = Op->Reason.TrapNumber,
.si_code = Op->Reason.si_code,
.err_code = Op->Reason.ErrorRegister,
};
uint64_t Constant{};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(x1, Constant);
str(x1, MemOperand(STATE, offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData)));
switch (Op->Reason.Signal) {
case SIGILL:
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
switch (Op->Header.Size) {
+95 -130
View File
@@ -10,7 +10,7 @@ namespace FEXCore::CPU {
using namespace vixl;
using namespace vixl::aarch64;
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
DEF_OP(VectorZero) {
if (HostSupportsSVE) {
const auto Dst = GetDst(Node).Z().VnD();
@@ -82,76 +82,38 @@ DEF_OP(VectorImm) {
}
DEF_OP(VMov) {
auto Op = IROp->C<IR::IROp_VMov>();
const uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VMov>();
const auto OpSize = IROp->Size;
const auto Dst = GetDst(Node);
const auto Source = GetSrc(Op->Source.ID());
switch (OpSize) {
case 1: {
if (HostSupportsSVE) {
eor(VTMP1.Z().VnD(), VTMP1.Z().VnD(), VTMP1.Z().VnD());
} else {
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
}
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
mov(VTMP1.V16B(), 0, Source.V16B(), 0);
if (HostSupportsSVE) {
mov(Dst.Z().VnD(), VTMP1.Z().VnD());
} else {
mov(Dst, VTMP1);
}
mov(Dst, VTMP1);
break;
}
case 2: {
if (HostSupportsSVE) {
eor(VTMP1.Z().VnD(), VTMP1.Z().VnD(), VTMP1.Z().VnD());
} else {
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
}
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
mov(VTMP1.V8H(), 0, Source.V8H(), 0);
if (HostSupportsSVE) {
mov(Dst.Z().VnD(), VTMP1.Z().VnD());
} else {
mov(Dst, VTMP1);
}
mov(Dst, VTMP1);
break;
}
case 4: {
if (HostSupportsSVE) {
eor(VTMP1.Z().VnD(), VTMP1.Z().VnD(), VTMP1.Z().VnD());
} else {
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
}
eor(VTMP1.V16B(), VTMP1.V16B(), VTMP1.V16B());
mov(VTMP1.V4S(), 0, Source.V4S(), 0);
if (HostSupportsSVE) {
mov(Dst.Z().VnD(), VTMP1.Z().VnD());
} else {
mov(Dst, VTMP1);
}
mov(Dst, VTMP1);
break;
}
case 8: {
if (HostSupportsSVE) {
eor(VTMP1.Z().VnD(), VTMP1.Z().VnD(), VTMP1.Z().VnD());
mov(VTMP1.V8B(), Source.V8B());
mov(Dst.Z().VnB(), VTMP1.Z().VnB());
} else {
mov(Dst.V8B(), Source.V8B());
}
mov(Dst.V8B(), Source.V8B());
break;
}
case 16: {
if (HostSupportsSVE) {
eor(VTMP1.Z().VnD(), VTMP1.Z().VnD(), VTMP1.Z().VnD());
mov(VTMP1.V16B(), Source.V16B());
mov(Dst.Z().VnB(), VTMP1.Z().VnB());
mov(Dst.V16B(), Source.V16B());
} else {
if (Dst.GetCode() != Source.GetCode()) {
mov(Dst.V16B(), Source.V16B());
@@ -160,6 +122,10 @@ DEF_OP(VMov) {
break;
}
case 32: {
// NOTE: If, in the distant future we support larger moves, or registers
// (*cough* AVX-512 *cough*) make sure to change this to treat
// 256-bit moves with zero extending behavior instead of doing only
// a regular SVE move into a 512-bit register.
if (Dst.GetCode() != Source.GetCode()) {
mov(Dst.Z().VnD(), Source.Z().VnD());
}
@@ -621,20 +587,70 @@ DEF_OP(VAddP) {
}
DEF_OP(VAddV) {
auto Op = IROp->C<IR::IROp_VAddV>();
const uint8_t OpSize = IROp->Size;
const uint8_t Elements = OpSize / Op->Header.ElementSize;
// Vector
switch (Op->Header.ElementSize) {
case 1:
case 2:
case 4:
addv(GetDst(Node).VCast(Op->Header.ElementSize * 8, 1), GetSrc(Op->Vector.ID()).VCast(OpSize * 8, Elements));
break;
case 8:
addp(GetDst(Node).VCast(OpSize * 8, 1), GetSrc(Op->Vector.ID()).VCast(OpSize * 8, Elements));
break;
default: LOGMAN_MSG_A_FMT("Unknown Element Size: {}", Op->Header.ElementSize); break;
const auto Op = IROp->C<IR::IROp_VAddV>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Elements = OpSize / ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetDst(Node);
const auto Vector = GetSrc(Op->Vector.ID());
if (HostSupportsSVE && Is256Bit) {
// SVE doesn't have an equivalent ADDV instruction, so we make do
// by performing two Adv. SIMD ADDV operations on the high and low
// 128-bit lanes and then sum them up.
const auto Mask = PRED_TMP_32B.Zeroing();
const auto CompactPred = p0;
// Select all our upper elements to run ADDV over them.
not_(CompactPred.VnB(), Mask, PRED_TMP_16B.VnB());
compact(VTMP1.Z().VnD(), CompactPred, Vector.Z().VnD());
switch (ElementSize) {
case 1:
addv(VTMP2.B(), Vector.V16B());
addv(VTMP1.B(), VTMP1.V16B());
add(Dst.V16B(), VTMP1.V16B(), VTMP2.V16B());
break;
case 2:
addv(VTMP2.H(), Vector.V8H());
addv(VTMP1.H(), VTMP1.V8H());
add(Dst.V8H(), VTMP1.V8H(), VTMP2.V8H());
break;
case 4:
addv(VTMP2.S(), Vector.V4S());
addv(VTMP1.S(), VTMP1.V4S());
add(Dst.V4S(), VTMP1.V4S(), VTMP2.V4S());
break;
case 8:
addp(VTMP2.D(), Vector.V2D());
addp(VTMP1.D(), VTMP1.V2D());
add(Dst.V2D(), VTMP1.V2D(), VTMP2.V2D());
break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
} else {
const auto OpSizeBits = OpSize * 8;
const auto ElementSizeBits = ElementSize * 8;
switch (ElementSize) {
case 1:
case 2:
case 4:
addv(Dst.VCast(ElementSizeBits, 1), Vector.VCast(OpSizeBits, Elements));
break;
case 8:
addp(Dst.VCast(OpSizeBits, 1), Vector.VCast(OpSizeBits, Elements));
break;
default:
LOGMAN_MSG_A_FMT("Unknown Element Size: {}", ElementSize);
break;
}
}
}
@@ -3952,12 +3968,16 @@ DEF_OP(VDupElement) {
}
DEF_OP(VExtr) {
auto Op = IROp->C<IR::IROp_VExtr>();
const uint8_t OpSize = IROp->Size;
const auto Op = IROp->C<IR::IROp_VExtr>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
// AArch64 ext op has bit arrangement as [Vm:Vn] so arguments need to be swapped
const auto Dst = GetDst(Node);
auto UpperBits = GetSrc(Op->VectorLower.ID());
auto LowerBits = GetSrc(Op->VectorUpper.ID());
const auto ElementSize = Op->Header.ElementSize;
auto Index = Op->Index;
if (Index >= OpSize) {
@@ -3970,70 +3990,17 @@ DEF_OP(VExtr) {
Index -= OpSize;
}
if (OpSize == 8) {
ext(GetDst(Node).V8B(), LowerBits.V8B(), UpperBits.V8B(), Index * Op->Header.ElementSize);
}
else {
ext(GetDst(Node).V16B(), LowerBits.V16B(), UpperBits.V16B(), Index * Op->Header.ElementSize);
}
}
const auto CopyFromByte = Index * ElementSize;
DEF_OP(VSLI) {
auto Op = IROp->C<IR::IROp_VSLI>();
const uint8_t OpSize = IROp->Size;
const uint8_t BitShift = Op->ByteShift * 8;
if (BitShift < 64) {
// Move to Pair [TMP2:TMP1]
mov(TMP1, GetSrc(Op->Vector.ID()).V2D(), 0);
mov(TMP2, GetSrc(Op->Vector.ID()).V2D(), 1);
// Left shift low 64bits
lsl(TMP3, TMP1, BitShift);
// Extract high 64bits from [TMP2:TMP1]
extr(TMP1, TMP2, TMP1, 64 - BitShift);
mov(GetDst(Node).V2D(), 0, TMP3);
mov(GetDst(Node).V2D(), 1, TMP1);
}
else {
if (Op->ByteShift >= OpSize) {
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
}
else {
mov(TMP1, GetSrc(Op->Vector.ID()).V2D(), 0);
lsl(TMP1, TMP1, BitShift - 64);
mov(GetDst(Node).V2D(), 0, xzr);
mov(GetDst(Node).V2D(), 1, TMP1);
}
}
}
DEF_OP(VSRI) {
auto Op = IROp->C<IR::IROp_VSRI>();
const uint8_t OpSize = IROp->Size;
const uint8_t BitShift = Op->ByteShift * 8;
if (BitShift < 64) {
// Move to Pair [TMP2:TMP1]
mov(TMP1, GetSrc(Op->Vector.ID()).V2D(), 0);
mov(TMP2, GetSrc(Op->Vector.ID()).V2D(), 1);
// Extract Low 64bits [TMP2:TMP2] >> BitShift
extr(TMP1, TMP2, TMP1, BitShift);
// Right shift high bits
lsr(TMP2, TMP2, BitShift);
mov(GetDst(Node).V2D(), 0, TMP1);
mov(GetDst(Node).V2D(), 1, TMP2);
}
else {
if (Op->ByteShift >= OpSize) {
eor(GetDst(Node).V16B(), GetDst(Node).V16B(), GetDst(Node).V16B());
}
else {
mov(TMP1, GetSrc(Op->Vector.ID()).V2D(), 1);
lsr(TMP1, TMP1, BitShift - 64);
mov(GetDst(Node).V2D(), 0, TMP1);
mov(GetDst(Node).V2D(), 1, xzr);
if (HostSupportsSVE && Is256Bit) {
movprfx(VTMP2.Z().VnD(), LowerBits.Z().VnD());
ext(VTMP2.Z().VnB(), VTMP2.Z().VnB(), UpperBits.Z().VnB(), CopyFromByte);
mov(Dst.Z().VnD(), VTMP2.Z().VnD());
} else {
if (OpSize == 8) {
ext(Dst.V8B(), LowerBits.V8B(), UpperBits.V8B(), CopyFromByte);
} else {
ext(Dst.V16B(), LowerBits.V16B(), UpperBits.V16B(), CopyFromByte);
}
}
}
@@ -5304,8 +5271,6 @@ void Arm64JITCore::RegisterVectorHandlers() {
REGISTER_OP(VINSELEMENT, VInsElement);
REGISTER_OP(VDUPELEMENT, VDupElement);
REGISTER_OP(VEXTR, VExtr);
REGISTER_OP(VSLI, VSLI);
REGISTER_OP(VSRI, VSRI);
REGISTER_OP(VUSHRI, VUShrI);
REGISTER_OP(VSSHRI, VSShrI);
REGISTER_OP(VSHLI, VShlI);
@@ -337,6 +337,8 @@ private:
///< Memory ops
DEF_OP(LoadContext);
DEF_OP(StoreContext);
DEF_OP(LoadRegister);
DEF_OP(StoreRegister);
DEF_OP(LoadContextIndexed);
DEF_OP(StoreContextIndexed);
DEF_OP(SpillRegister);
@@ -345,8 +347,6 @@ private:
DEF_OP(StoreFlag);
DEF_OP(LoadMem);
DEF_OP(StoreMem);
DEF_OP(VLoadMemElement);
DEF_OP(VStoreMemElement);
DEF_OP(CacheLineClear);
DEF_OP(CacheLineZero);
@@ -430,8 +430,6 @@ private:
DEF_OP(VInsElement);
DEF_OP(VDupElement);
DEF_OP(VExtr);
DEF_OP(VSLI);
DEF_OP(VSRI);
DEF_OP(VUShrI);
DEF_OP(VSShrI);
DEF_OP(VShlI);
@@ -82,11 +82,7 @@ DEF_OP(LoadContext) {
break;
}
case 32: {
if (Op->Offset % 32 == 0) {
vmovaps(ToYMM(Dst), yword [STATE + Op->Offset]);
} else {
vmovups(ToYMM(Dst), yword [STATE + Op->Offset]);
}
vmovups(ToYMM(Dst), yword [STATE + Op->Offset]);
break;
}
default:
@@ -156,11 +152,7 @@ DEF_OP(StoreContext) {
break;
}
case 32: {
if (Op->Offset % 32 == 0) {
vmovaps(yword [STATE + Op->Offset], ToYMM(Value));
} else {
vmovups(yword [STATE + Op->Offset], ToYMM(Value));
}
vmovups(yword [STATE + Op->Offset], ToYMM(Value));
break;
}
default:
@@ -170,6 +162,120 @@ DEF_OP(StoreContext) {
}
}
DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
switch (OpSize) {
case 1: {
movzx(GetSrc<RA_32>(Node), byte [STATE + Op->Offset]);
break;
}
case 2: {
movzx(GetSrc<RA_32>(Node), word [STATE + Op->Offset]);
break;
}
case 4: {
mov(GetSrc<RA_32>(Node), dword [STATE + Op->Offset]);
break;
}
case 8: {
mov(GetSrc<RA_64>(Node), qword [STATE + Op->Offset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister size: {}", OpSize);
break;
}
}
else {
const auto Dst = GetSrc(Node);
switch (OpSize) {
case 1: {
movzx(rax, byte [STATE + Op->Offset]);
vmovq(Dst, rax);
break;
}
case 2: {
movzx(rax, word [STATE + Op->Offset]);
vmovq(Dst, rax);
break;
}
case 4: {
vmovd(Dst, dword [STATE + Op->Offset]);
break;
}
case 8: {
vmovq(Dst, qword [STATE + Op->Offset]);
break;
}
case 16: {
if (Op->Offset % 16 == 0) {
vmovaps(Dst, xword [STATE + Op->Offset]);
} else {
vmovups(Dst, xword [STATE + Op->Offset]);
}
break;
}
case 32: {
vmovups(ToYMM(Dst), yword [STATE + Op->Offset]);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister size: {}", OpSize);
break;
}
}
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
const auto regOffs = Op->Offset & 7;
switch (OpSize) {
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
mov(dword [STATE + Op->Offset], GetSrc<RA_32>(Op->Value.ID()));
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
mov(qword [STATE + Op->Offset], GetSrc<RA_64>(Op->Value.ID()));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto Value = GetSrc(Op->Value.ID());
switch (OpSize) {
case 16: {
if (Op->Offset % 16 == 0) {
vmovaps(xword [STATE + Op->Offset], Value);
} else {
vmovups(xword [STATE + Op->Offset], Value);
}
break;
}
case 32: {
vmovups(yword [STATE + Op->Offset], ToYMM(Value));
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContext size: {}", OpSize);
break;
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(LoadContextIndexed) {
const auto Op = IROp->C<IR::IROp_LoadContextIndexed>();
const auto OpSize = IROp->Size;
@@ -269,11 +375,7 @@ DEF_OP(LoadContextIndexed) {
}
break;
case 32:
if (Op->BaseOffset % 32 == 0) {
vmovaps(ToYMM(Dst), yword [STATE + rax]);
} else {
vmovups(ToYMM(Dst), yword [STATE + rax]);
}
vmovups(ToYMM(Dst), yword [STATE + rax]);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadContextIndexed size: {}", OpSize);
@@ -369,11 +471,7 @@ DEF_OP(StoreContextIndexed) {
}
break;
case 32:
if (Op->BaseOffset % 32 == 0) {
vmovaps(yword [STATE + rax], ToYMM(Value));
} else {
vmovups(yword [STATE + rax], ToYMM(Value));
}
vmovups(yword [STATE + rax], ToYMM(Value));
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreContextIndexed size: {}", OpSize);
@@ -432,7 +530,7 @@ DEF_OP(SpillRegister) {
break;
}
case 32: {
vmovaps(yword [rsp + SlotOffset], ToYMM(Src));
vmovups(yword [rsp + SlotOffset], ToYMM(Src));
break;
}
default:
@@ -488,7 +586,7 @@ DEF_OP(FillRegister) {
break;
}
case 32: {
vmovaps(ToYMM(Dst), yword [rsp + SlotOffset]);
vmovups(ToYMM(Dst), yword [rsp + SlotOffset]);
break;
}
default:
@@ -668,14 +766,6 @@ DEF_OP(StoreMem) {
}
}
DEF_OP(VLoadMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(VStoreMemElement) {
LOGMAN_MSG_A_FMT("Unimplemented");
}
DEF_OP(CacheLineClear) {
auto Op = IROp->C<IR::IROp_CacheLineClear>();
@@ -706,8 +796,8 @@ void X86JITCore::RegisterMemoryHandlers() {
#define REGISTER_OP(op, x) OpHandlers[FEXCore::IR::IROps::OP_##op] = &X86JITCore::Op_##x
REGISTER_OP(LOADCONTEXT, LoadContext);
REGISTER_OP(STORECONTEXT, StoreContext);
REGISTER_OP(LOADREGISTER, Unhandled); // SRA specific, not supported on this backend
REGISTER_OP(STOREREGISTER, Unhandled);
REGISTER_OP(LOADREGISTER, LoadRegister);
REGISTER_OP(STOREREGISTER, StoreRegister);
REGISTER_OP(LOADCONTEXTINDEXED, LoadContextIndexed);
REGISTER_OP(STORECONTEXTINDEXED, StoreContextIndexed);
REGISTER_OP(SPILLREGISTER, SpillRegister);
@@ -718,8 +808,6 @@ void X86JITCore::RegisterMemoryHandlers() {
REGISTER_OP(STOREMEM, StoreMem);
REGISTER_OP(LOADMEMTSO, LoadMem);
REGISTER_OP(STOREMEMTSO, StoreMem);
REGISTER_OP(VLOADMEMELEMENT, VLoadMemElement);
REGISTER_OP(VSTOREMEMELEMENT, VStoreMemElement);
REGISTER_OP(CACHELINECLEAR, CacheLineClear);
REGISTER_OP(CACHELINEZERO, CacheLineZero);
#undef REGISTER_OP
@@ -50,11 +50,19 @@ DEF_OP(Break) {
add(rsp, SpillSlots * MaxSpillSlotSize);
}
mov(byte [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.FaultToTopAndGeneratedException)], 1);
mov(byte [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.Signal)], Op->Reason.Signal);
mov(dword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.TrapNo)], Op->Reason.TrapNumber);
mov(dword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.err_code)], Op->Reason.ErrorRegister);
mov(dword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData.si_code)], Op->Reason.si_code);
Core::CpuStateFrame::SynchronousFaultDataStruct State = {
.FaultToTopAndGeneratedException = 1,
.Signal = Op->Reason.Signal,
.TrapNo = Op->Reason.TrapNumber,
.si_code = Op->Reason.si_code,
.err_code = Op->Reason.ErrorRegister,
};
uint64_t Constant{};
memcpy(&Constant, &State, sizeof(State));
mov(TMP1, Constant);
mov(qword [STATE + offsetof(FEXCore::Core::CpuStateFrame, SynchronousFaultData)], TMP1);
switch (Op->Reason.Signal) {
case SIGILL:
File diff suppressed because it is too large. Load diff
+9 -8
View File
@@ -8,6 +8,7 @@
#include <utility>
#include <vector>
#include <mutex>
#include <tsl/robin_map.h>
namespace FEXCore {
namespace Context {
@@ -17,7 +18,7 @@ namespace Context {
class LookupCache {
public:
struct LookupCacheEntry {
struct LookupCacheEntry {
uintptr_t HostCode;
uintptr_t GuestCode;
};
@@ -54,7 +55,7 @@ public:
return L1Entry.HostCode;
}
}
// Try L3
auto HostCode = BlockList.find(Address);
@@ -62,7 +63,7 @@ public:
CacheBlockMapping(Address, HostCode->second);
return HostCode->second;
}
// Failed to find
return 0;
}
@@ -73,7 +74,7 @@ public:
// Returns true if new pages are marked as containing code
bool AddBlockExecutableRange(uint64_t Address, uint64_t Start, uint64_t Length) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length -1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
@@ -88,7 +89,7 @@ public:
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void *HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
LOGMAN_THROW_AA_FMT(Inserted, "Duplicate block mapping added");
@@ -158,7 +159,7 @@ public:
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::DebugStore,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
// Some care is taken so that L1 lookups can be done without locks, and even tearing is unlikely to lead to a crash.
@@ -167,7 +168,7 @@ public:
std::recursive_mutex WriteLock;
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Do L1
@@ -239,7 +240,7 @@ private:
std::map<BlockLinkTag, std::function<void()>> BlockLinks;
std::map<uint64_t, uint64_t> BlockList;
tsl::robin_map<uint64_t, uint64_t> BlockList;
constexpr static size_t CODE_SIZE = 128 * 1024 * 1024;
constexpr static size_t SIZE_PER_PAGE = 4096 * sizeof(LookupCacheEntry);
File diff suppressed because it is too large. Load diff
+30 -3
View File
@@ -76,10 +76,10 @@ public:
OrderedNode* flagsOpSrcSigned{};
FEXCore::Context::Context *CTX{};
// Used during new op bringup
bool ShouldDump {false};
struct JumpTargetInfo {
OrderedNode* BlockEntry;
bool HaveEmitted;
@@ -298,6 +298,8 @@ public:
void WriteSegmentReg(OpcodeArgs);
void EnterOp(OpcodeArgs);
void SGDTOp(OpcodeArgs);
// SSE
void MOVAPSOp(OpcodeArgs);
void MOVUPSOp(OpcodeArgs);
@@ -404,6 +406,26 @@ public:
// ADX Ops
void ADXOp(OpcodeArgs);
// AVX Ops
template <IROps IROp, size_t ElementSize>
void AVXVectorALUOp(OpcodeArgs);
void VANDNOp(OpcodeArgs);
void VMOVAPS_VMOVAPD_Op(OpcodeArgs);
void VMOVUPS_VMOVUPD_Op(OpcodeArgs);
void VMOVHPOp(OpcodeArgs);
void VMOVLPOp(OpcodeArgs);
void VMOVDDUPOp(OpcodeArgs);
void VMOVSHDUPOp(OpcodeArgs);
void VMOVSLDUPOp(OpcodeArgs);
void VMOVVectorNTOp(OpcodeArgs);
void VZEROOp(OpcodeArgs);
// X87 Ops
template<size_t width>
void FLD(OpcodeArgs);
@@ -521,7 +543,7 @@ public:
void X87FRSTORF64(OpcodeArgs);
void X87FXAMF64(OpcodeArgs);
void X87LDENVF64(OpcodeArgs);
template<size_t width, bool Integer, FCOMIFlags whichflags, bool poptwice>
void FCOMIF64(OpcodeArgs);
@@ -664,6 +686,11 @@ private:
// Non-temporal streaming
ACCESS_STREAM,
};
OrderedNode *LoadGPRRegister(uint32_t GPR, int8_t Size = -1, uint8_t Offset = 0);
OrderedNode *LoadXMMRegister(uint32_t XMM);
void StoreGPRRegister(uint32_t GPR, OrderedNode *const Src, int8_t Size = -1, uint8_t Offset = 0);
void StoreXMMRegister(uint32_t XMM, OrderedNode *const Src);
OrderedNode *GetRelocatedPC(FEXCore::X86Tables::DecodedOp const& Op, int64_t Offset = 0);
OrderedNode *LoadSource(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
OrderedNode *LoadSource_WithOpSize(FEXCore::IR::RegisterClassType Class, FEXCore::X86Tables::DecodedOp const& Op, FEXCore::X86Tables::DecodedOperand const& Operand, uint8_t OpSize, uint32_t Flags, int8_t Align, bool LoadData = true, bool ForceLoad = false, MemoryAccessType AccessType = MemoryAccessType::ACCESS_DEFAULT);
@@ -221,7 +221,8 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *XMM0 = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm.avx.data[0]));
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
@@ -33,11 +33,46 @@ void OpDispatchBuilder::MOVVectorNTOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::VMOVVectorNTOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1, true, false, MemoryAccessType::ACCESS_STREAM);
// TODO: When stores and loads gain the ability to explicitly express
// whether a vector extension or an insert is desirable, ensure
// the 128-bit case here is a zero extend on store if the destination
// is a register.
StoreResult(FPRClass, Op, Src, 1, MemoryAccessType::ACCESS_STREAM);
}
void OpDispatchBuilder::MOVAPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
StoreResult(FPRClass, Op, Src, -1);
}
void OpDispatchBuilder::VMOVAPS_VMOVAPD_Op(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto Is128BitDest = GetDstSize(Op) == Core::CPUState::XMM_SSE_REG_SIZE;
if (Op->Dest.IsGPR() && Is128BitDest) {
// Perform 32 byte store to clear the upper lane.
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 32, -1);
} else {
StoreResult(FPRClass, Op, Src, -1);
}
}
void OpDispatchBuilder::VMOVUPS_VMOVUPD_Op(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1);
const auto Is128BitDest = GetDstSize(Op) == Core::CPUState::XMM_SSE_REG_SIZE;
if (Op->Dest.IsGPR() && Is128BitDest) {
// Perform 32 byte store to clear the upper lane.
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 32, 1);
} else {
StoreResult(FPRClass, Op, Src, 1);
}
}
void OpDispatchBuilder::MOVUPSOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 1);
StoreResult(FPRClass, Op, Src, 1);
@@ -68,6 +103,19 @@ void OpDispatchBuilder::MOVHPDOp(OpcodeArgs) {
}
}
void OpDispatchBuilder::VMOVHPOp(OpcodeArgs) {
if (Op->Dest.IsGPR()) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 8, 1, 0, Src1, Src2);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
} else {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Result = _VInsElement(16, 8, 0, 1, Src, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 8, 8);
}
}
void OpDispatchBuilder::MOVLPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
if (Op->Dest.IsGPR()) {
@@ -78,9 +126,10 @@ void OpDispatchBuilder::MOVLPOp(OpcodeArgs) {
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 16, 16);
}
else {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, 8, 16);
auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
auto Result = _VInsElement(16, 8, 0, 0, Dest, Src);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 8, 16);
StoreResult(FPRClass, Op, Result, -1);
}
}
else {
@@ -88,24 +137,62 @@ void OpDispatchBuilder::MOVLPOp(OpcodeArgs) {
}
}
void OpDispatchBuilder::VMOVLPOp(OpcodeArgs) {
if (Op->Dest.IsGPR()) {
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 16);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 8, 0, 0, Src1, Src2);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
} else {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Src, 8, 8);
}
}
void OpDispatchBuilder::MOVSHDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 4, 3, 3, Src, Src);
Result = _VInsElement(16, 4, 2, 3, Result, Src);
Result = _VInsElement(16, 4, 1, 1, Result, Src);
OrderedNode *Result = _VInsElement(16, 4, 2, 3, Src, Src);
Result = _VInsElement(16, 4, 0, 1, Result, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VMOVSHDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto SrcSize = GetSrcSize(Op);
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
OrderedNode *Result = _VInsElement(SrcSize, 4, 2, 3, Src, Src);
Result = _VInsElement(SrcSize, 4, 0, 1, Result, Src);
if (Is256Bit) {
Result = _VInsElement(SrcSize, 4, 4, 5, Result, Src);
Result = _VInsElement(SrcSize, 4, 6, 7, Result, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
}
void OpDispatchBuilder::MOVSLDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, 8);
OrderedNode *Result = _VInsElement(16, 4, 3, 2, Src, Src);
Result = _VInsElement(16, 4, 2, 2, Result, Src);
Result = _VInsElement(16, 4, 1, 0, Result, Src);
Result = _VInsElement(16, 4, 0, 0, Result, Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VMOVSLDUPOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
const auto SrcSize = GetSrcSize(Op);
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
OrderedNode *Result = _VInsElement(SrcSize, 4, 3, 2, Src, Src);
Result = _VInsElement(SrcSize, 4, 1, 0, Result, Src);
if (Is256Bit) {
Result = _VInsElement(SrcSize, 4, 5, 4, Result, Src);
Result = _VInsElement(SrcSize, 4, 7, 6, Result, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Result, 32, -1);
}
void OpDispatchBuilder::MOVSSOp(OpcodeArgs) {
if (Op->Dest.IsGPR() && Op->Src[0].IsGPR()) {
// MOVSS xmm1, xmm2
@@ -301,6 +388,48 @@ void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 1>(OpcodeArgs);
template
void OpDispatchBuilder::VectorALUOp<IR::OP_VUQSUB, 2>(OpcodeArgs);
template <IROps IROp, size_t ElementSize>
void OpDispatchBuilder::AVXVectorALUOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
const auto Is128Bit = Size == Core::CPUState::XMM_SSE_REG_SIZE;
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
auto ALUOp = _VAdd(Size, ElementSize, Src1, Src2);
// Overwrite our IR's op type
ALUOp.first->Header.Op = IROp;
OrderedNode* Result = ALUOp;
if (Is128Bit) {
// 128-bit variants need to zero the upper lane.
Result = _VMov(Size, ALUOp);
}
StoreResult(FPRClass, Op, Result, -1);
}
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 1>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 2>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 4>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VADD, 8>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VFADD, 4>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VFADD, 8>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VAND, 16>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VOR, 16>(OpcodeArgs);
template
void OpDispatchBuilder::AVXVectorALUOp<IR::OP_VXOR, 16>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorALUROp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
@@ -321,8 +450,9 @@ void OpDispatchBuilder::VectorALUROp<IR::OP_VFSUB, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize>
void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// If OpSize == ElementSize then it only does the lower scalar op
@@ -332,9 +462,9 @@ void OpDispatchBuilder::VectorScalarALUOp(OpcodeArgs) {
OrderedNode* Result = ALUOp;
if (Size != ElementSize) {
if (DstSize != ElementSize) {
// Insert the lower bits
Result = _VInsElement(Size, ElementSize, 0, 0, Dest, Result);
Result = _VInsElement(DstSize, ElementSize, 0, 0, Dest, ALUOp);
}
StoreResult(FPRClass, Op, Result, -1);
@@ -368,11 +498,12 @@ void OpDispatchBuilder::VectorScalarALUOp<IR::OP_VFMAX, 8>(OpcodeArgs);
template<FEXCore::IR::IROps IROp, size_t ElementSize, bool Scalar>
void OpDispatchBuilder::VectorUnaryOp(OpcodeArgs) {
auto Size = GetSrcSize(Op);
auto DstSize = GetDstSize(Op);
if constexpr (Scalar) {
Size = ElementSize;
}
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
auto ALUOp = _VFSqrt(Size, ElementSize, Src);
// Overwrite our IR's op type
@@ -380,7 +511,7 @@ void OpDispatchBuilder::VectorUnaryOp(OpcodeArgs) {
if constexpr (Scalar) {
// Insert the lower bits
auto Result = _VInsElement(GetSrcSize(Op), ElementSize, 0, 0, Dest, ALUOp);
auto Result = _VInsElement(DstSize, ElementSize, 0, 0, Dest, ALUOp);
StoreResult(FPRClass, Op, Result, -1);
}
else {
@@ -441,14 +572,8 @@ void OpDispatchBuilder::MOVQOp(OpcodeArgs) {
const auto gpr = Op->Dest.Data.GPR.GPR;
const auto gprIndex = gpr - X86State::REG_XMM_0;
const auto fprLowOffset = CTX->HostFeatures.SupportsAVX ? offsetof(Core::CPUState, xmm.avx.data[gprIndex][0])
: offsetof(Core::CPUState, xmm.sse.data[gprIndex][0]);
const auto fprHighOffset = CTX->HostFeatures.SupportsAVX ? offsetof(Core::CPUState, xmm.avx.data[gprIndex][1])
: offsetof(Core::CPUState, xmm.sse.data[gprIndex][1]);
_StoreContext(8, FPRClass, Src, fprLowOffset);
auto Const = _Constant(0);
_StoreContext(8, GPRClass, Const, fprHighOffset);
auto Reg = _VMov(16, Src);
StoreXMMRegister(gprIndex, Reg);
}
else {
// This is simple, just store the result
@@ -659,6 +784,22 @@ void OpDispatchBuilder::ANDNOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Dest, -1);
}
void OpDispatchBuilder::VANDNOp(OpcodeArgs) {
const auto Size = GetSrcSize(Op);
const auto Is128Bit = Size == Core::CPUState::XMM_SSE_REG_SIZE;
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags, -1);
Src1 = _VNot(Size, Size, Src1);
OrderedNode *Dest = _VAnd(Size, Size, Src1, Src2);
if (Is128Bit) {
Dest = _VMov(16, Dest);
}
StoreResult(FPRClass, Op, Dest, -1);
}
template<size_t ElementSize>
void OpDispatchBuilder::PINSROp(OpcodeArgs) {
auto Size = GetDstSize(Op);
@@ -926,7 +1067,11 @@ void OpDispatchBuilder::PSRLDQ(OpcodeArgs) {
auto Size = GetDstSize(Op);
auto Result = _VSRI(Size, 16, Dest, Shift);
OrderedNode *Result = _VectorZero(Size);
if (Shift < Size) {
Result = _VExtr(Size, 1, Result, Dest, Shift);
}
StoreResult(FPRClass, Op, Result, -1);
}
@@ -938,7 +1083,10 @@ void OpDispatchBuilder::PSLLDQ(OpcodeArgs) {
auto Size = GetDstSize(Op);
auto Result = _VSLI(Size, 16, Dest, Shift);
OrderedNode *Result = _VectorZero(Size);
if (Shift < Size) {
Result = _VExtr(Size, 1, Dest, Result, Size - Shift);
}
StoreResult(FPRClass, Op, Result, -1);
}
@@ -982,6 +1130,23 @@ void OpDispatchBuilder::MOVDDUPOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Res, -1);
}
void OpDispatchBuilder::VMOVDDUPOp(OpcodeArgs) {
const auto SrcSize = GetSrcSize(Op);
const auto IsSrcGPR = Op->Src[0].IsGPR();
const auto Is256Bit = SrcSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto MemSize = Is256Bit ? 32 : 8;
OrderedNode *Src = IsSrcGPR ? LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags, -1)
: LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], MemSize, Op->Flags, -1);
OrderedNode *Res = _VInsElement(SrcSize, 8, 1, 0, Src, Src);
if (Is256Bit) {
Res = _VInsElement(SrcSize, 8, 3, 2, Res, Src);
}
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, Res, 32, -1);
}
template<size_t DstElementSize>
void OpDispatchBuilder::CVTGPR_To_FPR(OpcodeArgs) {
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
@@ -1087,13 +1252,16 @@ void OpDispatchBuilder::Vector_CVT_Float_To_Int<8, true, false>(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void OpDispatchBuilder::Scalar_CVT_Float_To_Float(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags, -1);
const auto DstSize = GetDstSize(Op);
OrderedNode *Dest = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, DstSize, Op->Flags, -1);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
Src = _Float_FToF(DstElementSize, SrcElementSize, Src);
Src = _VInsElement(16, DstElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Src, -1);
auto Result = _VInsElement(DstSize, DstElementSize, 0, 0, Dest, Src);
StoreResult(FPRClass, Op, Result, -1);
}
template
@@ -1103,8 +1271,9 @@ void OpDispatchBuilder::Scalar_CVT_Float_To_Float<8, 4>(OpcodeArgs);
template<size_t DstElementSize, size_t SrcElementSize>
void OpDispatchBuilder::Vector_CVT_Float_To_Float(OpcodeArgs) {
const auto Size = GetDstSize(Op);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
size_t Size = GetDstSize(Op);
if constexpr (DstElementSize > SrcElementSize) {
Src = _Vector_FToF(Size, SrcElementSize << 1, Src, SrcElementSize);
@@ -1181,13 +1350,12 @@ void OpDispatchBuilder::XMM_To_MMX_Vector_CVT_Float_To_Int<8, true>(OpcodeArgs);
void OpDispatchBuilder::MASKMOVOp(OpcodeArgs) {
// Until we get correct PHI nodes this is required to be a loop unroll
const auto GPRSize = CTX->GetGPRSize();
const auto Size = uint32_t{GetSrcSize(Op)} * 8;
OrderedNode *Src = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, -1);
OrderedNode *Dest = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, -1);
OrderedNode *MemDest = _LoadContext(GPRSize, GPRClass, GPROffset(X86State::REG_RDI));
auto MemDest = LoadGPRRegister(X86State::REG_RDI);
const size_t NumElements = Size / 64;
for (size_t Element = 0; Element < NumElements; ++Element) {
@@ -1398,16 +1566,9 @@ void OpDispatchBuilder::FXSaveOp(OpcodeArgs) {
}
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto GetXMMOffset = [this](size_t i) {
if (CTX->HostFeatures.SupportsAVX) {
return offsetof(Core::CPUState, xmm.avx.data[i]);
} else {
return offsetof(Core::CPUState, xmm.sse.data[i]);
}
};
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *XMMReg = _LoadContext(16, FPRClass, GetXMMOffset(i));
OrderedNode *XMMReg = LoadXMMRegister(i);
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
_StoreMem(FPRClass, 16, MemLocation, XMMReg, 16);
@@ -1455,18 +1616,11 @@ void OpDispatchBuilder::FXRStoreOp(OpcodeArgs) {
}
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
const auto GetXMMOffset = [this](size_t i) {
if (CTX->HostFeatures.SupportsAVX) {
return offsetof(Core::CPUState, xmm.avx.data[i]);
} else {
return offsetof(Core::CPUState, xmm.sse.data[i]);
}
};
for (unsigned i = 0; i < NumRegs; ++i) {
OrderedNode *MemLocation = _Add(Mem, _Constant(i * 16 + 160));
auto XMMReg = _LoadMem(FPRClass, 16, MemLocation, 16);
_StoreContext(16, FPRClass, XMMReg, GetXMMOffset(i));
StoreXMMRegister(i, XMMReg);
}
}
@@ -1607,11 +1761,9 @@ void OpDispatchBuilder::MOVQ2DQ(OpcodeArgs) {
// This instruction is a bit special in that if the source is MMX then it zexts to 128bit
if constexpr (ToXMM) {
const auto Index = Op->Dest.Data.GPR.GPR - FEXCore::X86State::REG_XMM_0;
const auto Offset = CTX->HostFeatures.SupportsAVX ? offsetof(FEXCore::Core::CPUState, xmm.avx.data[Index][0])
: offsetof(FEXCore::Core::CPUState, xmm.sse.data[Index][0]);
Src = _VMov(16, Src);
_StoreContext(16, FPRClass, Src, Offset);
StoreXMMRegister(Index, Src);
}
else {
// This is simple, just store the result
@@ -2371,7 +2523,8 @@ void OpDispatchBuilder::VectorVariableBlend(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags, -1);
// The mask is hardcoded to be xmm0 in this instruction
OrderedNode *Mask = _LoadContext(16, FPRClass, offsetof(FEXCore::Core::CPUState, xmm.avx.data[0]));
auto Mask = LoadXMMRegister(0);
// Each element is selected by the high bit of that element size
// Dest[ElementIdx] = Xmm0[ElementIndex][HighBit] ? Src : Dest;
//
@@ -2606,4 +2759,28 @@ void OpDispatchBuilder::MPSADBWOp(OpcodeArgs) {
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::VZEROOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
const auto IsVZEROALL = DstSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto NumRegs = CTX->Config.Is64BitMode ? 16U : 8U;
if (IsVZEROALL) {
// NOTE: Despite the name being VZEROALL, this will still only ever
// zero out up to the first 16 registers (even on AVX-512, where we have 32 registers)
OrderedNode* ZeroVector = _VectorZero(DstSize);
for (uint32_t i = 0; i < NumRegs; i++) {
StoreXMMRegister(i, ZeroVector);
}
} else {
// Likewise, VZEROUPPER will only ever zero only up to the first 16 registers
for (uint32_t i = 0; i < NumRegs; i++) {
OrderedNode* Reg = LoadXMMRegister(i);
OrderedNode* Dst = _VMov(16, Reg);
StoreXMMRegister(i, Dst);
}
}
}
}
@@ -67,7 +67,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_6, PF_F2, 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
// GROUP 7
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -76,7 +76,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -85,7 +85,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -94,7 +94,7 @@ void InitializeSecondaryGroupTables() {
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 1), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 2), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
@@ -295,7 +295,7 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x20, 4, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x24, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2A, 1, X86InstInfo{"CVTSI2SS", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{0x2B, 1, X86InstInfo{"MOVNTSS", TYPE_INST, GenFlagsSameSize(SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x2C, 1, X86InstInfo{"CVTTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
{0x2D, 1, X86InstInfo{"CVTSS2SI", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR, 0, nullptr}},
@@ -305,18 +305,18 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x40, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x50, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 1, X86InstInfo{"RSQRTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x53, 1, X86InstInfo{"RCPSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 1, X86InstInfo{"RSQRTSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x53, 1, X86InstInfo{"RCPSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x54, 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSS2SD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSS2SD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5B, 1, X86InstInfo{"CVTTPS2DQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x60, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x68, 7, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -383,16 +383,16 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x40, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x50, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x51, 1, X86InstInfo{"SQRTSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x52, 6, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x58, 1, X86InstInfo{"ADDSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x59, 1, X86InstInfo{"MULSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5A, 1, X86InstInfo{"CVTSD2SS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5B, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5C, 1, X86InstInfo{"SUBSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5D, 1, X86InstInfo{"MINSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5E, 1, X86InstInfo{"DIVSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x5F, 1, X86InstInfo{"MAXSD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x60, 16, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -17,23 +17,23 @@ void InitializeVEXTables() {
static constexpr U16U8InfoStruct VEXTable[] = {
// Map 0 (Reserved)
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMODUPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x10), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x10), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x10), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMODUPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x11), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x11), 1, X86InstInfo{"VMOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x13), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x13), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x14), 1, X86InstInfo{"VUNPCKLPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x14), 1, X86InstInfo{"VUNPCKLPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -41,12 +41,12 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x15), 1, X86InstInfo{"VUNPCKHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x15), 1, X86InstInfo{"VUNPCKHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x16), 1, X86InstInfo{"VMOVSHDUP", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x16), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b01, 0x16), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
{OPD(1, 0b10, 0x16), 1, X86InstInfo{"VMOVSHDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x17), 1, X86InstInfo{"VMOVHPS", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x17), 1, X86InstInfo{"VMOVHPD", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x50), 1, X86InstInfo{"VMOVMSKPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x50), 1, X86InstInfo{"VMOVMSKPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -62,17 +62,17 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x53), 1, X86InstInfo{"VRCPPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b10, 0x53), 1, X86InstInfo{"VRCPSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x54), 1, X86InstInfo{"VANDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x54), 1, X86InstInfo{"VANDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x55), 1, X86InstInfo{"VANDNPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x55), 1, X86InstInfo{"VANDNPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x55), 1, X86InstInfo{"VANDNPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x55), 1, X86InstInfo{"VANDNPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x56), 1, X86InstInfo{"VORPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x56), 1, X86InstInfo{"VORPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x56), 1, X86InstInfo{"VORPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x56), 1, X86InstInfo{"VORPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x57), 1, X86InstInfo{"VXORPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x57), 1, X86InstInfo{"VDORPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x57), 1, X86InstInfo{"VXORPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x57), 1, X86InstInfo{"VXORPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x60), 1, X86InstInfo{"VPUNPCKLBW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x61), 1, X86InstInfo{"VPUNPCKLWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -95,7 +95,7 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x75), 1, X86InstInfo{"VPCMPEQW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x76), 1, X86InstInfo{"VPCMPEQD", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x77), 1, X86InstInfo{"VZERO*", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x77), 1, X86InstInfo{"VZERO*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT), 0, nullptr}},
{OPD(1, 0b00, 0xC2), 1, X86InstInfo{"VCMPccPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xC2), 1, X86InstInfo{"VCMPccPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -112,17 +112,17 @@ void InitializeVEXTables() {
// This table doesn't state which VEX.pp is for which instruction
// XXX: Confirm all the above encoding opcodes
{OPD(1, 0b00, 0x28), 1, X86InstInfo{"VMOVAPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x28), 1, X86InstInfo{"VMOVAPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x28), 1, X86InstInfo{"VMOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x28), 1, X86InstInfo{"VMOVAPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0x29), 1, X86InstInfo{"VMOVAPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x29), 1, X86InstInfo{"VMOVAPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x29), 1, X86InstInfo{"VMOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x29), 1, X86InstInfo{"VMOVAPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2A), 1, X86InstInfo{"VCVTSI2SS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2A), 1, X86InstInfo{"VCVTSI2SD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x2B), 1, X86InstInfo{"VMOVNTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x2B), 1, X86InstInfo{"VMOVNTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x2C), 1, X86InstInfo{"VCVTTSS2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x2C), 1, X86InstInfo{"VCVTTSD2SI", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -136,8 +136,8 @@ void InitializeVEXTables() {
{OPD(1, 0b00, 0x2F), 1, X86InstInfo{"VUCOMISS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x2F), 1, X86InstInfo{"VUCOMISD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b00, 0x58), 1, X86InstInfo{"VADDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x58), 1, X86InstInfo{"VADDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x58), 1, X86InstInfo{"VADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x58), 1, X86InstInfo{"VADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -179,8 +179,8 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x6D), 1, X86InstInfo{"VPUNPCKHQDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x6E), 1, X86InstInfo{"VMOV*", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x6F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x6F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7C), 1, X86InstInfo{"VHADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7C), 1, X86InstInfo{"VHADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -188,11 +188,11 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0x7D), 1, X86InstInfo{"VHSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0x7D), 1, X86InstInfo{"VHSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7E), 1, X86InstInfo{"VMOV*", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7E), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0x7F), 1, X86InstInfo{"VMOVDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b10, 0x7F), 1, X86InstInfo{"VMOVDQU", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b00, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
{OPD(1, 0b01, 0xAE), 1, X86InstInfo{"", TYPE_VEX_GROUP_15, FLAGS_NONE, 0, nullptr}}, // VEX Group 15
@@ -205,19 +205,19 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0xD1), 1, X86InstInfo{"VPSRLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD2), 1, X86InstInfo{"VPSRLD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD3), 1, X86InstInfo{"VPSRLQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD4), 1, X86InstInfo{"VPADDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD5), 1, X86InstInfo{"VPMULLW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD6), 1, X86InstInfo{"VMOVQ", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD7), 1, X86InstInfo{"VPMOVMSKB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_DST_GPR | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(1, 0b01, 0xD8), 1, X86InstInfo{"VPSUBUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xD9), 1, X86InstInfo{"VPSUBUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDA), 1, X86InstInfo{"VPMINUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDB), 1, X86InstInfo{"VPAND", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDB), 1, X86InstInfo{"VPAND", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDC), 1, X86InstInfo{"VPADDUSB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDD), 1, X86InstInfo{"VPADDUSW", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDE), 1, X86InstInfo{"VPMAXUB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDF), 1, X86InstInfo{"VPANDN", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xDF), 1, X86InstInfo{"VPANDN", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE0), 1, X86InstInfo{"VPAVGB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE1), 1, X86InstInfo{"VPSRAW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -230,16 +230,16 @@ void InitializeVEXTables() {
{OPD(1, 0b10, 0xE6), 1, X86InstInfo{"VCVTDQ2PD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b11, 0xE6), 1, X86InstInfo{"VCVTPD2DQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE7), 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE7), 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xE8), 1, X86InstInfo{"VPSUBSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xE9), 1, X86InstInfo{"VPSUBSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEB), 1, X86InstInfo{"VPOR", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEA), 1, X86InstInfo{"VPMINSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEB), 1, X86InstInfo{"VPOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEC), 1, X86InstInfo{"VPADDSB", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xED), 1, X86InstInfo{"VPADDSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xEE), 1, X86InstInfo{"VPMAXSW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 0b01, 0xEF), 1, X86InstInfo{"VPXOR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b11, 0xF0), 1, X86InstInfo{"VLDDQU", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -255,9 +255,9 @@ void InitializeVEXTables() {
{OPD(1, 0b01, 0xF9), 1, X86InstInfo{"VPSUBW", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFA), 1, X86InstInfo{"VPSUBD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFB), 1, X86InstInfo{"VPSUBQ", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFC), 1, X86InstInfo{"VPADDB", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFD), 1, X86InstInfo{"VPADDW", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFE), 1, X86InstInfo{"VPADDD", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFC), 1, X86InstInfo{"VPADDB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFD), 1, X86InstInfo{"VPADDW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(1, 0b01, 0xFE), 1, X86InstInfo{"VPADDD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
// VEX Map 2
{OPD(2, 0b01, 0x00), 1, X86InstInfo{"VPSHUFB", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -298,7 +298,7 @@ void InitializeVEXTables() {
{OPD(2, 0b01, 0x28), 1, X86InstInfo{"VPMULDQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x29), 1, X86InstInfo{"VPCMPEQQ", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2A), 1, X86InstInfo{"VMOVNTDQA", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(2, 0b01, 0x2B), 1, X86InstInfo{"VPACKUSDW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2C), 1, X86InstInfo{"VMASKMOVPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(2, 0b01, 0x2D), 1, X86InstInfo{"VMASKMOVPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
+4 -1
View File
@@ -199,7 +199,7 @@ namespace FEXCore {
const uint8_t GPRSize = CTX->GetGPRSize();
emit->_StoreContext(GPRSize, IR::GPRClass, emit->_Constant(Entrypoint), offsetof(Core::CPUState, gregs[X86State::REG_R11]));
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, gregs[X86State::REG_R11]), IR::GPRClass, IR::GPRFixedClass, GPRSize);
emit->_ExitFunction(emit->_Constant(GuestThunkEntrypoint));
}, CTX->ThunkHandler.get(), (void*)args->target_addr);
@@ -431,6 +431,9 @@ namespace FEXCore {
FEX_DEFAULT_VISIBILITY
void FinalizeHostTrampolineForGuestFunction(HostToGuestTrampolinePtr* TrampolineAddress, void* HostPacker) {
if (TrampolineAddress == nullptr) return;
auto& Trampoline = GetInstanceInfo(TrampolineAddress);
LOGMAN_THROW_A_FMT(Trampoline.CallCallback == (uintptr_t)&ThunkHandler_impl::CallCallback,
+12 -30
View File
@@ -355,7 +355,9 @@
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContext to XMM\""
]
},
@@ -370,7 +372,9 @@
"EmitValidation": [
"WalkFindRegClass($Value) == $Class",
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($Offset >= offsetof(Core::CPUState, gregs[0]) && $Offset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContext to GPR\"",
"!($Offset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $Offset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContext to XMM\""
]
},
@@ -381,7 +385,9 @@
"DestSize": "ByteSize",
"EmitValidation": [
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't LoadContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't LoadContextIndexed to XMM\""
]
},
"StoreContextIndexed SSA:$Value, GPR:$Index, u8:#ByteSize, u32:$BaseOffset, u32:$Stride, RegisterClass:$Class": {
@@ -393,7 +399,9 @@
"EmitValidation": [
"WalkFindRegClass($Value) == $Class",
"($Class == GPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8)) || $Class == FPRClass",
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass"
"($Class == FPRClass && (#ByteSize == 1 || #ByteSize == 2 || #ByteSize == 4 || #ByteSize == 8 || #ByteSize == 16 || #ByteSize == 32)) || $Class == GPRClass",
"!($BaseOffset >= offsetof(Core::CPUState, gregs[0]) && $BaseOffset < offsetof(Core::CPUState, gregs[16])) && \"Can't StoreContextIndexed to GPR\"",
"!($BaseOffset >= offsetof(Core::CPUState, xmm.avx.data[0]) && $BaseOffset < offsetof(Core::CPUState, xmm.avx.data[16])) && \"Can't StoreContextIndexed to XMM\""
]
},
@@ -471,22 +479,6 @@
]
},
"FPR = VLoadMemElement u8:#RegisterSize, u8:#ElementSize, FPR:$Value, GPR:$Addr, u8:$Index, u8:$Align{1}": {
"Desc": ["Loads an element of size #ElementSize in to $Value from $Addr at $Index"
],
"OpClass": "Memory",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"VStoreMemElement u8:#RegisterSize, u8:#ElementSize, FPR:$Value, GPR:$Addr, u8:$Index, u8:$Align": {
"Desc": ["Stores an element of size #ElementSize from $Value[$Index] to $Addr"
],
"HasSideEffects": true,
"DestSize": "ElementSize",
"NumElements": "RegisterSize / ElementSize"
},
"CacheLineClear GPR:$Addr": {
"Desc": ["Does a 64 byte cacheline clear at the address specified",
"Only clears the data cachelines. Doesn't do any zeroing"
@@ -1022,16 +1014,6 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VSLI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$ByteShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VSRI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$ByteShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VShlI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
+6
View File
@@ -162,6 +162,12 @@ class IRParser: public FEXCore::IR::IREmitter {
else if (Arg == "FPR") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRClass};
}
else if (Arg == "GPRFixed") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRFixedClass};
}
else if (Arg == "FPRFixed") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRFixedClass};
}
else if (Arg == "GPRPair") {
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRPairClass};
}
-9
View File
@@ -37,15 +37,6 @@ void PassManager::AddDefaultPasses(FEXCore::Context::Context *ctx, bool InlineCo
InsertPass(CreateSyscallOptimization());
InsertPass(CreatePassDeadCodeElimination());
// only do SRA if enabled and JIT
if (InlineConstants && StaticRegisterAllocation)
InsertPass(CreateStaticRegisterAllocationPass(ctx->HostFeatures.SupportsAVX));
}
else {
// only do SRA if enabled and JIT
if (InlineConstants && StaticRegisterAllocation)
InsertPass(CreateStaticRegisterAllocationPass(ctx->HostFeatures.SupportsAVX));
}
// If the IR is compacted post-RA then the node indexing gets messed up and the backend isn't able to find the register assigned to a node
-1
View File
@@ -21,7 +21,6 @@ std::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction(FEXCore::Utils::IntrusiveP
std::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass,
bool OptimizeSRA,
bool SupportsAVX);
std::unique_ptr<FEXCore::IR::Pass> CreateStaticRegisterAllocationPass(bool SupportsAVX);
std::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass();
namespace Validation {
@@ -1,131 +0,0 @@
/*
$info$
tags: ir|opts
desc: Replaces Load/StoreContext with Load/StoreReg for SRA regs
$end_info$
*/
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <memory>
#include <stddef.h>
#include <stdint.h>
namespace FEXCore::IR {
class StaticRegisterAllocationPass final : public FEXCore::IR::Pass {
public:
explicit StaticRegisterAllocationPass(bool SupportsAVX_) : SupportsAVX{SupportsAVX_} {}
bool Run(IREmitter *IREmit) override;
private:
bool SupportsAVX;
bool IsStaticAllocGpr(uint32_t Offset, RegisterClassType Class) const {
const auto begin = offsetof(Core::CPUState, gregs[0]);
const auto end = offsetof(Core::CPUState, gregs[16]);
if (Offset >= begin && Offset < end) {
const auto reg = (Offset - begin) / Core::CPUState::GPR_REG_SIZE;
LOGMAN_THROW_AA_FMT(Class.Val == IR::GPRClass.Val, "unexpected Class {}", Class);
// 0..15 -> 16 in total
return reg < Core::CPUState::NUM_GPRS;
}
return false;
}
bool IsStaticAllocFpr(uint32_t Offset, RegisterClassType Class, bool AllowGpr) const {
const auto [begin, end] = [this]() -> std::pair<ptrdiff_t, ptrdiff_t> {
if (SupportsAVX) {
return {
offsetof(Core::CPUState, xmm.avx.data[0][0]),
offsetof(Core::CPUState, xmm.avx.data[16][0]),
};
} else {
return {
offsetof(Core::CPUState, xmm.sse.data[0][0]),
offsetof(Core::CPUState, xmm.sse.data[16][0]),
};
}
}();
if (Offset >= begin && Offset < end) {
const auto size = SupportsAVX ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
const auto reg = (Offset - begin) / size;
LOGMAN_THROW_AA_FMT(Class.Val == IR::FPRClass.Val || (AllowGpr && Class.Val == IR::GPRClass.Val), "unexpected Class {}, AllowGpr {}", Class, AllowGpr);
// 0..15 -> 16 in total
return reg < Core::CPUState::NUM_XMMS;
}
return false;
}
};
/**
* @brief This pass replaces Load/Store Context with Load/Store Register for Statically Mapped registers. It also does some validation.
*
*/
bool StaticRegisterAllocationPass::Run(IREmitter *IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::SRA");
auto CurrentIR = IREmit->ViewIR();
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
IREmit->SetWriteCursor(CodeNode);
if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->CW<IR::IROp_LoadContext>();
if (IsStaticAllocGpr(Op->Offset, Op->Class) || IsStaticAllocFpr(Op->Offset, Op->Class, true)) {
auto GeneralClass = Op->Class;
if (IsStaticAllocFpr(Op->Offset, GeneralClass, true) && GeneralClass == GPRClass) {
GeneralClass = FPRClass;
}
auto StaticClass = GeneralClass == GPRClass ? GPRFixedClass : FPRFixedClass;
OrderedNode *sraReg = IREmit->_LoadRegister(false, Op->Offset, GeneralClass, StaticClass, Op->Header.Size);
if (GeneralClass != Op->Class) {
sraReg = IREmit->_VExtractToGPR(Op->Header.Size, Op->Header.Size, sraReg, 0);
}
IREmit->ReplaceAllUsesWith(CodeNode, sraReg);
}
} if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
if (IsStaticAllocGpr(Op->Offset, Op->Class) || IsStaticAllocFpr(Op->Offset, Op->Class, true)) {
auto val = IREmit->UnwrapNode(Op->Value);
auto GeneralClass = Op->Class;
if (IsStaticAllocFpr(Op->Offset, GeneralClass, true) && GeneralClass == GPRClass) {
val = IREmit->_VCastFromGPR(Op->Header.Size, Op->Header.Size, val);
GeneralClass = FPRClass;
}
auto StaticClass = GeneralClass == GPRClass ? GPRFixedClass : FPRFixedClass;
IREmit->_StoreRegister(val, false, Op->Offset, GeneralClass, StaticClass, Op->Header.Size);
IREmit->Remove(CodeNode);
}
}
}
}
return true;
}
std::unique_ptr<FEXCore::IR::Pass> CreateStaticRegisterAllocationPass(bool SupportsAVX) {
return std::make_unique<StaticRegisterAllocationPass>(SupportsAVX);
}
}
+3 -1
View File
@@ -74,7 +74,9 @@ namespace Handler {
LAYER_GLOBAL_MAIN, ///< /usr/share/fex-emu/Config.json by default
LAYER_MAIN,
LAYER_ARGUMENTS,
LAYER_GLOBAL_STEAM_APP,
LAYER_GLOBAL_APP,
LAYER_LOCAL_STEAM_APP,
LAYER_LOCAL_APP,
LAYER_ENVIRONMENT,
LAYER_TOP,
@@ -272,7 +274,7 @@ namespace Type {
*
* @return unique_ptr for that layer
*/
FEX_DEFAULT_VISIBILITY std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, bool Global);
FEX_DEFAULT_VISIBILITY std::unique_ptr<FEXCore::Config::Layer> CreateAppLayer(const std::string& Filename, FEXCore::Config::LayerType Type);
/**
* @brief iCreate an environment configuration loader
+8 -4
View File
@@ -218,12 +218,13 @@ namespace FEXCore::Core {
uint32_t SignalHandlerRefCounter{};
struct SynchronousFaultDataStruct {
struct alignas(8) SynchronousFaultDataStruct {
bool FaultToTopAndGeneratedException{};
uint8_t Signal;
uint32_t TrapNo;
uint32_t err_code;
uint32_t si_code;
uint8_t TrapNo;
uint8_t si_code;
uint16_t err_code;
uint32_t _pad : 16;
} SynchronousFaultData;
InternalThreadState* Thread;
@@ -237,6 +238,9 @@ namespace FEXCore::Core {
static_assert(offsetof(CpuStateFrame, Pointers) + sizeof(CpuStateFrame::Pointers) <= 32760, "JITPointers maximum pointer needs to be less than architecture maximum 32768");
static_assert(std::is_standard_layout<CpuStateFrame>::value, "This needs to be standard layout");
static_assert(sizeof(CpuStateFrame::SynchronousFaultData) == 8, "This needs to be 8 bytes");
static_assert(std::alignment_of_v<CpuStateFrame::SynchronousFaultDataStruct> == 8, "This needs to be 8 bytes");
static_assert(offsetof(CpuStateFrame, SynchronousFaultData) % 8 == 0, "This needs to be aligned");
FEX_DEFAULT_VISIBILITY std::string_view const& GetFlagName(unsigned Flag);
FEX_DEFAULT_VISIBILITY std::string_view const& GetGRegName(unsigned Reg);
+1
View File
@@ -27,6 +27,7 @@ class HostFeatures final {
bool SupportsSHA{};
bool SupportsBMI1{};
bool SupportsBMI2{};
bool SupportsPMULL_128Bit{};
// Float exception behaviour
bool SupportsFlushInputsToZero{};
@@ -8,8 +8,8 @@
#include <FEXCore/Utils/InterruptableConditionVariable.h>
#include <FEXCore/Utils/Threads.h>
#include <map>
#include <unordered_map>
#include <tsl/robin_map.h>
#include <shared_mutex>
namespace FEXCore {
@@ -101,7 +101,7 @@ namespace FEXCore::Core {
std::unique_ptr<FEXCore::CPU::CPUBackend> CPUBackend;
std::unique_ptr<FEXCore::LookupCache> LookupCache;
std::unordered_map<uint64_t, LocalIREntry> DebugStore;
tsl::robin_map<uint64_t, LocalIREntry> DebugStore;
std::unique_ptr<FEXCore::Frontend::Decoder> FrontendDecoder;
std::unique_ptr<FEXCore::IR::PassManager> PassManager;
@@ -115,7 +115,7 @@ namespace FEXCore::Core {
std::shared_mutex ObjectCacheRefCounter{};
bool DestroyedByParent{false}; // Should the parent destroy this thread, or it destory itself
alignas(16) FEXCore::Core::CpuStateFrame BaseFrameState{};
};
+1
View File
@@ -1,4 +1,5 @@
#pragma once
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
+5
View File
@@ -280,6 +280,11 @@ public:
return Wrapper.GetNode(GetListData());
}
///< Gets an OrderedNode from the IRListView as an OrderedNodeWrapper.
[[nodiscard]] OrderedNodeWrapper WrapNode(OrderedNode *Node) const {
return Node->Wrapped(GetListData());
}
private:
struct BlockRange {
using iterator = NodeIterator;
+1 -1
+2 -3
View File
@@ -82,9 +82,8 @@ def IsSupportedDistro():
if Distro[0] == "ubuntu":
# We only support what is available in ppa:fex-emu/fex
return Distro[1] == "20.04" or \
Distro[1] == "21.04" or \
Distro[1] == "21.10" or \
Distro[1] == "22.04"
Distro[1] == "22.04" or \
Distro[1] == "22.10"
return False
+12 -2
View File
@@ -109,8 +109,18 @@ namespace FEX::Config {
}
}
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, true));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, false));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, FEXCore::Config::LayerType::LAYER_GLOBAL_APP));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, FEXCore::Config::LayerType::LAYER_LOCAL_APP));
auto SteamID = getenv("SteamAppId");
if (SteamID) {
// If a SteamID exists then let's search for Steam application configs as well.
// We want to key off both the SteamAppId number /and/ the executable since we may not want to thunk all binaries.
auto SteamAppName = fmt::format("Steam_{}_{}", SteamID, ProgramName.string());
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(SteamAppName, FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(SteamAppName, FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP));
}
return std::make_pair(Program, ProgramName);
}
return {};
+61 -11
View File
@@ -100,33 +100,78 @@ namespace FEXServerClient {
return GetServerLockFolder() + "RootFS.lock";
}
std::string GetServerSocketFile() {
FEX_CONFIG_OPT(ServerSocketPath, SERVERSOCKETPATH);
if (ServerSocketPath().empty()) {
return fmt::format("{}/{}.FEXServer.socket", std::filesystem::temp_directory_path().string(), ::geteuid());
std::string GetServerMountFolder() {
// We need a FEXServer mount directory that has some tricky requirements.
// - We don't want to use `/tmp/` if possible.
// - systemd services use `PrivateTmp` feature to gives services their own tmp.
// - We will use this as a fallback path /only/.
// - Can't be `[$XDG_DATA_HOME,$HOME]/.fex-emu/`
// - Might be mounted with a filesystem (sshfs) which can't handle mount points inside it.
//
// Directories it can be in:
// - $XDG_RUNTIME_DIR if set
// - Is typically `/run/user/<UID>/`
// - systemd `PrivateTmp` feature doesn't touch this.
// - If this path doesn't exist then fallback to `/tmp/` as a last resort.
// - pressure-vessel explicitly creates an internal XDG_RUNTIME_DIR inside its chroot.
// - This is okay since pressure-vessel rbinds the FEX rootfs from the host to `/run/pressure-vessel/interpreter-root`.
std::string Folder{};
auto XDGRuntimeEnv = getenv("XDG_RUNTIME_DIR");
if (XDGRuntimeEnv) {
// If the XDG runtime directory works then use that.
Folder = XDGRuntimeEnv;
}
else {
// Fallback to `/tmp/` if XDG_RUNTIME_DIR doesn't exist.
// Might not be ideal but we don't have much of a choice.
Folder = std::filesystem::temp_directory_path().string();
}
return ServerSocketPath;
if (FEXCore::Config::FindContainer() == "pressure-vessel") {
// In pressure-vessel the mount point changes location.
// This is due to pressure-vesssel being a chroot environment.
// It by default maps the host-filesystem to `/run/host/` so we need to redirect.
// After pressure-vessel is fully set up it will set the `FEX_ROOTFS` environment variable,
// which the FEXInterpreter will pick up on.
Folder = "/run/host/" + Folder;
}
return Folder;
}
std::string GetServerSocketName() {
return fmt::format("{}.FEXServer.Socket", ::geteuid());
}
int GetServerFD() {
return ServerFD;
}
int ConnectToServer() {
auto ServerSocketFile = GetServerSocketFile();
int ConnectToServer(ConnectionOption ConnectionOption) {
auto ServerSocketName = GetServerSocketName();
// Create the initial unix socket
int SocketFD = socket(AF_UNIX, SOCK_STREAM, 0);
if (SocketFD == -1) {
LogMan::Msg::EFmt("Couldn't open AF_UNIX socket {} {}", errno, strerror(errno));
return -1;
}
// AF_UNIX has a special feature for named socket paths.
// If the name of the socket begins with `\0` then it is an "abstract" socket address.
// The entirety of the name is used as a path to a socket that doesn't have any filesystem backing.
struct sockaddr_un addr{};
addr.sun_family = AF_UNIX;
strncpy(addr.sun_path, ServerSocketFile.data(), std::min(ServerSocketFile.size(), sizeof(addr.sun_path)));
size_t SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
addr.sun_path[0] = 0; // Abstract AF_UNIX sockets start with \0
strncpy(addr.sun_path + 1, ServerSocketName.data(), SizeOfSocketString);
// Include final null character.
size_t SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
if (connect(SocketFD, reinterpret_cast<struct sockaddr*>(&addr), sizeof(addr)) == -1) {
if (connect(SocketFD, reinterpret_cast<struct sockaddr*>(&addr), SizeOfAddr) == -1) {
if (ConnectionOption == ConnectionOption::Default || errno != ECONNREFUSED) {
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket {} {} {}", ServerSocketName, errno, strerror(errno));
}
close(SocketFD);
return -1;
}
@@ -153,7 +198,7 @@ namespace FEXServerClient {
}
int ConnectToAndStartServer(char *InterpreterPath) {
int ServerFD = ConnectToServer();
int ServerFD = ConnectToServer(ConnectionOption::NoPrintConnectionError);
if (ServerFD == -1) {
// Couldn't connect to the server. Start one
@@ -210,7 +255,7 @@ namespace FEXServerClient {
while (poll(&PollFD, 1, -1) == -1 && errno == EINTR);
for (size_t i = 0; i < 5; ++i) {
ServerFD = ConnectToServer();
ServerFD = ConnectToServer(ConnectionOption::Default);
if (ServerFD != -1) {
break;
@@ -218,6 +263,11 @@ namespace FEXServerClient {
std::this_thread::sleep_for(std::chrono::seconds(1));
}
if (ServerFD == -1) {
// Still couldn't connect to the socket.
LogMan::Msg::EFmt("Couldn't connect to FEXServer socket {} after launching the process", GetServerSocketName());
}
}
}
return ServerFD;
+7 -2
View File
@@ -50,7 +50,8 @@ namespace FEXServerClient {
std::string GetServerLockFolder();
std::string GetServerLockFile();
std::string GetServerRootFSLockFile();
std::string GetServerSocketFile();
std::string GetServerMountFolder();
std::string GetServerSocketName();
int GetServerFD();
bool SetupClient(char *InterpreterPath);
@@ -62,12 +63,16 @@ namespace FEXServerClient {
*/
int ConnectToAndStartServer(char *InterpreterPath);
enum class ConnectionOption {
Default,
NoPrintConnectionError,
};
/**
* @brief Connect to a FEXServer instance if it exists
*
* @return socket FD for communicating with server
*/
int ConnectToServer();
int ConnectToServer(ConnectionOption ConnectionOption = ConnectionOption::Default);
/**
* @name Packet request functions
+82 -13
View File
@@ -5,6 +5,7 @@
#include "Common/FDUtils.h"
#include "FEXCore/Utils/Allocator.h"
#include "Tests/LinuxSyscalls/Syscalls.h"
#include "Tests/VDSO_Emulation.h"
#include "Linux/Utils/ELFParser.h"
#include "Linux/Utils/ELFSymbolDatabase.h"
@@ -20,6 +21,8 @@
#include <FEXCore/Core/CodeLoader.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Core/UContext.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXHeaderUtils/Syscalls.h>
@@ -177,19 +180,21 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
static std::string ResolveRootfsFile(std::string const &File, std::string RootFS) {
// If the path is relative then just run that
if (std::filesystem::path(File).is_relative()) {
if (File[0] != '/') {
return File;
}
std::string RootFSLink = RootFS + File;
while (std::filesystem::is_symlink(RootFSLink)) {
char Filename[PATH_MAX];
while(FEX::HLE::IsSymlink(RootFSLink)) {
// Do some special handling if the RootFS's linker is a symlink
// Ubuntu's rootFS by default provides an absolute location symlink to the linker
// Resolve this around back to the rootfs
auto SymlinkTarget = std::filesystem::read_symlink(RootFSLink);
if (SymlinkTarget.is_absolute()) {
RootFSLink = RootFS + SymlinkTarget.string();
auto SymlinkSize = FEX::HLE::GetSymlink(RootFSLink, Filename, PATH_MAX - 1);
if (SymlinkSize > 0 && Filename[0] == '/') {
RootFSLink = RootFS;
RootFSLink += std::string_view(Filename, SymlinkSize);
}
else {
break;
@@ -529,24 +534,35 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
AuxVariables.emplace_back(auxv_t{17, getauxval(AT_CLKTCK)}); // AT_CLKTIK
AuxVariables.emplace_back(auxv_t{6, 0x1000}); // AT_PAGESIZE
AuxRandom = &AuxVariables.emplace_back(auxv_t{25, ~0ULL}); // AT_RANDOM
AuxVariables.emplace_back(auxv_t{23, 0}); // AT_SECURE
AuxVariables.emplace_back(auxv_t{23, getauxval(AT_SECURE)}); // AT_SECURE
AuxVariables.emplace_back(auxv_t{8, 0}); // AT_FLAGS
AuxVariables.emplace_back(auxv_t{5, MainElf.phdrs.size()}); // AT_PHNUM
AuxVariables.emplace_back(auxv_t{16, HWCap}); // AT_HWCAP
AuxVariables.emplace_back(auxv_t{26, HWCap2}); // AT_HWCAP2
AuxVariables.emplace_back(auxv_t{51, CalculateSignalStackSize()}); // AT_MINSIGSTKSZ
AuxPlatform = &AuxVariables.emplace_back(auxv_t{24, ~0ULL}); // AT_PLATFORM
if (Is64BitMode()) {
AuxVariables.emplace_back(auxv_t{4, 0x38}); // AT_PHENT
// we don't support vsyscall so we don't set those
//AuxVariables.emplace_back(auxv_t{32, 0}); // AT_SYSINFO - Entry point to syscall
}
else {
AuxVariables.emplace_back(auxv_t{4, 0x20}); // AT_PHENT
// we don't support vsyscall so we don't set those
//AuxVariables.emplace_back(auxv_t{32, 0}); // AT_SYSINFO - Entry point to syscall
auto VSyscallEntry = FEX::VDSO::GetVSyscallEntry(VDSOBase);
if (!VSyscallEntry) [[unlikely]] {
// If the VDSO thunk doesn't exist then we might not have a vsyscall entry.
// Newer glibc requires vsyscall to exist now. So let's allocate a buffer and stick a vsyscall in to it.
auto VSyscallPage = Mapper(nullptr, FHU::FEX_PAGE_SIZE, PROT_READ | PROT_WRITE, MAP_ANONYMOUS | MAP_PRIVATE, -1, 0);
constexpr static uint8_t VSyscallCode[] = {
0xcd, 0x80, // int 0x80
0xc3, // ret
};
memcpy(VSyscallPage, VSyscallCode, sizeof(VSyscallCode));
mprotect(VSyscallPage, FHU::FEX_PAGE_SIZE, PROT_READ);
VSyscallEntry = reinterpret_cast<uint64_t>(VSyscallPage);
}
AuxVariables.emplace_back(auxv_t{32, VSyscallEntry}); // AT_SYSINFO - Entry point to syscall
}
if (VDSOBase) {
@@ -769,7 +785,7 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
return BaseOffset;
}
bool Is64BitMode() {
bool Is64BitMode() const {
return MainElf.type == ::ELFLoader::ELFContainer::TYPE_X86_64;
}
@@ -795,6 +811,57 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
// 0 - MONITOR/MWAIT available in CPL3
// 1 - FSGSBASE instructions available in CPL3
HWCap2 = 0;
// We need to know if we support AVX for AT_MINSIGSTKSZ
SupportsAVX = !!(res_1.ecx & (1U << 28));
}
uint64_t CalculateSignalStackSize() const {
// We must calculate the required signal stack size that the "kernel" consumes.
// For FEX this means the amount of state we store in to the guest stack, not including the amount
// that FEX stores in to the host stack as well.
//
// This needs to match what we do in FEXCore's dispatcher (Which should at some point be moved to the frontend).
//
// This roughly means that we need to calculate the combined size of:
// - xstate or _libc_fstate depending on AVX support
// - ucontext_t
// - siginfo_t
// Size of state requiring to be stored is different between 32-bit and 64-bit.
uint64_t Result{};
if (Is64BitMode()) {
Result += sizeof(FEXCore::x86_64::ucontext_t);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86_64::ucontext_t));
if (SupportsAVX) {
Result += sizeof(FEXCore::x86_64::xstate);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86_64::xstate));
}
else {
Result += sizeof(FEXCore::x86_64::_libc_fpstate);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86_64::_libc_fpstate));
}
Result += sizeof(siginfo_t);
Result = FEXCore::AlignUp(Result, alignof(siginfo_t));
}
else {
Result += sizeof(FEXCore::x86::ucontext_t);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86::ucontext_t));
if (SupportsAVX) {
Result += sizeof(FEXCore::x86::xstate);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86::xstate));
}
else {
Result += sizeof(FEXCore::x86::_libc_fpstate);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86::_libc_fpstate));
}
Result += sizeof(FEXCore::x86::siginfo_t);
Result = FEXCore::AlignUp(Result, alignof(FEXCore::x86::siginfo_t));
}
return Result;
}
constexpr static uint64_t BRK_SIZE = 8 * 1024 * 1024;
@@ -815,13 +882,15 @@ class ELFCodeLoader2 final : public FEXCore::CodeLoader {
void* VDSOBase{};
uint64_t HWCap{};
uint64_t HWCap2{};
bool SupportsAVX{};
auxv_t *AuxRandom{};
auxv_t *AuxPlatform{};
static constexpr std::string_view platform_name_x86_64 = "x86_64";
static constexpr std::string_view platform_name_i686 = "i686";
static constexpr size_t platform_string_max_size = std::max(platform_name_x86_64.size(), platform_name_i686.size());
// Need to include null character.
static constexpr size_t platform_string_max_size = std::max(platform_name_x86_64.size(), platform_name_i686.size()) + 1;
FEX_CONFIG_OPT(AdditionalArguments, ADDITIONALARGUMENTS);
};
+47 -45
View File
@@ -191,10 +191,9 @@ bool IsInterpreterInstalled() {
// The interpreter is installed if both the binfmt_misc handlers are available
// Or if we were originally executed with FD. Which means the interpreter is installed
std::error_code ec{};
return ExecutedWithFD ||
(std::filesystem::exists("/proc/sys/fs/binfmt_misc/FEX-x86", ec) &&
std::filesystem::exists("/proc/sys/fs/binfmt_misc/FEX-x86_64", ec));
(access("/proc/sys/fs/binfmt_misc/FEX-x86", F_OK) == 0 &&
access("/proc/sys/fs/binfmt_misc/FEX-x86_64", F_OK) == 0);
}
int main(int argc, char **argv, char **const envp) {
@@ -428,38 +427,39 @@ int main(int argc, char **argv, char **const envp) {
});
}
if (AOTIRLoad() || AOTIRCapture() || AOTIRGenerate()) {
const bool AOTEnabled = AOTIRLoad() || AOTIRCapture() || AOTIRGenerate();
if (AOTEnabled) {
LogMan::Msg::IFmt("Warning: AOTIR is experimental, and might lead to crashes. "
"Capture doesn't work with programs that fork.");
FEXCore::Context::SetAOTIRLoader(CTX, [](const std::string &fileid) -> int {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
return open(filepath.c_str(), O_RDONLY);
});
FEXCore::Context::SetAOTIRWriter(CTX, [](const std::string& fileid) -> std::unique_ptr<std::ofstream> {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto AOTWrite = std::make_unique<std::ofstream>(filepath, std::ios::out | std::ios::binary);
if (*AOTWrite) {
std::filesystem::resize_file(filepath, 0);
AOTWrite->seekp(0);
LogMan::Msg::IFmt("AOTIR: Storing {}", fileid);
} else {
LogMan::Msg::IFmt("AOTIR: Failed to store {}", fileid);
}
return AOTWrite;
});
FEXCore::Context::SetAOTIRRenamer(CTX, [](const std::string& fileid) -> void {
auto TmpFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto NewFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
// Rename the temporary file to atomically update the file
std::filesystem::rename(TmpFilepath, NewFilepath);
});
}
FEXCore::Context::SetAOTIRLoader(CTX, [](const std::string &fileid) -> int {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
return open(filepath.c_str(), O_RDONLY);
});
FEXCore::Context::SetAOTIRWriter(CTX, [](const std::string& fileid) -> std::unique_ptr<std::ofstream> {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto AOTWrite = std::make_unique<std::ofstream>(filepath, std::ios::out | std::ios::binary);
if (*AOTWrite) {
std::filesystem::resize_file(filepath, 0);
AOTWrite->seekp(0);
LogMan::Msg::IFmt("AOTIR: Storing {}", fileid);
} else {
LogMan::Msg::IFmt("AOTIR: Failed to store {}", fileid);
}
return AOTWrite;
});
FEXCore::Context::SetAOTIRRenamer(CTX, [](const std::string& fileid) -> void {
auto TmpFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir.tmp");
auto NewFilepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".aotir");
// Rename the temporary file to atomically update the file
std::filesystem::rename(TmpFilepath, NewFilepath);
});
if (AOTIRGenerate()) {
for(auto &Section: Loader.Sections) {
FEX::AOT::AOTGenSection(CTX, Section);
@@ -468,21 +468,23 @@ int main(int argc, char **argv, char **const envp) {
FEXCore::Context::RunUntilExit(CTX);
}
std::filesystem::create_directories(std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir", ec);
if (!ec) {
FEXCore::Context::WriteFilesWithCode(CTX, [](const std::string& fileid, const std::string& filename) {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".path");
int fd = open(filepath.c_str(), O_CREAT | O_EXCL | O_WRONLY, 0644);
if (fd != -1) {
write(fd, filename.c_str(), filename.size());
close(fd);
}
});
}
if (AOTEnabled) {
std::filesystem::create_directories(std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir", ec);
if (!ec) {
FEXCore::Context::WriteFilesWithCode(CTX, [](const std::string& fileid, const std::string& filename) {
auto filepath = std::filesystem::path(FEXCore::Config::GetDataDirectory()) / "aotir" / (fileid + ".path");
int fd = open(filepath.c_str(), O_CREAT | O_EXCL | O_WRONLY, 0644);
if (fd != -1) {
write(fd, filename.c_str(), filename.size());
close(fd);
}
});
}
if (AOTIRCapture() || AOTIRGenerate()) {
FEXCore::Context::FinalizeAOTIRCache(CTX);
LogMan::Msg::IFmt("AOTIR Cache Stored");
if (AOTIRCapture() || AOTIRGenerate()) {
FEXCore::Context::FinalizeAOTIRCache(CTX);
LogMan::Msg::IFmt("AOTIR Cache Stored");
}
}
auto ProgramStatus = FEXCore::Context::GetProgramStatus(CTX);
@@ -622,6 +622,11 @@ namespace FEX::EmulatedFile {
EmulatedFDManager::EmulatedFDManager(FEXCore::Context::Context *ctx)
: CTX {ctx} {
FDReadCreators["/proc/cpuinfo"] = [&](FEXCore::Context::Context *ctx, int32_t fd, const char *pathname, int32_t flags, mode_t mode) -> int32_t {
// Only allow a single thread to initialize the cpu_info.
// Jit in-case multiple threads try to initialize at once.
// Check if deferred cpuinfo initialization has occured.
std::call_once(cpu_info_initialized, [&]() { cpu_info = GenerateCPUInfo(ctx, ThreadsConfig()); });
int FD = GenTmpFD();
write(FD, (void*)&cpu_info.at(0), cpu_info.size());
lseek(FD, 0, SEEK_SET);
@@ -700,8 +705,6 @@ namespace FEX::EmulatedFile {
if (CPUCores > 1) {
cpus_online += "-" + std::to_string(CPUCores - 1);
}
cpu_info = GenerateCPUInfo(ctx, CPUCores);
}
EmulatedFDManager::~EmulatedFDManager() {
@@ -734,7 +737,7 @@ namespace FEX::EmulatedFile {
}
std::error_code ec;
bool exists = std::filesystem::exists(Path, ec);
bool exists = access(Path.c_str(), F_OK) == 0;
if (ec) {
return -1;
}
@@ -27,6 +27,7 @@ namespace FEX::EmulatedFile {
private:
FEXCore::Context::Context *CTX;
std::string cpus_online{};
std::once_flag cpu_info_initialized{};
std::string cpu_info{};
using FDReadStringFunc = std::function<int32_t(FEXCore::Context::Context *ctx, int32_t fd, const char *pathname, int32_t flags, mode_t mode)>;
std::unordered_map<std::string, FDReadStringFunc> FDReadCreators;
+33 -9
View File
@@ -249,17 +249,40 @@ FileManager::FileManager(FEXCore::Context::Context *ctx)
}
};
// We try to load ThunksDB from {FEX global config, FEX user config, Defined ThunksConfig option, AppConfig Global, AppConfig Local}
// We try to load ThunksDB from:
// - FEX global config
// - FEX user config
// - Defined ThunksConfig option
// - Steam AppConfig Global
// - AppConfig Global
// - Steam AppConfig Local
// - AppConfig Local
// This doesn't support the classic thunks interface.
auto AppName = AppConfigName();
std::vector<std::string> ConfigPaths {
FEXCore::Config::GetConfigFileLocation(true),
FEXCore::Config::GetConfigFileLocation(false),
ThunkConfigFile,
FEXCore::Config::GetApplicationConfig(AppConfigName(), true),
FEXCore::Config::GetApplicationConfig(AppConfigName(), false),
};
auto SteamID = getenv("SteamAppId");
if (SteamID) {
// If a SteamID exists then let's search for Steam application configs as well.
// We want to key off both the SteamAppId number /and/ the executable since we may not want to thunk all binaries.
auto SteamAppName = fmt::format("Steam_{}_{}", SteamID, AppName);
// Steam application configs interleaved with non-steam for priority sorting.
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(SteamAppName, true));
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(AppName, true));
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(SteamAppName, false));
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(AppName, false));
}
else {
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(AppName, true));
ConfigPaths.emplace_back(FEXCore::Config::GetApplicationConfig(AppName, false));
}
for (const auto &Path : ConfigPaths) {
std::vector<char> FileData;
if (LoadFile(FileData, Path)) {
@@ -332,7 +355,6 @@ FileManager::~FileManager() {
}
std::string FileManager::GetEmulatedPath(const char *pathname, bool FollowSymlink) {
auto RootFSPath = LDPath();
if (!pathname || // If no pathname
pathname[0] != '/' || // If relative
strcmp(pathname, "/") == 0) { // If we are getting root
@@ -344,17 +366,19 @@ std::string FileManager::GetEmulatedPath(const char *pathname, bool FollowSymlin
return thunkOverlay->second;
}
auto RootFSPath = LDPath();
if (RootFSPath.empty()) { // If RootFS doesn't exist
return {};
}
std::string Path = RootFSPath + pathname;
if (FollowSymlink) {
std::error_code ec;
while(std::filesystem::is_symlink(Path, ec)) {
auto SymlinkTarget = std::filesystem::read_symlink(Path);
if (SymlinkTarget.is_absolute()) {
Path = RootFSPath + SymlinkTarget.string();
char Filename[PATH_MAX];
while(FEX::HLE::IsSymlink(Path)) {
auto SymlinkSize = FEX::HLE::GetSymlink(Path, Filename, PATH_MAX - 1);
if (SymlinkSize > 0 && Filename[0] == '/') {
Path = RootFSPath;
Path += std::string_view(Filename, SymlinkSize);
}
else {
break;
@@ -15,6 +15,7 @@ $end_info$
#include <stddef.h>
#include <string>
#include <sys/stat.h>
#include <unistd.h>
#include <vector>
#include <unordered_map>
@@ -27,6 +28,18 @@ struct Context;
}
namespace FEX::HLE {
[[maybe_unused]]
static bool IsSymlink(const std::string &Filename) {
// Checks to see if a filepath is a symlink.
struct stat Buffer{};
int Result = lstat(Filename.c_str(), &Buffer);
return Result == 0 && S_ISLNK(Buffer.st_mode);
}
[[maybe_unused]]
static ssize_t GetSymlink(const std::string &Filename, char *ResultBuffer, size_t ResultBufferSize) {
return readlink(Filename.c_str(), ResultBuffer, ResultBufferSize);
}
struct open_how;
+356
View File
@@ -15,6 +15,7 @@ extern "C" {
#include "fex-drm/panfrost_drm.h"
#include "fex-drm/msm_drm.h"
#include "fex-drm/nouveau_drm.h"
#include "fex-drm/radeon_drm.h"
#include "fex-drm/vc4_drm.h"
#include "fex-drm/v3d_drm.h"
#include "fex-drm/virtgpu_drm.h"
@@ -713,6 +714,360 @@ fex_drm_amdgpu_gem_metadata {
};
}
namespace RADEON {
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_gem_create")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_gem_create {
compat_uint64_t size;
compat_uint64_t alignment;
__u32 handle;
__u32 initial_domain;
__u32 flags;
fex_drm_radeon_gem_create() = delete;
operator drm_radeon_gem_create() const {
drm_radeon_gem_create val{};
val.size = size;
val.alignment = alignment;
val.handle = handle;
val.initial_domain = initial_domain;
val.flags = flags;
return val;
}
fex_drm_radeon_gem_create(struct drm_radeon_gem_create val) {
size = val.size;
alignment = val.alignment;
handle = val.handle;
initial_domain = val.initial_domain;
flags = val.flags;
}
};
struct
FEX_PACKED
FEX_ANNOTATE("alias-x86_32-drm_radeon_init")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_init_t {
enum {
} func;
compat_ulong_t sarea_priv_offset;
int32_t is_pci;
int32_t cp_mode;
int32_t gart_size;
int32_t ring_size;
int32_t usec_timeout;
uint32_t fb_bpp;
uint32_t front_offset, front_pitch;
uint32_t back_offset, back_pitch;
uint32_t depth_bpp;
uint32_t depth_offset, depth_pitch;
compat_ulong_t fb_offset;
compat_ulong_t mmio_offset;
compat_ulong_t ring_offset;
compat_ulong_t ring_rptr_offset;
compat_ulong_t buffers_offset;
compat_ulong_t gart_textures_offset;
fex_drm_radeon_init_t() = delete;
operator drm_radeon_init_t() const {
drm_radeon_init_t val{};
val.sarea_priv_offset = sarea_priv_offset;
val.is_pci = is_pci;
val.cp_mode = cp_mode;
val.gart_size = gart_size;
val.ring_size = ring_size;
val.usec_timeout = usec_timeout;
val.fb_bpp = fb_bpp;
val.front_offset = front_offset;
val.front_pitch = front_pitch;
val.back_offset = back_offset;
val.back_pitch = back_pitch;
val.depth_bpp = depth_bpp;
val.depth_offset = depth_offset;
val.depth_pitch = depth_pitch;
val.fb_offset = fb_offset;
val.mmio_offset = mmio_offset;
val.ring_offset = ring_offset;
val.ring_rptr_offset = ring_rptr_offset;
val.buffers_offset = buffers_offset;
val.gart_textures_offset = gart_textures_offset;
return val;
}
fex_drm_radeon_init_t(drm_radeon_init_t val) {
sarea_priv_offset = val.sarea_priv_offset;
is_pci = val.is_pci;
cp_mode = val.cp_mode;
gart_size = val.gart_size;
ring_size = val.ring_size;
usec_timeout = val.usec_timeout;
fb_bpp = val.fb_bpp;
front_offset = val.front_offset;
front_pitch = val.front_pitch;
back_offset = val.back_offset;
back_pitch = val.back_pitch;
depth_bpp = val.depth_bpp;
depth_offset = val.depth_offset;
depth_pitch = val.depth_pitch;
fb_offset = val.fb_offset;
mmio_offset = val.mmio_offset;
ring_offset = val.ring_offset;
ring_rptr_offset = val.ring_rptr_offset;
buffers_offset = val.buffers_offset;
gart_textures_offset = val.gart_textures_offset;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_clear")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_clear_t {
uint32_t flags;
uint32_t clear_color;
uint32_t clear_depth;
uint32_t color_mask;
uint32_t depth_mask;
compat_ptr<drm_radeon_clear_rect_t> depth_boxes;
fex_drm_radeon_clear_t() = delete;
operator drm_radeon_clear_t() const {
drm_radeon_clear_t val{};
val.flags = flags;
val.clear_color = clear_color;
val.clear_depth = clear_depth;
val.color_mask = color_mask;
val.depth_mask = depth_mask;
val.depth_boxes = depth_boxes;
return val;
}
fex_drm_radeon_clear_t(drm_radeon_clear_t val)
: depth_boxes {val.depth_boxes} {
flags = val.flags;
clear_color = val.clear_color;
clear_depth = val.clear_depth;
color_mask = val.color_mask;
depth_mask = val.depth_mask;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_stipple")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_stipple_t {
compat_ptr<uint32_t> mask;
fex_drm_radeon_stipple_t() = delete;
operator drm_radeon_stipple_t() const {
drm_radeon_stipple_t val{};
val.mask = mask;
return val;
}
fex_drm_radeon_stipple_t(drm_radeon_stipple_t val)
: mask {val.mask} {
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_texture")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_texture_t {
uint32_t offset;
int32_t pitch;
int32_t format;
int32_t width;
int32_t height;
compat_ptr<drm_radeon_tex_image_t> image;
fex_drm_radeon_texture_t() = delete;
operator drm_radeon_texture_t() const {
drm_radeon_texture_t val{};
val.offset = offset;
val.pitch = pitch;
val.format = format;
val.width = width;
val.height = height;
val.image = image;
return val;
}
fex_drm_radeon_texture_t(drm_radeon_texture_t val)
: image {val.image} {
offset = val.offset;
pitch = val.pitch;
format = val.format;
width = val.width;
height = val.height;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_vertex2")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_vertex2_t {
int32_t idx;
int32_t discard;
int32_t nr_states;
compat_ptr<drm_radeon_state_t> state;
int32_t nr_prims;
compat_ptr<drm_radeon_prim_t> prim;
fex_drm_radeon_vertex2_t() = delete;
operator drm_radeon_vertex2_t() const {
drm_radeon_vertex2_t val;
val.idx = idx;
val.discard = discard;
val.nr_states = nr_states;
val.state = state;
val.nr_prims = nr_prims;
val.prim = prim;
return val;
}
fex_drm_radeon_vertex2_t(drm_radeon_vertex2_t val)
: state {val.state}
, prim {val.prim} {
idx = val.idx;
discard = val.discard;
nr_states = val.nr_states;
nr_prims = val.nr_prims;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_cmd_buffer")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_cmd_buffer_t {
int32_t bufsz;
compat_ptr<char> buf;
int32_t nbox;
compat_ptr<drm_clip_rect> boxes;
fex_drm_radeon_cmd_buffer_t() = delete;
operator drm_radeon_cmd_buffer_t() const {
drm_radeon_cmd_buffer_t val;
val.bufsz = bufsz;
val.buf = buf;
val.nbox = nbox;
val.boxes = boxes;
return val;
}
fex_drm_radeon_cmd_buffer_t(drm_radeon_cmd_buffer_t val)
: buf {val.buf}
, boxes {val.boxes} {
val.bufsz = bufsz;
val.nbox = nbox;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_getparam")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_getparam_t {
int32_t param;
compat_ptr<void> value;
fex_drm_radeon_getparam_t() = delete;
operator drm_radeon_getparam_t() const {
drm_radeon_getparam_t val;
val.param = param;
val.value = value;
return val;
}
fex_drm_radeon_getparam_t(drm_radeon_getparam_t val)
: value {val.value} {
val.param = param;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_mem_alloc")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_mem_alloc_t {
int32_t region;
int32_t alignment;
int32_t size;
compat_ptr<int32_t> region_offset;
fex_drm_radeon_mem_alloc_t() = delete;
operator drm_radeon_mem_alloc_t() const {
drm_radeon_mem_alloc_t val;
val.region = region;
val.alignment = alignment;
val.size = size;
val.region_offset = region_offset;
return val;
}
fex_drm_radeon_mem_alloc_t(drm_radeon_mem_alloc_t val)
: region_offset {val.region_offset} {
val.region = region;
val.alignment = alignment;
val.size = size;
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_irq_emit")
FEX_ANNOTATE("fex-match")
fex_drm_radeon_irq_emit_t {
compat_ptr<int32_t> irq_seq;
fex_drm_radeon_irq_emit_t() = delete;
operator drm_radeon_irq_emit_t() const {
drm_radeon_irq_emit_t val;
val.irq_seq = irq_seq;
return val;
}
fex_drm_radeon_irq_emit_t(drm_radeon_irq_emit_t val)
: irq_seq {val.irq_seq} {
}
};
struct
FEX_ANNOTATE("alias-x86_32-drm_radeon_setparam")
FEX_ANNOTATE("fex-match")
FEX_PACKED
fex_drm_radeon_setparam_t {
uint32_t param;
compat_int64_t value;
fex_drm_radeon_setparam_t() = delete;
operator drm_radeon_setparam_t() const {
drm_radeon_setparam_t val;
val.param = param;
val.value = value;
return val;
}
fex_drm_radeon_setparam_t(drm_radeon_setparam_t val) {
param = val.param;
value = val.value;
}
};
}
namespace MSM {
struct
FEX_ANNOTATE("alias-x86_32-drm_msm_timespec")
@@ -1061,6 +1416,7 @@ fex_drm_v3d_submit_csd {
#include "Tests/LinuxSyscalls/x32/Ioctl/lima_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/panfrost_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/nouveau_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/radeon_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/vc4_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/v3d_drm.inl"
@@ -0,0 +1,44 @@
_CUSTOM_META(DRM_IOCTL_RADEON_CP_INIT, DRM_IOW(DRM_COMMAND_BASE + DRM_RADEON_CP_INIT, FEX::HLE::x32::RADEON::fex_drm_radeon_init_t))
_BASIC_META(DRM_IOCTL_RADEON_CP_START)
_BASIC_META(DRM_IOCTL_RADEON_CP_STOP)
_BASIC_META(DRM_IOCTL_RADEON_CP_RESET)
_BASIC_META(DRM_IOCTL_RADEON_CP_IDLE)
_BASIC_META(DRM_IOCTL_RADEON_RESET)
_BASIC_META(DRM_IOCTL_RADEON_FULLSCREEN)
_BASIC_META(DRM_IOCTL_RADEON_SWAP)
_CUSTOM_META(DRM_IOCTL_RADEON_CLEAR, DRM_IOW(DRM_COMMAND_BASE + DRM_RADEON_CLEAR, FEX::HLE::x32::RADEON::fex_drm_radeon_clear_t))
_BASIC_META(DRM_IOCTL_RADEON_VERTEX)
_BASIC_META(DRM_IOCTL_RADEON_INDICES)
_CUSTOM_META(DRM_IOCTL_RADEON_STIPPLE, DRM_IOW( DRM_COMMAND_BASE + DRM_RADEON_STIPPLE, FEX::HLE::x32::RADEON::fex_drm_radeon_stipple_t))
_BASIC_META(DRM_IOCTL_RADEON_INDIRECT)
_CUSTOM_META(DRM_IOCTL_RADEON_TEXTURE, DRM_IOWR(DRM_COMMAND_BASE + DRM_RADEON_TEXTURE, FEX::HLE::x32::RADEON::fex_drm_radeon_texture_t))
_CUSTOM_META(DRM_IOCTL_RADEON_VERTEX2, DRM_IOW(DRM_COMMAND_BASE + DRM_RADEON_VERTEX2, FEX::HLE::x32::RADEON::fex_drm_radeon_vertex2_t))
_CUSTOM_META(DRM_IOCTL_RADEON_CMDBUF, DRM_IOW(DRM_COMMAND_BASE + DRM_RADEON_CMDBUF, FEX::HLE::x32::RADEON::fex_drm_radeon_cmd_buffer_t))
_CUSTOM_META(DRM_IOCTL_RADEON_GETPARAM, DRM_IOWR(DRM_COMMAND_BASE + DRM_RADEON_GETPARAM, FEX::HLE::x32::RADEON::fex_drm_radeon_getparam_t))
_BASIC_META(DRM_IOCTL_RADEON_FLIP)
_CUSTOM_META(DRM_IOCTL_RADEON_ALLOC, DRM_IOWR(DRM_COMMAND_BASE + DRM_RADEON_ALLOC, FEX::HLE::x32::RADEON::fex_drm_radeon_mem_alloc_t))
_BASIC_META(DRM_IOCTL_RADEON_FREE)
_BASIC_META(DRM_IOCTL_RADEON_INIT_HEAP)
_CUSTOM_META(DRM_IOCTL_RADEON_IRQ_EMIT, DRM_IOWR(DRM_COMMAND_BASE + DRM_RADEON_IRQ_EMIT, FEX::HLE::x32::RADEON::fex_drm_radeon_irq_emit_t))
_BASIC_META(DRM_IOCTL_RADEON_IRQ_WAIT)
_BASIC_META(DRM_IOCTL_RADEON_CP_RESUME)
_CUSTOM_META(DRM_IOCTL_RADEON_SETPARAM, DRM_IOW(DRM_COMMAND_BASE + DRM_RADEON_SETPARAM, FEX::HLE::x32::RADEON::fex_drm_radeon_setparam_t))
_BASIC_META(DRM_IOCTL_RADEON_SURF_ALLOC)
_BASIC_META(DRM_IOCTL_RADEON_SURF_FREE)
_BASIC_META(DRM_IOCTL_RADEON_GEM_INFO)
_CUSTOM_META(DRM_IOCTL_RADEON_GEM_CREATE, DRM_IOWR(DRM_COMMAND_BASE + DRM_RADEON_GEM_CREATE, FEX::HLE::x32::RADEON::fex_drm_radeon_gem_create))
_BASIC_META(DRM_IOCTL_RADEON_GEM_MMAP)
_BASIC_META(DRM_IOCTL_RADEON_GEM_PREAD)
_BASIC_META(DRM_IOCTL_RADEON_GEM_PWRITE)
_BASIC_META(DRM_IOCTL_RADEON_GEM_SET_DOMAIN)
_BASIC_META(DRM_IOCTL_RADEON_GEM_WAIT_IDLE)
_BASIC_META(DRM_IOCTL_RADEON_CS)
_BASIC_META(DRM_IOCTL_RADEON_INFO)
_BASIC_META(DRM_IOCTL_RADEON_GEM_SET_TILING)
_BASIC_META(DRM_IOCTL_RADEON_GEM_GET_TILING)
_BASIC_META(DRM_IOCTL_RADEON_GEM_BUSY)
_BASIC_META(DRM_IOCTL_RADEON_GEM_VA)
_BASIC_META(DRM_IOCTL_RADEON_GEM_OP)
_BASIC_META(DRM_IOCTL_RADEON_GEM_USERPTR)
@@ -178,6 +178,141 @@ namespace FEX::HLE::x32 {
return -EPERM;
}
uint32_t RADEON_Handler(int fd, uint32_t cmd, uint32_t args) {
switch (_IOC_NR(cmd)) {
case _IOC_NR(FEX_DRM_IOCTL_RADEON_CP_INIT): {
RADEON::fex_drm_radeon_init_t *val = reinterpret_cast<RADEON::fex_drm_radeon_init_t*>(args);
drm_radeon_init_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_CP_INIT, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_CLEAR): {
RADEON::fex_drm_radeon_clear_t *val = reinterpret_cast<RADEON::fex_drm_radeon_clear_t*>(args);
drm_radeon_clear_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_CLEAR, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_STIPPLE): {
RADEON::fex_drm_radeon_stipple_t *val = reinterpret_cast<RADEON::fex_drm_radeon_stipple_t*>(args);
drm_radeon_stipple_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_STIPPLE, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_TEXTURE): {
RADEON::fex_drm_radeon_texture_t *val = reinterpret_cast<RADEON::fex_drm_radeon_texture_t*>(args);
drm_radeon_texture_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_TEXTURE, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_VERTEX2): {
RADEON::fex_drm_radeon_vertex2_t *val = reinterpret_cast<RADEON::fex_drm_radeon_vertex2_t*>(args);
drm_radeon_vertex2_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_VERTEX2, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_CMDBUF): {
RADEON::fex_drm_radeon_cmd_buffer_t *val = reinterpret_cast<RADEON::fex_drm_radeon_cmd_buffer_t*>(args);
drm_radeon_cmd_buffer_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_CMDBUF, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_GETPARAM): {
RADEON::fex_drm_radeon_getparam_t *val = reinterpret_cast<RADEON::fex_drm_radeon_getparam_t*>(args);
drm_radeon_getparam_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_GETPARAM, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_ALLOC): {
RADEON::fex_drm_radeon_mem_alloc_t *val = reinterpret_cast<RADEON::fex_drm_radeon_mem_alloc_t*>(args);
drm_radeon_mem_alloc_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_ALLOC, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_IRQ_EMIT): {
RADEON::fex_drm_radeon_irq_emit_t *val = reinterpret_cast<RADEON::fex_drm_radeon_irq_emit_t*>(args);
drm_radeon_irq_emit_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_IRQ_EMIT, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_SETPARAM): {
RADEON::fex_drm_radeon_setparam_t *val = reinterpret_cast<RADEON::fex_drm_radeon_setparam_t*>(args);
drm_radeon_setparam_t Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_SETPARAM, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
case _IOC_NR(FEX_DRM_IOCTL_RADEON_GEM_CREATE): {
RADEON::fex_drm_radeon_gem_create *val = reinterpret_cast<RADEON::fex_drm_radeon_gem_create*>(args);
drm_radeon_gem_create Host_val = *val;
uint64_t Result = ioctl(fd, DRM_IOCTL_RADEON_GEM_CREATE, &Host_val);
if (Result != -1) {
*val = Host_val;
}
SYSCALL_ERRNO();
break;
}
#define _BASIC_META(x) case _IOC_NR(x):
#define _BASIC_META_VAR(x, args...) case _IOC_NR(x):
#define _CUSTOM_META(name, ioctl_num)
#define _CUSTOM_META_OFFSET(name, ioctl_num, offset)
// DRM
#include "Tests/LinuxSyscalls/x32/Ioctl/radeon_drm.inl"
{
uint64_t Result = ::ioctl(fd, cmd, args);
SYSCALL_ERRNO();
break;
}
default:
UnhandledIoctl("RADEON", fd, cmd, args);
return -EPERM;
break;
}
#undef _BASIC_META
#undef _BASIC_META_VAR
#undef _CUSTOM_META
#undef _CUSTOM_META_OFFSET
return -EPERM;
}
uint32_t MSM_Handler(int fd, uint32_t cmd, uint32_t args) {
switch (_IOC_NR(cmd)) {
case _IOC_NR(FEX_DRM_IOCTL_MSM_WAIT_FENCE): {
@@ -435,6 +570,9 @@ namespace FEX::HLE::x32 {
if (strcmp(Version.name, "amdgpu") == 0) {
FDToHandler.SetFDHandler(fd, AMDGPU_Handler);
}
else if (strcmp(Version.name, "radeon") == 0) {
FDToHandler.SetFDHandler(fd, RADEON_Handler);
}
else if (strcmp(Version.name, "msm") == 0) {
FDToHandler.SetFDHandler(fd, MSM_Handler);
}
@@ -611,6 +749,7 @@ namespace FEX::HLE::x32 {
#include "Tests/LinuxSyscalls/x32/Ioctl/lima_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/panfrost_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/nouveau_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/radeon_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/vc4_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/v3d_drm.inl"
#include "Tests/LinuxSyscalls/x32/Ioctl/virtio_drm.inl"
+4 -6
View File
@@ -52,8 +52,6 @@ namespace FEX::HLE::x32 {
}
void x32SyscallHandler::RegisterSyscallHandlers() {
Definitions.resize(FEX::HLE::x32::SYSCALL_x86_MAX);
auto cvt = [](auto in) {
union {
decltype(in) val;
@@ -63,10 +61,10 @@ namespace FEX::HLE::x32 {
return raw.raw;
};
for (auto &Def : Definitions) {
Def.NumArgs = 255;
Def.Ptr = cvt(&UnimplementedSyscall);
}
Definitions.resize(FEX::HLE::x32::SYSCALL_x86_MAX, SyscallFunctionDefinition {
.NumArgs = 255,
.Ptr = cvt(&UnimplementedSyscall),
});
FEX::HLE::RegisterEpoll(this);
FEX::HLE::RegisterFD(this);
+4 -6
View File
@@ -36,7 +36,6 @@ namespace FEX::HLE::x64 {
}
void x64SyscallHandler::RegisterSyscallHandlers() {
Definitions.resize(FEX::HLE::x64::SYSCALL_x64_MAX);
auto cvt = [](auto in) {
union {
decltype(in) val;
@@ -46,11 +45,10 @@ namespace FEX::HLE::x64 {
return raw.raw;
};
// Clear all definitions
for (auto &Def : Definitions) {
Def.NumArgs = 255;
Def.Ptr = cvt(&UnimplementedSyscall);
}
Definitions.resize(FEX::HLE::x64::SYSCALL_x64_MAX, SyscallFunctionDefinition {
.NumArgs = 255,
.Ptr = cvt(&UnimplementedSyscall),
});
FEX::HLE::RegisterEpoll(this);
FEX::HLE::RegisterFD(this);
+16
View File
@@ -8,6 +8,7 @@
#include <FEXHeaderUtils/Syscalls.h>
#include <dlfcn.h>
#include <elf.h>
#include <fcntl.h>
#include <filesystem>
#include <sys/mman.h>
@@ -291,6 +292,21 @@ namespace FEX::VDSO {
return VDSOBase;
}
uint64_t GetVSyscallEntry(const void* VDSOBase) {
if (!VDSOBase) {
return 0;
}
// Extract the vsyscall location from the VDSO header.
auto Header = reinterpret_cast<const Elf32_Ehdr*>(VDSOBase);
if (Header->e_entry) {
return reinterpret_cast<uint64_t>(VDSOBase) + Header->e_entry;
}
return 0;
}
std::vector<FEXCore::IR::ThunkDefinition> const& GetVDSOThunkDefinitions() {
return VDSODefinitions;
}
+2
View File
@@ -5,5 +5,7 @@ namespace FEX::VDSO {
using MapperFn = std::function<void *(void *addr, size_t length, int prot, int flags, int fd, off_t offset)>;
void* LoadVDSOThunks(bool Is64Bit, MapperFn Mapper);
uint64_t GetVSyscallEntry(const void* VDSOBase);
std::vector<FEXCore::IR::ThunkDefinition> const& GetVDSOThunkDefinitions();
}
+1
View File
@@ -141,6 +141,7 @@ namespace {
}
}
}
std::sort(NamedRootFS.begin(), NamedRootFS.end());
}
void INotifyThread() {
+11 -2
View File
@@ -40,8 +40,17 @@ int main(int argc, char **argv, char **envp) {
if (Options.is_set_by_user("app")) {
// Load the application config if one was provided
auto ProgramName = std::filesystem::path(Options["app"]).filename();
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, true));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, false));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, FEXCore::Config::LayerType::LAYER_GLOBAL_APP));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(ProgramName, FEXCore::Config::LayerType::LAYER_LOCAL_APP));
auto SteamID = getenv("SteamAppId");
if (SteamID) {
// If a SteamID exists then let's search for Steam application configs as well.
// We want to key off both the SteamAppId number /and/ the executable since we may not want to thunk all binaries.
auto SteamAppName = fmt::format("Steam_{}_{}", SteamID, ProgramName.string());
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(SteamAppName, FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP));
FEXCore::Config::AddLayer(FEXCore::Config::CreateAppLayer(SteamAppName, FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP));
}
}
// Reload the meta layer
+6
View File
@@ -109,6 +109,12 @@ namespace {
* @brief Deparents itself by forking and terminating the parent process.
*/
void DeparentSelf() {
auto SystemdEnv = getenv("SYSTEMD_EXEC_PID");
if (SystemdEnv) {
// If FEXServer was launched through systemd then don't deparent, otherwise systemd kills the entire server.
return;
}
pid_t pid = fork();
if (pid != 0) {
+12 -13
View File
@@ -83,7 +83,7 @@ namespace ProcessPipe {
if (!std::filesystem::exists(ServerFolder, ec)) {
// Doesn't exist, create the the folder as a user convenience
if (!std::filesystem::create_directories(ServerFolder, ec)) {
LogMan::Msg::AFmt("Couldn't create server pipe folder at: {}", ServerFolder);
LogMan::Msg::EFmt("Couldn't create server pipe folder at: {}", ServerFolder);
return false;
}
}
@@ -137,7 +137,7 @@ namespace ProcessPipe {
}
else if (Ret == -1) {
// Unhandled error.
LogMan::Msg::AFmt("Unable to create FEXServer named lock file at: {} {} {}", ServerLockPath, errno, strerror(errno));
LogMan::Msg::EFmt("Unable to create FEXServer named lock file at: {} {} {}", ServerLockPath, errno, strerror(errno));
return false;
}
else {
@@ -169,7 +169,7 @@ namespace ProcessPipe {
if (Ret == -1) {
// This shouldn't occur
LogMan::Msg::AFmt("Unable to downgrade a write lock to a read lock {} {} {}", ServerLockPath, errno, strerror(errno));
LogMan::Msg::EFmt("Unable to downgrade a write lock to a read lock {} {} {}", ServerLockPath, errno, strerror(errno));
close(ServerLockFD);
ServerLockFD = -1;
return false;
@@ -179,28 +179,27 @@ namespace ProcessPipe {
}
bool InitializeServerSocket() {
auto ServerSocketFile = FEXServerClient::GetServerSocketFile();
// Unlink the socket file if it exists
// We are being asked to create a daemon, not error check
// We don't care if this failed or not
unlink(ServerSocketFile.c_str());
auto ServerSocketName = FEXServerClient::GetServerSocketName();
// Create the initial unix socket
ServerSocketFD = socket(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0);
if (ServerSocketFD == -1) {
LogMan::Msg::AFmt("Couldn't create AF_UNIX socket: %d %s\n", errno, strerror(errno));
LogMan::Msg::EFmt("Couldn't create AF_UNIX socket: {} {}\n", errno, strerror(errno));
return false;
}
struct sockaddr_un addr{};
addr.sun_family = AF_UNIX;
strncpy(addr.sun_path, ServerSocketFile.data(), std::min(ServerSocketFile.size(), sizeof(addr.sun_path)));
size_t SizeOfSocketString = std::min(ServerSocketName.size() + 1, sizeof(addr.sun_path) - 1);
addr.sun_path[0] = 0; // Abstract AF_UNIX sockets start with \0
strncpy(addr.sun_path + 1, ServerSocketName.data(), SizeOfSocketString);
// Include final null character.
size_t SizeOfAddr = sizeof(addr.sun_family) + SizeOfSocketString;
// Bind the socket to the path
int Result = bind(ServerSocketFD, reinterpret_cast<struct sockaddr*>(&addr), sizeof(addr));
int Result = bind(ServerSocketFD, reinterpret_cast<struct sockaddr*>(&addr), SizeOfAddr);
if (Result == -1) {
LogMan::Msg::AFmt("Couldn't bind AF_UNIX socket '%s': %d %s\n", addr.sun_path, errno, strerror(errno));
LogMan::Msg::EFmt("Couldn't bind AF_UNIX socket '{}': {} {}\n", addr.sun_path, errno, strerror(errno));
close(ServerSocketFD);
ServerSocketFD = -1;
return false;
+11 -1
View File
@@ -14,8 +14,15 @@ namespace SquashFS {
constexpr int USER_PERMS = S_IRWXU | S_IRWXG | S_IRWXO;
int ServerRootFSLockFD {-1};
int FuseMountPID{};
std::string MountFolder{};
void ShutdownImagePID() {
if (FuseMountPID) {
tgkill(FuseMountPID, FuseMountPID, SIGINT);
}
}
bool InitializeSquashFSPipe() {
std::string RootFSLockFile = FEXServerClient::GetServerRootFSLockFile();
@@ -99,7 +106,7 @@ namespace SquashFS {
bool MountRootFSImagePath(std::string SquashFS, bool EroFS) {
pid_t ParentTID = ::getpid();
MountFolder = fmt::format("{}/.FEXMount{}-XXXXXX", std::filesystem::temp_directory_path().string(), ParentTID);
MountFolder = fmt::format("{}/.FEXMount{}-XXXXXX", FEXServerClient::GetServerMountFolder(), ParentTID);
char *MountFolderStr = MountFolder.data();
// Make the temporary mount folder
@@ -145,6 +152,7 @@ namespace SquashFS {
}
}
else {
FuseMountPID = pid;
// Parent
// Wait for the child to exit
// This will happen with execvpe of squashmount or exit on failure
@@ -186,6 +194,8 @@ namespace SquashFS {
return;
}
SquashFS::ShutdownImagePID();
// Handle final mount removal
// fusermount for unmounting the mountpoint, then the {erfsfuse, squashfuse} will exit automatically
int pid = fork();
+1 -1
View File
@@ -1,7 +1,6 @@
cmake_minimum_required(VERSION 3.14)
project(guest-thunks)
option(BITNESS "Which bitness the thunks are building for" 64)
option(ENABLE_CLANG_THUNKS "Enable building thunks with clang" FALSE)
if (ENABLE_CLANG_THUNKS)
@@ -34,6 +33,7 @@ else()
set(GENERATOR_EXE thunkgen)
set(TARGET_TYPE OBJECT)
set(GENERATE_GUEST_INSTALL_TARGETS FALSE)
set(BITNESS 64)
endif()
# Syntax: generate(libxyz libxyz-interface.cpp)
+6
View File
@@ -40,6 +40,12 @@ Follow the steps in: https://github.com/FEX-Emu/FEX-ppa/blob/main/README.md
* Requires PPA GPG key signing access
* Wait the 20-30 minutes for Ubuntu PPA to build and publish the binaries
## ArchLinux AUR package
* Clone https://aur.archlinux.org/packages/fex-emu
* Requires maintainer or co-maintainer permissions
* Update PKGBUILD and .SRCINFO file.
* Push changes upstream
## Github releases page Steps
* Requires administrative rights
* Go to https://github.com/FEX-Emu/FEX/releases
+1 -2
View File
@@ -1,4 +1,4 @@
# FEX-2211
# FEX-2212
## External/FEXCore
See [FEXCore/Readme.md](../External/FEXCore/Readme.md) for more details
@@ -148,7 +148,6 @@ IR to IR Optimization
- [RedundantFlagCalculationElimination.cpp](../External/FEXCore/Source/Interface/IR/Passes/RedundantFlagCalculationElimination.cpp): This is not used right now, possibly broken
- [RegisterAllocationPass.cpp](../External/FEXCore/Source/Interface/IR/Passes/RegisterAllocationPass.cpp)
- [RegisterAllocationPass.h](../External/FEXCore/Source/Interface/IR/Passes/RegisterAllocationPass.h)
- [StaticRegisterAllocationPass.cpp](../External/FEXCore/Source/Interface/IR/Passes/StaticRegisterAllocationPass.cpp): Replaces Load/StoreContext with Load/StoreReg for SRA regs
- [SyscallOptimization.cpp](../External/FEXCore/Source/Interface/IR/Passes/SyscallOptimization.cpp): Removes unused arguments if known syscall number
- [ValueDominanceValidation.cpp](../External/FEXCore/Source/Interface/IR/Passes/ValueDominanceValidation.cpp): Sanity Checking
+3
View File
@@ -10,3 +10,6 @@ Test_32Bit_Primary/Primary_8C_2.asm
# Hecks with CS
# Causing our host runner a bit of pain
Test_32Bit_Primary/Primary_CF.asm
# Zen+ CI doesn't support UMIP so it returns "real" values
Test_32Bit_Secondary/07_XX_00.asm
@@ -0,0 +1,21 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0",
"RBX": "0x00000000FFFE0000"
},
"Mode": "32BIT"
}
%endif
sgdt [rel data]
movzx eax, word [rel data]
mov ebx, dword [rel data + 2]
hlt
data:
; Limit
dw 0
; Base
dd 0
+3
View File
@@ -12,3 +12,6 @@ Test_Secondary/15_F3_02.asm
Test_Secondary/15_F3_03.asm
Test_Secondary/15_F3_02_2.asm
Test_Secondary/15_F3_03_2.asm
# Zen+ CI doesn't support UMIP so it returns "real" values
Test_Secondary/07_XX_00.asm
+20
View File
@@ -0,0 +1,20 @@
%ifdef CONFIG
{
"RegData": {
"RAX": "0",
"RBX": "0xFFFFFFFFFFFE0000"
}
}
%endif
sgdt [rel data]
movzx rax, word [rel data]
mov rbx, qword [rel data + 2]
hlt
data:
; Limit
dw 0
; Base
dq 0
+42
View File
@@ -0,0 +1,42 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x3FF0000000000000", "0x4000000000000000", "0x4008000000000000", "0x4010000000000000"],
"XMM1": ["0x4014000000000000", "0x4018000000000000", "0x401C000000000000", "0x4020000000000000"],
"XMM2": ["0x4018000000000000", "0x4020000000000000", "0x4024000000000000", "0x4028000000000000"],
"XMM3": ["0x4018000000000000", "0x4020000000000000", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x4018000000000000", "0x4020000000000000", "0x4024000000000000", "0x4028000000000000"],
"XMM5": ["0x4018000000000000", "0x4020000000000000", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data]
; Registers
vmovapd ymm0, [rdx]
vmovapd ymm1, [rdx + 32]
vaddpd ymm2, ymm0, ymm1
vaddpd xmm3, xmm0, xmm1
; Memory operand
vaddpd ymm4, ymm0, [rdx + 32]
vaddpd xmm5, xmm0, [rdx + 32]
hlt
align 32
.data:
dq 0x3FF0000000000000 ; 1.0
dq 0x4000000000000000 ; 2.0
dq 0x4008000000000000 ; 3.0
dq 0x4010000000000000 ; 4.0
dq 0x4014000000000000 ; 5.0
dq 0x4018000000000000 ; 6.0
dq 0x401C000000000000 ; 7.0
dq 0x4020000000000000 ; 8.0
+42
View File
@@ -0,0 +1,42 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x400000003F800000", "0x4080000040400000", "0x40C0000040A00000", "0x4100000040E00000"],
"XMM1": ["0x4100000040E00000", "0x40C0000040A00000", "0x4080000040400000", "0x400000003F800000"],
"XMM2": ["0x4120000041000000", "0x4120000041000000", "0x4120000041000000", "0x4120000041000000"],
"XMM3": ["0x4120000041000000", "0x4120000041000000", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x4120000041000000", "0x4120000041000000", "0x4120000041000000", "0x4120000041000000"],
"XMM5": ["0x4120000041000000", "0x4120000041000000", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data]
; Registers
vmovapd ymm0, [rdx]
vmovapd ymm1, [rdx + 32]
vaddps ymm2, ymm0, ymm1
vaddps xmm3, xmm0, xmm1
; Memory operand
vaddps ymm4, ymm0, [rdx + 32]
vaddps xmm5, xmm0, [rdx + 32]
hlt
align 32
.data:
dq 0x400000003F800000 ; 2.0, 1.0
dq 0x4080000040400000 ; 4.0, 3.0
dq 0x40C0000040A00000 ; 6.0, 5.0
dq 0x4100000040E00000 ; 8.0, 7.0
dq 0x4100000040E00000 ; 8.0, 7.0
dq 0x40C0000040A00000 ; 6.0, 5.0
dq 0x4080000040400000 ; 4.0, 3.0
dq 0x400000003F800000 ; 2.0, 1.0
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0xCCCCCCCC75767778", "0x61626364DDDDDDDD", "0xEEEEEEEE55565758", "0x41424344FFFFFFFF"],
"XMM2": ["0x8C8C8C8830303030", "0x2020202088898885", "0x8E8C8C8A10101010", "0x000000008A898887"],
"XMM3": ["0x8C8C8C8830303030", "0x2020202088898885", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x8C8C8C8830303030", "0x2020202088898885", "0x8E8C8C8A10101010", "0x000000008A898887"],
"XMM5": ["0x8C8C8C8830303030", "0x2020202088898885", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data1]
lea rbx, [rel .data2]
vmovapd ymm0, [rdx]
vmovapd ymm1, [rbx]
; Register only
vandnpd ymm2, ymm0, ymm1
vandnpd xmm3, xmm0, xmm1
; With memory operand
vandnpd ymm4, ymm0, [rbx]
vandnpd xmm5, xmm0, [rbx]
hlt
align 32
.data1:
dq 0x4142434445464748
dq 0x5152535455565758
dq 0x6162636465666768
dq 0x7172737475767778
.data2:
dq 0xCCCCCCCC75767778
dq 0x61626364DDDDDDDD
dq 0xEEEEEEEE55565758
dq 0x41424344FFFFFFFF
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0xCCCCCCCC75767778", "0x61626364DDDDDDDD", "0xEEEEEEEE55565758", "0x41424344FFFFFFFF"],
"XMM2": ["0x8C8C8C8830303030", "0x2020202088898885", "0x8E8C8C8A10101010", "0x000000008A898887"],
"XMM3": ["0x8C8C8C8830303030", "0x2020202088898885", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x8C8C8C8830303030", "0x2020202088898885", "0x8E8C8C8A10101010", "0x000000008A898887"],
"XMM5": ["0x8C8C8C8830303030", "0x2020202088898885", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data1]
lea rbx, [rel .data2]
vmovapd ymm0, [rdx]
vmovapd ymm1, [rbx]
; Register only
vandnps ymm2, ymm0, ymm1
vandnps xmm3, xmm0, xmm1
; With memory operand
vandnps ymm4, ymm0, [rbx]
vandnps xmm5, xmm0, [rbx]
hlt
align 32
.data1:
dq 0x4142434445464748
dq 0x5152535455565758
dq 0x6162636465666768
dq 0x7172737475767778
.data2:
dq 0xCCCCCCCC75767778
dq 0x61626364DDDDDDDD
dq 0xEEEEEEEE55565758
dq 0x41424344FFFFFFFF
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0xCCCCCCCC75767778", "0x61626364DDDDDDDD", "0xEEEEEEEE55565758", "0x41424344FFFFFFFF"],
"XMM2": ["0x4040404445464748", "0x4142434455545558", "0x6062626445464748", "0x4142434475767778"],
"XMM3": ["0x4040404445464748", "0x4142434455545558", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x4040404445464748", "0x4142434455545558", "0x6062626445464748", "0x4142434475767778"],
"XMM5": ["0x4040404445464748", "0x4142434455545558", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data1]
lea rbx, [rel .data2]
vmovapd ymm0, [rdx]
vmovapd ymm1, [rbx]
; Register only
vandpd ymm2, ymm0, ymm1
vandpd xmm3, xmm0, xmm1
; With memory operand
vandpd ymm4, ymm0, [rbx]
vandpd xmm5, xmm0, [rbx]
hlt
align 32
.data1:
dq 0x4142434445464748
dq 0x5152535455565758
dq 0x6162636465666768
dq 0x7172737475767778
.data2:
dq 0xCCCCCCCC75767778
dq 0x61626364DDDDDDDD
dq 0xEEEEEEEE55565758
dq 0x41424344FFFFFFFF
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0xCCCCCCCC75767778", "0x61626364DDDDDDDD", "0xEEEEEEEE55565758", "0x41424344FFFFFFFF"],
"XMM2": ["0x4040404445464748", "0x4142434455545558", "0x6062626445464748", "0x4142434475767778"],
"XMM3": ["0x4040404445464748", "0x4142434455545558", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0x4040404445464748", "0x4142434455545558", "0x6062626445464748", "0x4142434475767778"],
"XMM5": ["0x4040404445464748", "0x4142434455545558", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
lea rdx, [rel .data1]
lea rbx, [rel .data2]
vmovapd ymm0, [rdx]
vmovapd ymm1, [rbx]
; Register only
vandps ymm2, ymm0, ymm1
vandps xmm3, xmm0, xmm1
; With memory operand
vandps ymm4, ymm0, [rbx]
vandps xmm5, xmm0, [rbx]
hlt
align 32
.data1:
dq 0x4142434445464748
dq 0x5152535455565758
dq 0x6162636465666768
dq 0x7172737475767778
.data2:
dq 0xCCCCCCCC75767778
dq 0x61626364DDDDDDDD
dq 0xEEEEEEEE55565758
dq 0x41424344FFFFFFFF
+26
View File
@@ -0,0 +1,26 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM2": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC"]
}
}
%endif
lea rdx, [rel .data]
; Load inputs
vmovapd ymm1, [rdx]
vmovapd xmm2, [rdx]
vmovapd ymm3, [rdx + 32]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
+49
View File
@@ -0,0 +1,49 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"RAX": "0x6162636465666768",
"RBX": "0x7172737475767778",
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
mov rdx, 0xe0000000
mov rax, 0x4142434445464748
mov [rdx + 8 * 0], rax
mov rax, 0x5152535455565758
mov [rdx + 8 * 1], rax
mov rax, 0x6162636465666768
mov [rdx + 8 * 2], rax
mov rax, 0x7172737475767778
mov [rdx + 8 * 3], rax
mov rax, 0xCCCCCCCCCCCCCCCC
mov [rdx + 8 * 4], rax
mov rax, 0xDDDDDDDDDDDDDDDD
mov [rdx + 8 * 5], rax
mov rax, 0xEEEEEEEEEEEEEEEE
mov [rdx + 8 * 6], rax
mov rax, 0xFFFFFFFFFFFFFFFF
mov [rdx + 8 * 7], rax
; Test truncation
vmovapd ymm2, [rdx + 8 * 4]
vmovapd xmm2, [rdx + 8 * 4]
; Test memory overwrite
vmovapd ymm0, [rdx]
vmovapd [rdx + 8 * 4], ymm0
vmovapd ymm1, ymm0
mov rax, [rdx + 8 * 6]
mov rbx, [rdx + 8 * 7]
hlt
+26
View File
@@ -0,0 +1,26 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM2": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC"]
}
}
%endif
lea rdx, [rel .data]
; Load inputs
vmovaps ymm1, [rdx]
vmovaps xmm2, [rdx]
vmovaps ymm3, [rdx + 32]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
+49
View File
@@ -0,0 +1,49 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"RAX": "0x6162636465666768",
"RBX": "0x7172737475767778",
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM1": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
mov rdx, 0xe0000000
mov rax, 0x4142434445464748
mov [rdx + 8 * 0], rax
mov rax, 0x5152535455565758
mov [rdx + 8 * 1], rax
mov rax, 0x6162636465666768
mov [rdx + 8 * 2], rax
mov rax, 0x7172737475767778
mov [rdx + 8 * 3], rax
mov rax, 0xCCCCCCCCCCCCCCCC
mov [rdx + 8 * 4], rax
mov rax, 0xDDDDDDDDDDDDDDDD
mov [rdx + 8 * 5], rax
mov rax, 0xEEEEEEEEEEEEEEEE
mov [rdx + 8 * 6], rax
mov rax, 0xFFFFFFFFFFFFFFFF
mov [rdx + 8 * 7], rax
; Test truncation
vmovaps ymm2, [rdx + 8 * 4]
vmovaps xmm2, [rdx + 8 * 4]
; Test memory overwrite
vmovaps ymm0, [rdx]
vmovaps [rdx + 8 * 4], ymm0
vmovaps ymm1, ymm0
mov rax, [rdx + 8 * 6]
mov rbx, [rdx + 8 * 7]
hlt
+41
View File
@@ -0,0 +1,41 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0xEEEEEEEEFFFFFFFF", "0xCCCCCCCCDDDDDDDD", "0xAAAAAAAABBBBBBBB", "0x0808080809090909"],
"XMM1": ["0xEEEEEEEEFFFFFFFF", "0xEEEEEEEEFFFFFFFF", "0xAAAAAAAABBBBBBBB", "0xAAAAAAAABBBBBBBB"],
"XMM2": ["0xEEEEEEEEFFFFFFFF", "0xEEEEEEEEFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0xEEEEEEEEFFFFFFFF", "0xEEEEEEEEFFFFFFFF", "0xAAAAAAAABBBBBBBB", "0xAAAAAAAABBBBBBBB"],
"XMM4": ["0xEEEEEEEEFFFFFFFF", "0xEEEEEEEEFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM5": ["0xEEEEEEEEFFFFFFFF", "0xEEEEEEEEFFFFFFFF", "0xAAAAAAAABBBBBBBB", "0xAAAAAAAABBBBBBBB"],
"XMM6": ["0xCCCCCCCCDDDDDDDD", "0xCCCCCCCCDDDDDDDD", "0x0000000000000000", "0x0000000000000000"]
}
}
%endif
lea rdx, [rel .data]
;; Register duplication
vmovapd ymm0, [rdx]
vmovddup ymm1, ymm0
; 128-bit
vmovddup xmm2, xmm0
;; Same register
vmovapd ymm3, ymm0
vmovddup ymm3, ymm3
; 128-bit
vmovapd ymm4, ymm0
vmovddup xmm4, xmm4
;; From memory
vmovddup ymm5, [rdx]
; 128-bit
vmovddup xmm6, [rdx + 8]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xEE, 0xEE, 0xEE, 0xEE, 0xDD, 0xDD, 0xDD, 0xDD, 0xCC, 0xCC, 0xCC, 0xCC
db 0xBB, 0xBB, 0xBB, 0xBB, 0xAA, 0xAA, 0xAA, 0xAA, 0x09, 0x09, 0x09, 0x09, 0x08, 0x08, 0x08, 0x08
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM2": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD"],
"XMM4": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM5": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM6": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
; Load inputs
vmovdqa ymm1, [rdx]
vmovdqa xmm2, [rdx]
vmovdqa ymm3, [rdx + 32]
; Test memory overwrite
mov rax, 0xCCCCCCCCCCCCCCCC
mov [rdx + 32], rax
mov rax, 0xDDDDDDDDDDDDDDDD
mov [rdx + 40], rax
mov rax, 0xEEEEEEEEEEEEEEEE
mov [rdx + 48], rax
mov rax, 0xFFFFFFFFFFFFFFFF
mov [rdx + 56], rax
vmovdqa ymm4, [rdx + 32]
vmovdqa [rdx], xmm4
vmovapd ymm5, [rdx]
vmovdqa [rdx], ymm4
vmovapd ymm6, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
+45
View File
@@ -0,0 +1,45 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM2": ["0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD"],
"XMM4": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM5": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"],
"XMM6": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
; Load inputs
vmovdqu ymm1, [rdx]
vmovdqu xmm2, [rdx]
vmovdqu ymm3, [rdx + 32]
; Test memory overwrite
mov rax, 0xCCCCCCCCCCCCCCCC
mov [rdx + 32], rax
mov rax, 0xDDDDDDDDDDDDDDDD
mov [rdx + 40], rax
mov rax, 0xEEEEEEEEEEEEEEEE
mov [rdx + 48], rax
mov rax, 0xFFFFFFFFFFFFFFFF
mov [rdx + 56], rax
vmovdqu ymm4, [rdx + 32]
vmovdqu [rdx], xmm4
vmovapd ymm5, [rdx]
vmovdqu [rdx], ymm4
vmovapd ymm6, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
+36
View File
@@ -0,0 +1,36 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xCCCCCCCCCCCCCCCC", "0xEEEEEEEEEEEEEEEE", "0x0000000000000000", "0x0000000000000000"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0xDDDDDDDDDDDDDDDD", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
;; Register as DST tests
; Load inputs
vmovapd ymm1, [rdx]
vmovapd ymm2, [rdx + 32]
vmovhpd xmm1, xmm2, [rdx + 48]
vmovhpd xmm3, xmm1, [rdx + 56]
;; Store to memory test
; Overwrite beginning of data, then yank it back into a vector
; Nothing in memory should be modified except the first 64 bits.
vmovhpd [rdx], xmm2
vmovapd ymm4, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
db 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
+36
View File
@@ -0,0 +1,36 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xCCCCCCCCCCCCCCCC", "0xEEEEEEEEEEEEEEEE", "0x0000000000000000", "0x0000000000000000"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM3": ["0xCCCCCCCCCCCCCCCC", "0xFFFFFFFFFFFFFFFF", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0xDDDDDDDDDDDDDDDD", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
;; Register as DST tests
; Load inputs
vmovapd ymm1, [rdx]
vmovapd ymm2, [rdx + 32]
vmovhps xmm1, xmm2, [rdx + 48]
vmovhps xmm3, xmm1, [rdx + 56]
;; Store to memory test
; Overwrite beginning of data, then yank it back into a vector
; Nothing in memory should be modified except the first 64 bits.
vmovhps [rdx], xmm2
vmovapd ymm4, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
db 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
+36
View File
@@ -0,0 +1,36 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xEEEEEEEEEEEEEEEE", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM3": ["0xFFFFFFFFFFFFFFFF", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0xCCCCCCCCCCCCCCCC", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
;; Register as DST tests
; Load inputs
vmovapd ymm1, [rdx]
vmovapd ymm2, [rdx + 32]
vmovlpd xmm1, xmm2, [rdx + 48]
vmovlpd xmm3, xmm1, [rdx + 56]
;; Store to memory test
; Overwrite beginning of data, then yank it back into a vector
; Nothing in memory should be modified except the first 64 bits.
vmovlpd [rdx], xmm2
vmovapd ymm4, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
db 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
+36
View File
@@ -0,0 +1,36 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM1": ["0xEEEEEEEEEEEEEEEE", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"],
"XMM2": ["0xCCCCCCCCCCCCCCCC", "0xDDDDDDDDDDDDDDDD", "0xEEEEEEEEEEEEEEEE", "0xFFFFFFFFFFFFFFFF"],
"XMM3": ["0xFFFFFFFFFFFFFFFF", "0xDDDDDDDDDDDDDDDD", "0x0000000000000000", "0x0000000000000000"],
"XMM4": ["0xCCCCCCCCCCCCCCCC", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF", "0xFFFFFFFFFFFFFFFF"]
}
}
%endif
lea rdx, [rel .data]
;; Register as DST tests
; Load inputs
vmovapd ymm1, [rdx]
vmovapd ymm2, [rdx + 32]
vmovlps xmm1, xmm2, [rdx + 48]
vmovlps xmm3, xmm1, [rdx + 56]
;; Store to memory test
; Overwrite beginning of data, then yank it back into a vector
; Nothing in memory should be modified except the first 64 bits.
vmovlps [rdx], xmm2
vmovapd ymm4, [rdx]
hlt
align 32
.data:
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
db 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xCC, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD, 0xDD
db 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xEE, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF
+42
View File
@@ -0,0 +1,42 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x6162636465666768", "0x7172737475767778", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
mov rdx, 0xe0000000
mov rax, 0x4142434445464748
mov [rdx + 8 * 0], rax
mov rax, 0x5152535455565758
mov [rdx + 8 * 1], rax
mov rax, 0x6162636465666768
mov [rdx + 8 * 2], rax
mov rax, 0x7172737475767778
mov [rdx + 8 * 3], rax
mov rax, 0x0
mov [rdx + 8 * 4], rax
mov [rdx + 8 * 5], rax
mov [rdx + 8 * 6], rax
mov [rdx + 8 * 7], rax
vmovaps xmm0, [rdx + 8 * 0]
vmovaps xmm1, [rdx + 8 * 2]
vmovaps ymm2, [rdx + 8 * 0]
vmovntdq [rdx + 8 * 4], xmm1
vmovaps xmm0, [rdx + 8 * 4]
vmovntdq [rdx + 8 * 4], ymm2
vmovaps ymm3, [rdx + 8 * 4]
hlt
+29
View File
@@ -0,0 +1,29 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x4142434445464748", "0x5152535455565758", "0x0000000000000000", "0x0000000000000000"],
"XMM1": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
mov rdx, 0xe0000000
mov rax, 0x4142434445464748
mov [rdx + 8 * 0], rax
mov rax, 0x5152535455565758
mov [rdx + 8 * 1], rax
mov rax, 0x6162636465666768
mov [rdx + 8 * 2], rax
mov rax, 0x7172737475767778
mov [rdx + 8 * 3], rax
vmovntdqa xmm0, [rdx]
vmovntdqa ymm1, [rdx]
hlt
+42
View File
@@ -0,0 +1,42 @@
%ifdef CONFIG
{
"HostFeatures": ["AVX"],
"RegData": {
"XMM0": ["0x6162636465666768", "0x7172737475767778", "0x0000000000000000", "0x0000000000000000"],
"XMM3": ["0x4142434445464748", "0x5152535455565758", "0x6162636465666768", "0x7172737475767778"]
},
"MemoryRegions": {
"0x100000000": "4096"
}
}
%endif
mov rdx, 0xe0000000
mov rax, 0x4142434445464748
mov [rdx + 8 * 0], rax
mov rax, 0x5152535455565758
mov [rdx + 8 * 1], rax
mov rax, 0x6162636465666768
mov [rdx + 8 * 2], rax
mov rax, 0x7172737475767778
mov [rdx + 8 * 3], rax
mov rax, 0x0
mov [rdx + 8 * 4], rax
mov [rdx + 8 * 5], rax
mov [rdx + 8 * 6], rax
mov [rdx + 8 * 7], rax
vmovaps xmm0, [rdx + 8 * 0]
vmovaps xmm1, [rdx + 8 * 2]
vmovaps ymm2, [rdx + 8 * 0]
vmovntpd [rdx + 8 * 4], xmm1
vmovaps xmm0, [rdx + 8 * 4]
vmovntpd [rdx + 8 * 4], ymm2
vmovaps ymm3, [rdx + 8 * 4]
hlt
Loaded 100 of 131 files, more files were not shown because too many files have changed in this diff. Show more