Compare commits

...
509 Commits
Author SHA1 Message Date
Ryan Houdek ee56a2cbde Docs: Update for release FEX-2403 2024-03-04 17:55:24 -08:00
Mai d66a83a98f Merge pull request #3482 from Sonicadvance1/fix_3479
Removes steamwebhelper config
2024-03-02 22:27:28 -05:00
Ryan Houdek 067b346e1b Removes steamwebhelper config
Fixes #3479

As of the Steam update on Feb 29 2024, this is no longer necessary and
actually causes steamwebhelper to abort.

Steam has started using SLR to launch steamwebhelper, which already
passes in --no-sandbox. Adding this argument twice seemingly breaks the
application since it starts complaining that it doesn't understand the
argument and closes with a SIGABRT.

With this removed Steam starts working again.
2024-03-02 13:32:08 -08:00
Ryan Houdek ea7d1697b1 Merge pull request #3480 from alyssarosenzweig/setf
Use SETF8/16 for 8/16-bit INC/DEC
2024-03-02 13:31:44 -08:00
Alyssa Rosenzweig 5fd91d34f1 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig 11880459a5 OpcodeDispatcher: use SETF for DEC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig 0ef0bb2c97 OpcodeDispatcher: use SETF for INC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig 72edee7c6f IR: add SETF8/SETF16 ir ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-01 19:40:53 -04:00
Alyssa Rosenzweig cb5644cf81 InstCountCI: add multi-inst 8-bit DEC cases
would've got a bug in an earlier version of this series

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-01 19:38:40 -04:00
Ryan Houdek 2dd922c3ce Merge pull request #3477 from neobrain/feature_update_vulkan
Library Forwarding: Update Vulkan definitions to v1.3.278
2024-03-01 10:38:14 -08:00
Alyssa Rosenzweig b7984e8651 Merge pull request #3471 from Sonicadvance1/another_3421_bug
ASM: Another sign extend bug in #3421
2024-03-01 10:57:32 -04:00
Tony Wasserka 31e976a5bc Library Forwarding: Update Vulkan definitions to v1.3.278 2024-02-29 19:08:38 +01:00
Ryan Houdek 009ae55ff0 Merge pull request #3475 from alyssarosenzweig/opt/lock-dec
Optimize lock dec
2024-02-29 08:44:24 -08:00
Ryan Houdek 98572b9e23 Merge pull request #3473 from Sonicadvance1/remove_mov_swap
Arm64: Stop moving source in atomic swap
2024-02-29 08:44:16 -08:00
Alyssa Rosenzweig 1ab234f6dd InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-29 09:28:21 -04:00
Alyssa Rosenzweig fed5e6d546 OpcodeDispatcher: use fetchadd for atomic DEC
Avoids a NEG.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-29 09:28:21 -04:00
Alyssa Rosenzweig f27e2246e2 Merge pull request #3468 from alyssarosenzweig/opt/miscs
Misc little opts
2024-02-29 09:18:17 -04:00
Mai 4779fb74de Merge pull request #3474 from Sonicadvance1/fix_early_exit_region_reserving
Fix reserving range check
2024-02-29 07:22:07 -05:00
Ryan Houdek eaf83aa6b4 Fix reserving range check
Fixes an issue where TestHarnessRunner was managing to reserve the space
below stack again, resulting in stack growth breaking. Would typically
only show up when using the vixl simulator under gdb for some reason.

This is likely the last bandage on this code before it gets completely
rewritten to be more readable.
2024-02-29 04:02:05 -08:00
Ryan Houdek 4f8b28e83b InstcountCI: Update for swap improvement 2024-02-29 03:07:19 -08:00
Ryan Houdek c318947695 Arm64: Stop moving source in atomic swap
ldswpal doesn't overwrite the source register and only reads the bits
required for the sized operation.
Not sure exactly why we were doing a copy here.

Removing it means improving Skyrim's hottest code block, as seen in #3472
2024-02-29 03:07:05 -08:00
Ryan Houdek b3489d7262 ASM: Another sign extend bug in #3421
This time found in MGRR. It flips the problem space on its head.
2024-02-29 02:06:57 -08:00
Alyssa Rosenzweig 44d738fa93 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:35:34 -04:00
Alyssa Rosenzweig 811487ad98 OpcodeDispatcher: use real branch for INT
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig 4f4e38ace2 OpcodeDispatcher: use real branch for rep cmps
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig edd6becc56 OpcodeDispatcher: use real branch for rep scas
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:35:12 -04:00
Alyssa Rosenzweig e47a94cae7 OpcodeDispatcher: skip mask with shld
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:18:29 -04:00
Alyssa Rosenzweig cc82dba1ca OpcodeDispatcher: use mvn for AF with constants
This reduces pointless constant usage. For now, it's no net change to
instcountci, but it should make it easier to get wins later. Hopefully.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:18:28 -04:00
Alyssa Rosenzweig 8232669b22 OpcodeDispatcher: simplify
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 10:00:13 -04:00
Alyssa Rosenzweig 28936073c4 OpcodeDispatcher: allow upper garbage on NEG
like SUB.

due to RA silliness, this is a loss for inst count but a win for cycles.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 09:59:57 -04:00
Alyssa Rosenzweig ef2559d911 OpcodeDispatcher: allow garbage with SCAS
it's just feeding SUB flags which allow it

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-28 09:16:08 -04:00
Ryan Houdek d24446ed13 Merge pull request #3466 from Sonicadvance1/fixed_opt
OpcodeDispatcher: Don't use AddShift with no shift
2024-02-28 04:51:10 -08:00
Ryan Houdek 67f13ba927 Adds sign extending address bug that was detected when testing #3421
Doesn't quite match the libc code directly because it uses `[gs:eax]`
with both having the sign bit set and we can't deal with that with ASM
tests. So match the behaviour in a different way.
2024-02-27 19:57:22 -08:00
Ryan Houdek dc5239c003 InstCountCI: Update for previous fix 2024-02-27 19:57:06 -08:00
Ryan Houdek f346f89678 OpcodeDispatcher: Don't use AddShift with no shift
This accidentally removed optimizations elsewhere that was only checking
for Add.
2024-02-27 19:56:12 -08:00
Ryan Houdek 2f9449cb5a Merge pull request #3465 from alyssarosenzweig/icci/pa
InstCountCI: enable preserve_all
2024-02-27 16:39:46 -08:00
Ryan Houdek 139367d248 Merge pull request #3463 from Sonicadvance1/update_xxhash
Update xxhash to v0.8.2
2024-02-27 16:39:38 -08:00
Alyssa Rosenzweig fcebad51bd InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-27 12:04:29 -04:00
Alyssa Rosenzweig b50292493a InstCountCI: enable preserve_all ABI
This is what we'll actually ship (I hope), so that's the config we want to
track long-term. It's also a lot more managable resulting asm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-27 12:03:58 -04:00
Ryan Houdek b74de53056 Merge pull request #3449 from Sonicadvance1/syscalls_passthrough_not_json
Linux: Converts passthrough syscalls to direct passthrough handlers
2024-02-27 07:50:04 -08:00
Alyssa Rosenzweig 49e798ab2b Merge pull request #3461 from alyssarosenzweig/opt/sbc
Optimize SBC
2024-02-27 11:29:45 -04:00
Ryan Houdek 151e2279af Linux: Converts passthrough syscalls to direct passthrough handlers
Reimagining of #3355 without any json generators or new concepts.

Fixes some mislabeling of system calls. Some getting inlined when they
shouldn't be, a lot not getting inlined when they can be.

This really cleans up the syscall implementation, all syscalls that can
be passthrough implementations require a very small two line
declaration.
Additionally cleans up a bit of implementation cruft where some
passthrough syscalls were using the glibc syscall handler, and some were
using the glibc implementation. We have had multiple issues in the past
where the glibc implementation does something subtly different than the
raw syscall and breaks things. Now all passthrough handlers do a system
call directly, removing at least one indirection and some ambiguity.

This makes it significantly easier to add new passthrough syscalls as
well. Only need to do a version check and add the three lines per
syscall. Which there are new syscalls incoming that we will want to add.

Tangible improvements:
- Syscalls are lower overhead than ever.
- When I'm adding more syscalls I have less chance of mucking it up.
2024-02-27 02:40:53 -08:00
Ryan Houdek 93ada89708 Linux: Move unimplement ustat and sysfs
AArch64 doesn't implement these and will return ENOSYS.
Moving them to NotImplemented so we can get a log if an application
tries to use these.
2024-02-27 02:39:36 -08:00
Mai f41674bb7d Merge pull request #3464 from Sonicadvance1/psychonauts_block
InstcountCI: Add a monster of a game block
2024-02-27 04:59:20 -05:00
Ryan Houdek 854fd70735 InstcountCI: Add a monster of a game block
Doing very little work with a bunch of instructions.
Hottest block in the Windows version of Psychonauts, it's just doing a
matrix swizzle but in the worst possible way.
2024-02-27 01:51:20 -08:00
Ryan Houdek 78a362581d Update xxhash to v0.8.2
Switches to using upstream cmake files.
2024-02-26 23:57:25 -08:00
Mai 8c0d5c6583 Merge pull request #3462 from Sonicadvance1/update_vixl4
Update vixl
2024-02-27 02:28:39 -05:00
Ryan Houdek 1c184997e7 InstcountCI: Update for vixl update 2024-02-26 23:17:52 -08:00
Ryan Houdek ccc699444d Update vixl
Removes a commit from our fork.
2024-02-26 23:16:31 -08:00
Ryan Houdek 946c805d84 Merge pull request #3459 from Sonicadvance1/fix_591
Capture a 64-bit process trying to jump to 32-bit syscall handler
2024-02-26 21:57:03 -08:00
Ryan Houdek 118b8b200e Merge pull request #3458 from Sonicadvance1/fix_635
Track unittest dependencies through to the custom target
2024-02-26 21:56:45 -08:00
Ryan Houdek aa9d7c5629 Merge pull request #3460 from Sonicadvance1/add_unittest_for_3421_bug
Adds a unittest for a bug from #3421
2024-02-26 21:56:17 -08:00
Ryan Houdek 0b34035085 Merge pull request #3439 from Sonicadvance1/allocate_first_4gb_of_64bit
FEXLoader: Allocate the second 4GB of virtual memory when executing 32-bit
2024-02-26 18:46:22 -08:00
Alyssa Rosenzweig b6bd826014 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:15 -04:00
Alyssa Rosenzweig 12cc980603 OpcodeDispatcher: shuffle adc flag order
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig f3d55dd721 OpcodeDispatcher: shuffle SBC flag order
avoids clobbering nzcv

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig 91cef6b76f OpcodeDispatcher: use native ADC even for 8/16-bit
we mask off the upper bits, and they agree in the lower bits.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:48:07 -04:00
Alyssa Rosenzweig 270cbf39b5 OpcodeDispatcher: specialize SALC
this gets rid of the awkward non-flag SBB case, which streamlines SBB. while
getting better codegen for the demon opcode (-:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 16:33:41 -04:00
Alyssa Rosenzweig 2e0be0a5e7 OpcodeDispatcher: allow more upper garbage with adc
missed this last series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:42:13 -04:00
Alyssa Rosenzweig d60c089697 OpcodeDispatcher: allow upper garbage with sbb
for the usual reasons

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig e76ebeab58 OpcodeDispatcher: use 1-op "src + CF"
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig 333271d490 OpcodeDispatcher: fuse sbb when flags calculated
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:35:14 -04:00
Alyssa Rosenzweig a750870abf OpcodeDispatcher: use fused sbcs calculations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:25:03 -04:00
Alyssa Rosenzweig 15db72ef60 IR: add Sbb, SbbWithFlags ops
For fusing sbc+sbcs

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-26 15:25:03 -04:00
Ryan Houdek 9687ac51f0 Merge pull request #3424 from Sonicadvance1/safer_clone_stack_handling
Linux: More safe stack cleanup for clone
2024-02-26 06:59:27 -08:00
Alyssa Rosenzweig 5f16f357af Merge pull request #3456 from alyssarosenzweig/opt/adc
Optimize ADC
2024-02-26 09:39:37 -04:00
Ryan Houdek 4f028b8614 Capture a 64-bit process trying to jump to 32-bit syscall handler
Fixes #591

Adds a simple unittest
2024-02-26 05:37:29 -08:00
Alyssa Rosenzweig 32a4abbea7 Merge pull request #3457 from alyssarosenzweig/bug/nzcv
RedundantFlagCalculationElimination: fix missing NEG case
2024-02-26 09:31:06 -04:00
Ryan Houdek c00c9b397e Adds a unittest for a bug from #3421
When the source arguments for LoadMem/StoreMem have bit 31 set then they
are incorrectly sign extending in some instances.

Detected this when testing #3421 but I don't have a proper fix for it.
2024-02-26 00:07:19 -08:00
Ryan Houdek 6b5d8bd8c0 Track unittest dependencies through to the custom target
Fixes #635
2024-02-25 19:27:52 -08:00
Alyssa Rosenzweig 7deb4976a3 RedundantFlagCalculationElimination: fix missing NEG case
can be predicated.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 17:37:16 -04:00
Alyssa Rosenzweig 2cfd71c159 InstCountCI: add dead ADC test
nothing else covers this case

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 10:50:25 -04:00
Alyssa Rosenzweig 0f26780de0 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 10:50:25 -04:00
Alyssa Rosenzweig 1e153e0c81 OpcodeDispatcher: allow garbage with adcs
for the usual reasons

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 10:50:02 -04:00
Alyssa Rosenzweig 6994fc3a01 IR,OpcodeDispatcher,JIT: fuse adcs flags
The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!

Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-25 10:49:32 -04:00
Alyssa Rosenzweig 0ef72bf118 Merge pull request #3451 from Sonicadvance1/fix_zero_reg_regression
Fixes zero register flag generation
2024-02-24 21:08:49 -04:00
Alyssa Rosenzweig 49ca0e2181 Merge pull request #3452 from Sonicadvance1/dxvk_mgrr_hotblock
Adds MGRR hottest block on render thread
2024-02-24 21:07:25 -04:00
Ryan Houdek 947ae1c243 Adds MGRR hottest block on render thread
Was about 7% CPU time in this looping block. Has some fairly obvious
performance improvements that can be done.
2024-02-24 16:49:46 -08:00
Ryan Houdek d703f3ccee Fixes zero register flag generation
Fixes 140976d322

Adds a unit test to ensure it keeps working.
2024-02-24 16:32:25 -08:00
Ryan Houdek d8a18687e8 Merge pull request #3443 from alyssarosenzweig/opt/add-too
Fuse add + cmn -> adds
2024-02-24 15:12:16 -08:00
Alyssa Rosenzweig 045549f166 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:55:07 -04:00
Alyssa Rosenzweig 80e632db8a OpcodeDispatcher: garbage collect
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 852e3c4e93 OpcodeDispatcher: fuse XADD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig e86547bbcb OpcodeDispatcher: fuse INC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 883cca2e8f OpcodeDispatcher: use AddWithFlags
give it the same treatment we just gave sub.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 8540332520 IR: add AddWithFlags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig df5bdefb8a OpcodeDispatcher: merge secondary ALU with primary ALU
It's the same, stop copypasting. This gets our flag and arithmetic opts (current
and future) applied to secondary ALU too.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 3d1fb7701c OpcodeDispatcher: optimize sub
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 140976d322 OpcodeDispatcher: prep primary ALU for better flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 9a11d3b1a2 OpcodeDispatcher: fuse NEG
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 96e652879f OpcodeDispatcher: fuse DEC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig cc1c1dd047 OpcodeDispatcher: return result from SUB flag calculate
for fusion

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig c1d572951f OpcodeDispatcher: drop unused GenerateFlags_SUB arg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig dd9d3264dd OpcodeDispatcher: smarten SUB flag generation
we don't need the result, we can use subs and come out ahead in practice. also a
step towards better fusion

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 2aaf957ad8 RedundantFlagCalculationElimination: DCE as we go
This is required to ensure single-iteration convergence with a sequence like:

  write C
  whatever = load C
  rmif C, whatever
  invalidate C

avoids regressing the "DEC dead" case with future work in the series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig d459b2f9b5 IR: propagate 0 into sub
now that we have to handle it, we may as well take advantage of it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 22ab7f2b3e IR: add SubWithFlags op (arm64 subs)
with 8/16-bit handling to keep everything uniform.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig f7e32373ce JIT: allow #0 in sub
turns into neg, this will be generated via SubWithFlags -> Sub opts.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 25d422e92b JIT: use GetZeroableReg for NZCVSelect
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig cfbeece09f JIT: use GetZeroableReg for CondAddNZCV
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig a597a09825 JIT: use GetZeroableReg for SubNZCV
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 99854ff310 JIT: add GetZeroableReg helper
for inlining constant zeroes in applicable sources

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-24 15:54:49 -04:00
Alyssa Rosenzweig 44d1502b5a Merge pull request #3450 from Sonicadvance1/vector_add_loop
InstcountCI: Adds a vector addition loop from bytemark
2024-02-24 15:54:44 -04:00
Ryan Houdek c3c635e36e InstcountCI: Adds a vector addition loop from bytemark
This was interesting because it caught that we are failing loadstores
with reg+reg address generation on vectors.
2024-02-24 03:44:02 -08:00
Ryan Houdek 7e2f20cabb Merge pull request #3446 from neobrain/fix_gdb_source_map
Syscalls: Fix SourcecodeMap generation for GDB JIT integration
2024-02-24 02:46:45 -08:00
Tony Wasserka 3425b07711 Syscalls: Fix SourcecodeMap generation for GDB JIT integration
This fixes a regression from 9dd715573, which accidentally changed the
filename and set up incorrect file opening flags.
2024-02-24 11:00:23 +01:00
Ryan Houdek 9b93495d45 Merge pull request #3441 from Sonicadvance1/remove_signaldelegation_tls
Moves SignalDelegator TLS tracking to the frontend
2024-02-24 01:23:43 -08:00
Ryan Houdek cf067994f3 Merge pull request #3440 from Sonicadvance1/addressing_limits
InstcountCI: Adds addressing limitations to instcountci
2024-02-24 01:09:22 -08:00
Ryan Houdek 0a64f8a9c5 Moves SignalDelegator TLS tracking to the frontend
FEXCore doesn't need track the TLS state of the SignalDelegator, this is
a frontend concept.

Removes the tracking from the backend and keeps it in the frontend.
2024-02-24 01:07:29 -08:00
Ryan Houdek 3ac7fe3f05 Linux: More safe stack cleanup for clone
Previously: Would keep one clone thread's stack active for teardown
delaying.

With aggressive cloning and teardown, this was unsafe.
Only reap the stack when told it is safe to do so.
2024-02-24 01:05:20 -08:00
Ryan Houdek be96cb7bd0 FEXLoader: Allocate the second 4GB of virtual memory when executing 32-bit
Spurred on by #3421. To ensure that applications don't take advantage of
small address wrap around, allocate the second 4GB of virtual memory.

Some context. Linux always reserves the first 16KB of virtual address
space (unless you tinker with some settings which nobody should do).

Example of 32-bit code:
lea eax, [0xffff_0000]
mov ebx, [eax + 0x1_0000]

The address calculated by the mov will wrap around to 0x0 which will
result in SIGSEGV. If FEX messes up zero extensions then it would try to
access 0x1_0000_0000 instead.

This could result in a 32-bit application potentially accessing some FEX
memory instead of crashing.
Add this safety net which will still SIGSEGV and we will be able to see
the crash.
2024-02-24 00:54:07 -08:00
Ryan Houdek 59ec88f48d Merge pull request #3438 from Sonicadvance1/move_tls_allocation
Moves JITSymbol allocation
2024-02-23 14:49:04 -08:00
Ryan Houdek 6ec628fa31 Merge pull request #3433 from bylaws/arm64ec-pt1
Arm64Emitter: Introduce ARM64EC SRA mappings
2024-02-23 14:48:43 -08:00
Alyssa Rosenzweig 90256730c3 Merge pull request #3447 from Sonicadvance1/fix_instcountci
Fix instcountci
2024-02-22 17:24:24 -04:00
Ryan Houdek 3a0f9db512 Fix instcountci 2024-02-22 12:44:10 -08:00
Alyssa Rosenzweig 5378ae2e76 Merge pull request #3436 from alyssarosenzweig/ir/af-simplify
Simplify CalculateAF
2024-02-22 08:17:07 -04:00
Ryan Houdek cc9c80d79f Merge pull request #3445 from alyssarosenzweig/instc/fmod
InstCountCI: add FMOD block
2024-02-21 19:27:07 -08:00
Ryan Houdek 60e8da05cd Merge pull request #3442 from alyssarosenzweig/instc/witcher3
InstCountCI: add The Witcher 3 block
2024-02-21 18:56:39 -08:00
Alyssa Rosenzweig aac7fa9b58 InstCountCI: add hot scalar FMOD block
both AFP and non-AFP versions, since this is greatly affected by AFP opts.

pulled from #2563

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-21 22:36:15 -04:00
Ryan Houdek bd4a81a2a1 Moves JITSymbol allocation
This isn't actually using TLS allocations. Instead it is an allocation
tied to the InternalThreadState object.
2024-02-21 17:57:34 -08:00
Alyssa Rosenzweig a6211f29e7 InstCountCI: add The Witcher 3 block
Closes: #2688

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-21 21:44:12 -04:00
Ryan Houdek da263834f8 InstcountCI: Adds addressing limitations to instcountci
Spurred on by #3421

This does a bunch of GPR and vector loads to showcase addressing
limitations between ARM and x86.

Tests:
- 8/16/32/64-bit GPR loads
   - Both 32-bit and 64-bit addressing modes
- 32/64/128-bit Vector loads
   - Both 32-bit and 64-bit addressing modes
- Duplicate the tests for 32-bit addressing mode with a 32-bit process
   - Since it should change behaviour.

Untested:
- 8/16-bit vector loadstores since those don't exist on x86
   - 16-bit x87 integer load exists but that doesn't go through a vector
     load in FEX.
2024-02-21 16:50:00 -08:00
Ryan Houdek 3d671cba10 Allocator: Early return if past the end of the allocation range
Fixes a bug where it would eventually hit the stack region and remap the
range as RW even for ranges that don't overlap the stack.
2024-02-21 15:37:46 -08:00
Ryan Houdek d4be2dc636 Merge pull request #3434 from bylaws/arm64ec-pt3
FEXCore: Expose AbsoluteLoopTopAddress to the frontend
2024-02-21 14:31:04 -08:00
Ryan Houdek 3e5694bd06 Merge pull request #3432 from bylaws/arm64ec-pt0
CMake: Define _M_ARM_64EC when building for ARM64EC
2024-02-21 14:30:09 -08:00
Ryan Houdek 66feea9e8e Merge pull request #3422 from neobrain/feature_libfwd_wayland32
Library Forwarding: Add support for 32-bit Wayland
2024-02-21 13:41:38 -08:00
Alyssa Rosenzweig 2bcd285851 Merge pull request #3430 from Sonicadvance1/tsc_scale
Implement small TSC scaling
2024-02-21 13:16:27 -04:00
Alyssa Rosenzweig 30ff225e80 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-21 12:49:14 -04:00
Alyssa Rosenzweig 8762bc1fa3 OpcodeDispatcher: simplify CalculateAF signature
- Res is unused
- SrcSize doesn't matter since we ignore the high bits, might as well always use
  32-bit, it doesn't matter

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-21 12:48:15 -04:00
Billy Laws 5b4162b712 FEXCore: Expose AbsoluteLoopTopAddress to the frontend
ARM64EC has a shared SRA mapping between ARM64 and X64 code, so there
needs to be a public way to enter the dispatcher without refilling SRA
from the in-memory context struct.
2024-02-21 11:46:24 +00:00
Billy Laws cb5c07f4b1 Arm64Emitter: Introduce ARM64EC SRA mappings
See https://learn.microsoft.com/en-us/cpp/build/arm64ec-windows-abi-conventions?view=msvc-170
note that since mm registers are volatile there is no need to match the
mapping for them when in JIT, so they can be used as scratch regs.
Disallowed regs are also wiped on context switches, so they cannot be
taken advantage of to e.g. avoid spilling.
2024-02-21 11:18:10 +00:00
Billy Laws 3364b48f3b CMake: Define _M_ARM_64EC when building for ARM64EC 2024-02-21 11:18:07 +00:00
Tony Wasserka 67530171e6 Library Forwarding/wayland: Clean up wl_interface exchange 2024-02-21 11:44:55 +01:00
Tony Wasserka 6f00611892 Library Forwarding/wayland: Add functions required by zink + Super Meat Boy on 32-bit 2024-02-21 11:44:55 +01:00
Tony Wasserka 67941f04eb Library Forwarding/wayland: Write arguments for callbacks to guest stack 2024-02-21 11:44:55 +01:00
Tony Wasserka 42531108b7 Library Forwarding/wayland: Enable 32-bit build 2024-02-21 11:44:55 +01:00
Tony Wasserka 7edab7ee3b Library Forwarding/wayland: For 32-bit guests, support repacking wl_interface/wl_message/wl_argument 2024-02-21 11:44:55 +01:00
Tony Wasserka b9f2389d74 Library Forwarding/wayland: Avoid using global state to substitute callback tables
The pointer tracked internally by Wayland can be queried via
wl_proxy_get_listener, so we don't need our own bookkeeping.

This also changes the callback table's element type to a fixed-size uint64_t,
which makes it work for 32-bit guests.
2024-02-21 11:44:55 +01:00
Tony Wasserka 3c0b041c44 Library Forwarding/wayland: Support reading message signatures on 32-bit guests 2024-02-21 11:44:54 +01:00
Tony Wasserka 80fcd640be unittests/Library Forwarding: Test interaction between ptr_passthrough and fixed-size integer mapping 2024-02-21 11:44:54 +01:00
Tony Wasserka 0908968e87 Library Forwarding/gen: Don't register types used exclusively with ptr_passthrough annotations
This allows forwarding APIs that sparsely use non-repackable types.
2024-02-21 11:44:54 +01:00
Ryan Houdek 8e0543d9af InstcountCI 2024-02-20 12:05:44 -08:00
Ryan Houdek b902b8edab Implement small TSC scaling
Games engines are expecting >1Ghz cycle counters. Scale them to work
around the issue.

Resolves the excessive busy waiting in Unreal Engine 5 games.
2024-02-20 12:05:44 -08:00
Alyssa Rosenzweig 9c38332e7e Merge pull request #3429 from alyssarosenzweig/cleanup/nzcv-helpers
Cleanup NZCV metadata
2024-02-20 10:18:15 -04:00
Alyssa Rosenzweig 0503c89ff6 OpcodeDispatcher: use NZCV update helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-19 14:12:54 -04:00
Alyssa Rosenzweig 6dd410698a OpcodeDispatcher: add helpers for updating NZCV metadata
to reduce error-prone copypaste

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-19 14:12:54 -04:00
Ryan Houdek bbac014d6d Merge pull request #3427 from Sonicadvance1/fix_vdso_crash
Fixes VDSO crash in 64-bit code
2024-02-18 18:47:41 -08:00
Ryan Houdek a723ff09c1 Fixes VDSO crash in 64-bit code
Ever since #3406 this has been crashing. Struct tail padding was saving
this before.
2024-02-17 16:43:09 -08:00
Mai 5769ffbba7 Merge pull request #3425 from Sonicadvance1/frontend_pointer
FEXCore: Add a frontend pointer to InternalThreadState
2024-02-16 19:41:41 -05:00
Mai a7c7fe4a35 Merge pull request #3426 from Sonicadvance1/thread_leak
Linux: Make sure to destroy thread object when thread shuts down
2024-02-16 19:41:23 -05:00
Ryan Houdek 80f20ad121 Linux: Make sure to destroy thread object when thread shuts down
This fixes a fairly large memor leak.
2024-02-16 15:08:29 -08:00
Ryan Houdek 808ced455d FEXCore: Add a frontend pointer to InternalThreadState
FEXCore is guaranteed to not touch this pointer and can be used by
frontends to store thread-specific data.
2024-02-15 02:06:16 -08:00
Mai 0505b30d34 Merge pull request #3423 from Sonicadvance1/fix_fd_leak
FileFormatCheck: Fixes FD leak
2024-02-14 16:01:29 -05:00
Ryan Houdek 16a5d1a6b1 FileFormatCheck: Fixes FD leak 2024-02-14 12:21:47 -08:00
Ryan Houdek 9cab746aa7 Merge pull request #3407 from neobrain/feature_libfwd_arguments_on_guest_stack
Library Forwarding: Allocate packed arguments on the guest stack if needed
2024-02-12 16:31:34 -08:00
Ryan Houdek b888bb5ce5 Merge pull request #3406 from neobrain/feature_libfwd_packed_arguments
Library Forwarding: Disable struct padding for packed arguments
2024-02-12 16:23:44 -08:00
Ryan Houdek 47218254a1 Merge pull request #3416 from alyssarosenzweig/opt/add-smol
Optimize 8/16-bit adds/subs
2024-02-12 15:52:05 -08:00
Ryan Houdek 7cb15582c2 Docs: Update for release FEX-2402 2024-02-12 10:37:55 -08:00
Ryan Houdek 930d2650f8 Merge pull request #3417 from Sonicadvance1/workaround_hang
ThreadManager: StealAndDropActiveLocks in the child forked process
2024-02-12 10:36:16 -08:00
Alyssa Rosenzweig fba5678476 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-12 12:36:48 -04:00
Alyssa Rosenzweig 68232366e4 OpcodeDispatcher: don't mask add/sub sources
not needed in the new approach

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-12 12:36:28 -04:00
Alyssa Rosenzweig d7ff1b78fb IR: handle 8/16-bit AddNZCV/SubNZCV
we can do it more effectively than the current s/w lowering.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-12 12:36:09 -04:00
Mai 780b48620b Merge pull request #3420 from Sonicadvance1/preserve_all_3419
Fix #3419
2024-02-10 23:24:38 -05:00
Ryan Houdek 4a0878fa92 Fix #3419 2024-02-10 19:55:51 -08:00
Ryan Houdek 67143ed1d1 ThreadManager: StealAndDropActiveLocks in the child forked process
The creation mutex could have been held if the parent thread was in the
middle of creating a thread when forking. This would result in a
deadlock once the fork child attempted to create another thread.

Forcefully dropping the lock in the fork child works around this
deadlock. This comes at the expense of potentially leaving resources
guarded by the thread creation mutex in an invalid state. Crashes caused
by this are easier to reason about than a delayed deadlock, though.
2024-02-10 19:51:18 -08:00
Ryan Houdek df3d6938ae Merge pull request #3410 from alyssarosenzweig/opt/nzcv-pass-2
Add NZCV+PF/AF optimization pass
2024-02-10 05:03:12 -08:00
Ryan Houdek 577372c203 Linux: Consolidate LockBeforeFork usage
Moves the CTX LockBeforeFork in to the Syscallhandler's LockBeforeFork.

This lets the syscall handler just call its own LockBeforeFork and
UnlockAfterFork functions rather than two on each call site.

Also moves the CTX->UnlockAfterFork in to the SyscallHandler's to be
consistent with the LockBeforeFork half.

No functional change.
2024-02-09 05:55:23 -08:00
Ryan Houdek d4aa64ebd1 Linux: Convert ThreadCreationMutex to forkable mutex
We will need to hold this mutex when forking.

No functional change.
2024-02-09 05:55:23 -08:00
Ryan Houdek ba41da7da0 Merge pull request #3414 from Sonicadvance1/fix_one_mutex_hang
Fixes one mutex hang
2024-02-09 05:54:40 -08:00
Alyssa Rosenzweig 806e5b8bcf Merge pull request #3415 from alyssarosenzweig/opt/testm1
Optimize test -1
2024-02-09 09:52:21 -04:00
Ryan Houdek 2480bab409 Fixes one mutex hang
When code invalidation is happening we currently have the issue that a
thread can acquire the code invalidation mutex in the middle of
invalidation. This is due to us acquiring and releasing the mutex
between each thread's code invalidation.

We need to hold the mutex for the entire duration for all thread's code
invalidation.
This fixes a rare hang on proton startup and resolves a consistent hang
on Proton application shutdown.

This now puts us on par with FEX-2312.1 with hanging.

This does not fix a relatively rare hang on fork (which also existed with FEX-2312.1).

This also does not fix the issue that the intersection of our mutexes
between frontend and backend are very convoluted. In part of the work
that is going to fix the rare fork mutex hang will change more of this.
2024-02-08 18:18:00 -08:00
Alyssa Rosenzweig de0b690672 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-08 14:11:56 -04:00
Alyssa Rosenzweig 300e2729a7 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-08 14:10:30 -04:00
Alyssa Rosenzweig ad7202e7d7 OpcodeDispatcher: optimize test -1
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-08 14:10:13 -04:00
Alyssa Rosenzweig 69466dce92 InstCountCI: add flag optimizations tests
Now that lots of instructions have optimal flag calculation in isolation, we
need to look at sequences of instructions together. These cases test various
interesting cases where concatenating the optimal instruction translations gives
something terrible for the whole block. These cases exercise the
new flag optimization pass.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 14:22:28 -04:00
Alyssa Rosenzweig 175a57dd27 OpcodeDispatcher: emit AndWithFlags directly for primary alu
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 14:22:28 -04:00
Alyssa Rosenzweig e2ce60148c OpcodeDispatcher: emit AndWithFlags directly for 2ndary alu
rely on opt pass to drop the flags.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 14:22:28 -04:00
Alyssa Rosenzweig 99660129f3 IR: implement AndWithFlags for 8/16-bit
easier to deal with in the JIT

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 14:22:28 -04:00
Alyssa Rosenzweig 308d9a751c RedundantFlagCalculationElimination: optimize rmif
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:26 -04:00
Alyssa Rosenzweig 4bd28c0ed8 RedundantFlagCalculationElimination: optimize condaddnzcv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:26 -04:00
Alyssa Rosenzweig 8397f3ac99 RedundantFlagCalculationElimination: refine AXFLAG
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:26 -04:00
Alyssa Rosenzweig 0452bc7212 RedundantFlagCalculationElimination: optimize condjump
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:26 -04:00
Alyssa Rosenzweig 3d7ed89ffb RedundantFlagCalculationElimination: optimize NZCVSelect
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:25 -04:00
Alyssa Rosenzweig 23ab0a978e RedundantFlagCalculationElimination: also handle InvalidateFlags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:25 -04:00
Alyssa Rosenzweig 7f47a9ef0e IR: add local dead flag elimination pass
RCLSE ignores NZCV and doesn't optimize stores which doesn't help us with PF/AF
either. So, we add a new pass for dead flag elimination (cannibalizing the old
and broken dead flag elimination pass). This is a simple local optimizer that
walks each block backwards, converging in linear time & constant space in a
single iteration.

Right now, it doesn't do a ton (other than a nice reduction in silliness in
the hot Sonic block), but it provides the framework to fuse comparisons.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-06 13:06:25 -04:00
Alyssa Rosenzweig 4331753ca0 Merge pull request #3408 from alyssarosenzweig/opt/tst
Optimize TST
2024-02-06 11:28:02 -04:00
Ryan Houdek cdcc432399 Merge pull request #3411 from pmatos/AccessNullptr
Clean up access to possible nullptr
2024-02-06 04:52:28 -08:00
Paulo Matos fa8bcfd67a Clean up access to possible nullptr
Patch suggested by @Sonicadvance1
2024-02-06 12:31:38 +00:00
Mai 5e5984a29b Merge pull request #3404 from Sonicadvance1/fix_early_thread_create_tracking
Linux: Decouple thread object creation and tracking
2024-02-05 17:43:48 -05:00
Tony Wasserka 435b67aa0c Merge pull request #3409 from FEX-Emu/revert-3394-PreserveAllOption
Revert "Add cmake option DISABLE_CLANG_PRESERVE_ALL"
2024-02-05 22:40:29 +01:00
Tony Wasserka a1343e9296 Revert "Add cmake option DISABLE_CLANG_PRESERVE_ALL" 2024-02-05 22:31:45 +01:00
Alyssa Rosenzweig 2dfb7727d5 InstCountCI: add test-with-self cases
these are useful x86 instructions for e.g. branch-if-not-zero, and have
opportunities for clever translations. Add the instcountci cases to make sure we
do something reasonable for them.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-05 15:41:29 -04:00
Alyssa Rosenzweig 235f32ce8c Merge pull request #3401 from Sonicadvance1/runtime_preserve_all
HostFeatures: Supports runtime disabling of preserve_all
2024-02-05 15:34:46 -04:00
Alyssa Rosenzweig 560565aa13 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-05 15:34:11 -04:00
Alyssa Rosenzweig 2e0cb2fbd4 OpcodeDispatcher: optimize TST
it's just an AndWithFlags setting the PF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-05 15:32:21 -04:00
Alyssa Rosenzweig 4790a7ba79 IR: add AndWithFlags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-05 15:32:21 -04:00
Mai a4a1d607d5 Merge pull request #3405 from Sonicadvance1/remove_unused_variable
SpinWaitLock: Removes unused variable in spin-loop fallback
2024-02-05 13:25:22 -05:00
Tony Wasserka df3e51fc8c Library Forwarding: Allocate packed arguments on the guest stack if needed
This is required for host-side calls to guest functions on 32-bit guests.
Since the host stack is allocated before FEX blocks memory inaccessible to
the guest, the guest would otherwise fail to read the packed argument data.
2024-02-05 18:10:34 +01:00
Tony Wasserka 06c29eab88 Library Forwarding: Disable struct padding for packed arguments
ARM64, x86 (64-bit), and x86 (32-bit) each have different alignment
requirements, so this change ensures that consistent data layout is
used for packing and unpacking.
2024-02-05 17:39:34 +01:00
Ryan Houdek 0139498072 SpinWaitLock: Removes unused variable in spin-loop fallback
Tmp was no longer being used, forgot to remove it.
2024-02-05 07:22:52 -08:00
Ryan Houdek 8cfbabde94 Linux: Decouple thread object creation and tracking
If the thread object is added to the tracking vector immediately then
there ends up being a race condition before the thread manages to fill
out the thread-specific data that only occurs at the start of the new
thread.

This manifests in a crash when a thread is allocating memory while
another thread is getting constructed. Easy fix is to defer the tracking
until the thread has setup its state.
2024-02-05 07:18:50 -08:00
Ryan Houdek 472a701e2b Merge pull request #3403 from Sonicadvance1/fix_spinlock_contended_lock
SpinLockWait: Fixes unexpected lock success
2024-02-05 06:51:42 -08:00
Mai a41ebe2d85 Merge pull request #3402 from Sonicadvance1/revert_the_revert_revert(q)revert_
Revert "Revert "FEXLoader: Moves thread management to the frontend""
2024-02-03 14:21:48 -05:00
Ryan Houdek cce6011205 SpinLockWait: Fixes unexpected lock success
With a contended unique lock, we forgot to reset the `Expected` value to
zero. This was causing a contended mutex to incorrectly succeed.

Noticed this when converting some pthread mutexes over to spinloops to
remove strace noise.

The reference wfe_mutex library I wrote didn't have this problem since
the implementation is slightly different.
2024-02-03 01:10:57 -08:00
Ryan Houdek c437129ed8 Revert "Revert "FEXLoader: Moves thread management to the frontend""
This reverts commit 5358af7794.
2024-02-03 00:57:36 -08:00
Ryan Houdek 557cb593e1 Merge pull request #3396 from alyssarosenzweig/opt/sha-stuff
Eliminate spilling in sha256rnds2 and sha1rnds4
2024-02-02 09:59:46 -08:00
Alyssa Rosenzweig afae24d870 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:55 -04:00
Alyssa Rosenzweig 8d3f0b6f02 OpcodeDispatcher: reassociate and sink W in sha1
We only need each part of W extracted in the corresponding round, so sink the
extract into the round to reduce pressure.

Further, W and E are added and then never used again. So, by reassociating we
can do the add upfront, killing W and E at the start and further reducing
pressure.

Eliminates spilling in sha1rnds4.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 60f7b9bcc4 OpcodeDispatcher: optimze sha1's 2/3 expr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig a487557173 OpcodeDispatcher: extract BitwiseAtLeastTwo
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 394b4888bb OpcodeDispatcher: reassociate and remat C0, G0
costs 2 moves and eliminates the rest of our spilling

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 142cbdd852 OpcodeDispatcher: expand, reassociate, and interleave sha256 calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 2f9102f78d OpcodeDispatcher: expand & interleave sha256 calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig c9824d04cb OpcodeDispatcher: sink sha256 extracts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 9c2a569539 OpcodeDispatcher: reexpress Major in sha256
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 515aa4ce3e OpcodeDispatcher: fuse eor+ror in sha256
This reduces instructions a ton.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig f616beb992 OpcodeDispatcher: CSE sha
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 0dcf1e12b8 OpcodeDispatcher: copyprop sha logic
prepare for clever

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Alyssa Rosenzweig 2cbf544ef5 OpcodeDispatcher: expand sha logic
no functional change, just preparing for cleverness.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-02-02 13:03:07 -04:00
Ryan Houdek 0eed73beeb HostFeatures: Supports runtime disabling of preserve_all
This is used for instcountci to ensure instruction counts don't change
when a compiler supports this feature or not. Always runtime disable
when running in instcountci.

CMake option from #3394 can still be useful so leaving that in place.
2024-02-02 08:59:04 -08:00
Mai 6993f4fd8d Merge pull request #3400 from Sonicadvance1/revert_runtime_longmode_switch
Revert #3303
2024-02-02 11:53:44 -05:00
Mai 920a8db369 Merge pull request #3397 from pmatos/XCHGOp
Improve XCHG operations
2024-02-02 11:53:22 -05:00
Ryan Houdek 9c37c0f1c3 Merge pull request #3392 from neobrain/feature_libfwd_fixed_size_ints
Library Forwarding: Handle cross-architecture differences of integer types
2024-02-02 08:42:41 -08:00
Paulo Matos 4623544f69 Improve XCHG operations
Marking loads as allowing upper garbage simplifies some operations.
Update InstCountCI as well.
2024-02-02 08:16:13 +00:00
Ryan Houdek 45587278c9 InstCountCI: Update for segment changes 2024-02-01 18:17:28 -08:00
Ryan Houdek ccf1402fe6 Revert "FEXCore: Accurately store segment descriptors"
This reverts commit 8648fb1485.
2024-02-01 18:14:30 -08:00
Ryan Houdek da0e1b515a Revert "OpcodeDispatcher: Initial support for runtime long-mode switch"
This reverts commit 9e5d7aa5fe.
2024-02-01 18:14:24 -08:00
Ryan Houdek c0ec4da849 Revert "Mingw: Update for GDT"
This reverts commit ae2f98e017.
2024-02-01 18:14:19 -08:00
Mai bd1e029f35 Merge pull request #3399 from Sonicadvance1/fix_json_argumenthandler
Config: Fixes JSON parsing of "ArgumentHandler" types
2024-02-01 19:59:45 -05:00
Ryan Houdek 690cb6fa48 Config: Fixes JSON parsing of "ArgumentHandler" types
When 4d109c9ce0 fixed parsing strenum
types in the json, it also added `ArgumentHandler` types to the json
parsing. This was incorrect as those types are already stored in the
json in their decoded numerical format.

Without this change, all config options with `ArgumentHandler` will
decode as "0" which is incorrect. The main killer here is that SMCChecks
gets disabled (visible in both FEXConfig and when applications are
running) which was causing spurious failures.
2024-02-01 16:20:57 -08:00
Ryan Houdek d34302a968 Merge pull request #3398 from neobrain/fix_flt_smc_warnings
FEXLinuxTests: Fix build warnings
2024-02-01 12:52:24 -08:00
Tony Wasserka ac77376441 FEXLinuxTests: Suppress "attribute ignored" warnings for nocf_check 2024-02-01 14:43:50 +01:00
Tony Wasserka f246e9864c FEXLinuxTests: Simplify a few extern "C" declarations 2024-02-01 14:43:50 +01:00
Tony Wasserka 78cb2cd9fc FEXLinuxTests: Fix warnings about unused results 2024-02-01 14:43:50 +01:00
Ryan Houdek cec1814a09 Merge pull request #3384 from pmatos/CDQOp-Opt
Optimize CDQOp
2024-01-31 17:51:23 -08:00
Mai 4d49ac7c3d Merge pull request #3387 from alyssarosenzweig/opt/rotates
Optimize rotates
2024-01-31 18:20:40 -05:00
Alyssa Rosenzweig 6d13d9fb56 Merge pull request #3395 from pmatos/StaticAnalysis
Code cleanup - mainly dead store removal; NFC
2024-01-31 17:24:48 -04:00
Mai ae7dc250db Merge pull request #3386 from alyssarosenzweig/opt/shift
Optimize shifts a bit
2024-01-31 14:11:58 -05:00
Mai f4086b25e6 Merge pull request #3385 from alyssarosenzweig/opt/bmi
Optimize bit manipulation instructions
2024-01-31 14:07:23 -05:00
Tony Wasserka 8e1aaa0559 Library Forwarding: Avoid de-sugaring pointee types
Doing so would accidentally resolve typedefs before. This code isn't needed
now that integers are mapped to fixed-size equivalents anyway.
2024-01-31 19:16:19 +01:00
Tony Wasserka a4c1aaa6bc Library Forwarding: Fix build for clang < 16 2024-01-31 19:14:57 +01:00
Ryan Houdek 37a611bcd6 Merge pull request #3394 from pmatos/PreserveAllOption
Add cmake option DISABLE_CLANG_PRESERVE_ALL
2024-01-31 00:58:30 -08:00
Paulo Matos e4560ed0c8 Code cleanup - mainly dead store removal; NFC
scan-build found a few dead stores that can be easily cleaned-up
2024-01-31 08:35:55 +00:00
Paulo Matos 6d58ea31b9 Add cmake option DISABLE_CLANG_PRESERVE_ALL
Forces disabling use of __attribute__((preserve_all)).
Until CI uses clang17, where this attribute was added, instcountci fails
when FEX is compiled with clang>=17.
2024-01-31 08:29:20 +00:00
Ryan Houdek 3db31a655c Merge pull request #3393 from Sonicadvance1/fix_relative_data
InstCountCI: Fixes test to not use relative data
2024-01-30 23:43:26 -08:00
Alyssa Rosenzweig e9a6bfc077 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 5a143c016c unittests: add ASM test for blsmsk flags
failing on main

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig fd1307eea3 unittests: add BLSI flag test
passing on main, but kinda by accident?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig d9f08d4bdf unittests: add BLSR flag unit test
failng on main

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig f3eee8f305 OpcodeDispatcher: optimize bextr's length sanitize
reordering the operations saves an immediate move.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig f66085f4a7 OpcodeDispatcher: optimize bextr's (1 << x) - 1
little algebraic trick I cribbed from llvm

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig c9461d9997 OpcodeDispatcher: optimize BEXTR flag setting
use native test.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig d5eb99fac8 OpcodeDispatcher: optimize popcount flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig f3175848b1 OpcodeDispatcher: use lzcnt flag gen for tzcnt
as far as flags go, they're identical: set ZF for zero output, set CF for output
= DestSize, undef the rest. merge the impls, so we get the optimized lzcnt impl
for tzcnt.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3dd597a591 OpcodeDispatcher: optimize lzcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig e8e05252f0 OpcodeDispatcher: optimize BLSI
and explain why the suss thing we did before was actually right all along.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 93cef53ec0 OpcodeDispatcher: optimize blsr flags
reorder to avoid nzcv clobber

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 3a19133267 OpcodeDispatcher: fix inverted BLSR carry
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 9b309b2102 OpcodeDispatcher: optimize blsmsk flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig fe88b904c9 OpcodeDispatcher: fix missing SF set with blsmsk 2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 2e63c6d547 OpcodeDispatcher: fix inverted CF with blsmsk
CF set if SRC = 0

per https://www.felixcloutier.com/x86/blsmsk

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:28:06 -04:00
Alyssa Rosenzweig 4cc59ba8f5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig 0bc9e1a409 OpcodeDispatcher: clobber OF with shift immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig 338f12845d OpcodeDispatcher: save a constant in shld
one weird trick

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig b3ae81f75f OpcodeDispatcher: allow garbage on shld shift
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig c1a1c37980 OpcodeDispatcher: mark ideas to improve SHLD
a bit tricky right now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:26:59 -04:00
Alyssa Rosenzweig 56dd9f23d1 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig fb6f850bb4 OpcodeDispatcher: remove rcl sub
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig b6d8749525 OpcodeDispatcher: remove select from rcl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig d3f1397325 OpcodeDispatcher: eliminate constants in RCR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 0a164428fa OpcodeDispatcher: eliminate select in RCR
the nzcv clobber I actually came ofr

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 7496175100 OpcodeDispatcher: optimize 32-bit rcl/rcr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 0616a9cef1 OpcodeDispatcher: eliminate move in rcr 1-bit
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 97f8775354 OpcodeDispatcher: optimize <32-bit rcr op1
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig c92099aa98 OpcodeDispatcher: fuse orlshl in rcr 1-bit
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 7c288b09f1 OpcodeDispatcher: rmif mask rcl smaller OF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 680af7b1b0 OpcodeDispatcher: rcr op 8x1 cleanup
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 349bc9efab OpcodeDispatcher: unify rcr op 1bit codepaths
get additional opt for <32-bit

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig ad5c3cb268 OpcodeDispatcher: rmif mask for OF in rcr smaller
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig be8d37ef3d OpcodeDispatcher: optimize 32-bit rol/ror imm
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 6ad2514bfe OpcodeDispatcher: rmif mask rcl smaller cf
better on flagm. extra moves on non-flagm but, meh.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 3fa6129a14 OpcodeDispatcher: rmif mask rcr smaller cf
and do some constant folding to do so more.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig a57cebaf58 OpcodeDispatcher: skip OF calc for constant rotate >= 2
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 34fdb14da1 OpcodeDispatcher: add and use AndConst
this skips the constant folding, which saves the branching in the rotate
immediate implementations.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 974baca09c OpcodeDispatcher: allow upper garbage with rcl/rcr smaller
we're masking immediately to something smaller

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig f22094a493 OpcodeDispatcher: use a branch for 8/16-bit rotate flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig d979b3a1da OpcodeDispatcher: note idea to further optimize rcl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Alyssa Rosenzweig 6d82c957fa OpcodeDispatcher: fuse orlshl in rcl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-30 22:22:57 -04:00
Ryan Houdek 28715a1df6 InstCountCI: Fixes test to not use relative data
JIT code and loaded code memory offsets aren't consistent run to run
which causes code generation consistency issues.
Instead use a register plus an immediate that can't fit in a single
16-bit encoding to get similar register allocation to RIP relative
addresses.

At some point it will probably be nice to support consistent placement
of both guest and host code so RIP relative can be used, but we don't have
support in our backend to do that.
2024-01-30 17:40:36 -08:00
Tony Wasserka b40d8f5679 unittests/Library Forwarding: Add tests for fixed-size integer mapping 2024-01-30 19:54:45 +01:00
Ryan Houdek 2884337d85 Merge pull request #3391 from neobrain/fix_libfwd_custom_repack_nonstruct
Library Forwarding: Don't attempt custom repacking for non-struct types
2024-01-30 08:52:21 -08:00
Ryan Houdek 3036f3b3ff Merge pull request #3390 from pmatos/TestHarnessArgs
Check that path arguments to TestHarnessRunner exist
2024-01-30 08:52:14 -08:00
Paulo Matos b02ed40b5a Check that path arguments to TestHarnessRunner exist 2024-01-30 16:37:31 +00:00
Tony Wasserka 0114c874a7 unittests/Library Forwarding: Respect silent flag when compiling SourceWithAST 2024-01-30 17:22:23 +01:00
Tony Wasserka 3ff3dc8769 Library Forwarding: Add unit test for diverging parameter type sizes 2024-01-30 17:22:23 +01:00
Tony Wasserka d04c94fe80 Library Forwarding/gen: Map integers to fixed-size equivalents on guest 2024-01-30 17:22:23 +01:00
Tony Wasserka 0ebf260ed0 Library Forwarding: Consider struct metadata equal if it only differs in integer member type names 2024-01-30 17:22:23 +01:00
Tony Wasserka 28ae84cf60 Library Forwarding: Allow implicit conversions between various integer types
Generally, implicit integer conversions are prohibited for data wrapped in
guest_layout/host_layout, but a few types are exceptional:
* char vs signed char vs unsigned char vs other 8-bit ints
* wchar_t vs other 32-bit ints
* size_t vs uint32_t (32-bit only)
* long long vs other 64-bit ints
* long vs long long (64-bit only)

These combinations have the same data size, so conversions between them are
explicitly allowed now.
2024-01-30 17:22:23 +01:00
Tony Wasserka 26007168b0 Library Forwarding: Extend function pointer interface to take separate guest parameter lists
Some types (notably size_t on 32-bit) have different sizes on the guest than on
the host. This template function must be aware of these differences, so a
second parameter list with fixed-size types must be provided to describe the
guest types.

Note that this information can't be queried through type traits: To a C++
compiler, size_t is indistuingishable from uint64_t. For this reason, the
correct guest type must indeed be provided externally.
2024-01-30 17:22:23 +01:00
Tony Wasserka a48237e218 Library Forwarding: Don't attempt custom repacking for non-struct types 2024-01-30 17:18:54 +01:00
Mai fa3352004e Merge pull request #3381 from alyssarosenzweig/opt/masking
Allow upper garbage on a bunch of instructions
2024-01-30 10:07:53 -05:00
Mai b9378856da Merge pull request #3389 from Sonicadvance1/panic_spill_libnss3
instcountci: Adds panicspill test from steamwebhelper
2024-01-30 10:05:31 -05:00
Paulo Matos 515e98f4cb Update instruction counts 2024-01-30 10:47:56 +00:00
Ryan Houdek ce2924731e vixl/simulator: Enlarge simulator stack size
Simulator stack size defaults to 8KB. This new unit test requires at
least 15360 stack size. Just push it up to 8MB.
2024-01-29 19:48:38 -08:00
Ryan Houdek c099fd01f1 instcountci: Adds panicspill test from steamwebhelper
This block of code causes our RA to panic spill and it gets hit my
steamwebhelper.
We should be able to spill without panic on this code at some point.
2024-01-29 19:48:38 -08:00
Ryan Houdek 36250b10f6 InstCountCI: Sanitize out adrp and adr
First time usage of adrp and adr, need to sanitize it.
2024-01-29 19:32:04 -08:00
Ryan Houdek bc67910ee4 Merge pull request #3382 from pmatos/TypoFix
Fix typos; NFC
2024-01-29 16:10:14 -08:00
Ryan Houdek 79526b9c9e Merge pull request #3379 from neobrain/fix_fexconfig_paths
FEXConfig: Initialize paths before trying to read configuration files
2024-01-29 16:01:12 -08:00
Mai 31a4158957 Merge pull request #3383 from alyssarosenzweig/opt/ptest
Optimize PTEST and VTESTP
2024-01-29 13:30:53 -05:00
Mai 58f3d3caf5 Merge pull request #3380 from alyssarosenzweig/opt/pdep
Optimize PDEP
2024-01-29 13:27:15 -05:00
Alyssa Rosenzweig be52676b3b InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:24:11 -04:00
Alyssa Rosenzweig ae48228943 OpcodeDispatcher: optimize vtestps/vtestpd
I don't really care about AVX but do the same thing we did for vptest.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:24:11 -04:00
Alyssa Rosenzweig e8e35e48c7 OpcodeDispatcher: optimize ptest with tst
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Alyssa Rosenzweig 8b8f27a88f OpcodeDispatcher: optimize ptest with umaxv
to check if the vector is zero, umaxv its elements and check if the reduced
scalar is zero.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:20:17 -04:00
Alyssa Rosenzweig 8e7906a665 IR: add UMaxV
will be used to accelerate ptest

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:19:22 -04:00
Paulo Matos 85105a6af3 Update instcountci json file 2024-01-29 17:18:02 +00:00
Paulo Matos 027fbbf051 Optimize CDQOp 2024-01-29 17:18:02 +00:00
Paulo Matos ca31a0404c ConstProp should generate 32bit constants when required 2024-01-29 17:15:47 +00:00
Paulo Matos f644959c7c Fixing some typos; NFC 2024-01-29 17:14:53 +00:00
Alyssa Rosenzweig 10cca02656 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:13:44 -04:00
Alyssa Rosenzweig 16a54742e6 OpcodeDispatcher: optimize 32-bit tzcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig ad9aa0bc87 OpcodeDispatcher: optimize 32-bit lzcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig bd2b3f35a3 OpcodeDispatcher: optimize 32-bit popcnt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig 50169ce640 OpcodeDispatcher: optimize 32-bit pext
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig baae2d68f9 OpcodeDispatcher: optimize 32-bit bextr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig ad8d038b8a OpcodeDispatcher: optimize 32-bit blsi
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig 820932e3c7 OpcodeDispatcher: optimize 32-bit blsmsk
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig 6f11f2e6f4 OpcodeDispatcher: optimize 32-bit blsr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:11:25 -04:00
Alyssa Rosenzweig 409b6ff6ef InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:08:30 -04:00
Alyssa Rosenzweig f5ad7682c3 OpcodeDispatcher: optimize 32-bit pdep
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:06:56 -04:00
Alyssa Rosenzweig 04805f351b JIT: rewrite pdep implementation
- use better algorithm that is O(# set bits) instead of O(# total bits)
- eliminate spilling by careful management of our temporaries
- fix nzcv clobber bug (whoops)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-29 13:06:56 -04:00
Ryan Houdek 8e3d4a3e02 Merge pull request #3375 from Sonicadvance1/implement_efault_support
Linux: Implements a fault safe memcpy routine
2024-01-26 02:15:42 -08:00
Ryan Houdek 0913741343 Linux: Disable EFAULT handler until we find something that uses it. 2024-01-26 01:54:09 -08:00
Ryan Houdek 929193c16c Linux: Implements support for EFAULT with ppoll's timeout
Only need to handle the timeout structure, the rest is handled in the
kernel itself.
2024-01-26 01:54:09 -08:00
Ryan Houdek 69ff984a82 Implements a FEXLinuxTest for testing ppoll EFAULT behaviour 2024-01-26 01:54:09 -08:00
Ryan Houdek c5ffc0664d Linux: Implements a fault safe memcpy routine
We are required in our syscall emulation to handle cases where pointers
are invalid. This means we need to pessimistically assume a memcpy will
fault when reading application memory.

This implements a signal handler based approach to catching the SIGSEGV
on memcpy and returning an EFAULT if it faults.
2024-01-25 13:50:55 -08:00
Tony Wasserka c1ef11f034 FEXConfig: Initialize paths before trying to read configuration files 2024-01-25 15:54:25 +01:00
Mai 750b0b70bc Merge pull request #3356 from Sonicadvance1/modify_code_lock
Jitarm64: Implements spin-loop futex for JIT blocks
2024-01-23 13:46:59 -05:00
Ryan Houdek 56d8080ec9 Merge pull request #3345 from Sonicadvance1/fix_syscall_registers
OpcodeDispatcher: Fixes syscall rcx/r11 generation
2024-01-22 15:21:13 -08:00
Ryan Houdek c0be974272 Merge pull request #3368 from bylaws/preprcr
FEXCore: Fix RCL/RCR shift wraparound behaviour
2024-01-21 13:44:49 -08:00
Ryan Houdek 0e97f8fb2b Merge pull request #3366 from bylaws/prep222
FEXCore: Use TMP1-4 for values that need preserving across spills
2024-01-21 13:19:34 -08:00
Billy Laws c4c10d0a16 instcountci: update 2024-01-21 18:15:50 +00:00
Billy Laws 576cf61a96 unittests: Test RC{R,L} wraparound behaviour 2024-01-21 18:15:50 +00:00
Billy Laws e323938173 FEXCore: Fix RCL/RCR shift wraparound behaviour
This ends up being cleaner to handle outside of
CalculateFlags_ShiftVariable as constant masking is only needed for
RCL/RCR.
2024-01-21 18:15:50 +00:00
Billy Laws 407e26bfee FEXCore: Use TMP1-4 for values that need preserving across spills
The ARM64EC SRA layout will use x0-3 for x86_64 registers, as such any
arguments passed to C ABI functions need to proxy their arguments
through the temporaries and move as appropriate.
2024-01-21 16:21:13 +00:00
Ryan Houdek d04f3a288e Merge pull request #3377 from neobrain/feature_thunk_flip_nullcheck_order
Library Forwarding: Flip order of nullptr check and custom repacking
2024-01-20 06:47:53 -08:00
Ryan Houdek 3bfb4bedd5 Merge pull request #3374 from neobrain/feature_thunk_test_repack_vs_assume_compatible
Library Forwarding: Test interaction of struct repacking and assume_compatible_data_layout
2024-01-20 00:15:51 -08:00
Tony Wasserka fb15cf0db4 Thunks: Flip order of nullptr check and custom repacking 2024-01-19 11:25:07 +01:00
Tony Wasserka b08d372c78 Thunks: Fix definition of GUEST_THUNK_LIBRARY
The -deps target is the wrong target to add this to, since its compile flags
are propagated to both Guest.cpp and Host.cpp. Instead, define the flag only
when processing files within a guest context.
2024-01-19 11:17:44 +01:00
Tony Wasserka acf48e29b2 FEXLinuxTests/Thunks: Test interaction of struct repacking and assume_compatible_data_layout 2024-01-19 11:17:44 +01:00
Ryan Houdek 162733130c Merge pull request #3376 from Guztaver/patch-1
Improvements to the Dockerfile
2024-01-19 01:53:49 -08:00
Gustavo Muniz f6ff6c8b44 Improvements to the Dockerfile
I added the `pkg-config` that is missing in the Dockerfile, and make the git clone the repository direct from the image, so the user don't need to clone it manually and after that pulls inside the Dockerfile!
2024-01-19 02:27:03 -03:00
Ryan Houdek 6efc4a967c Merge pull request #3373 from neobrain/feature_thunk_custom_repacking
Library Forwarding: Implement assisted struct repacking
2024-01-18 13:56:20 -08:00
Ryan Houdek a6c57f71e9 SpinWaitLock: Fixes potential extra wait that would occur on contended lock
We had a chance of doing an additional bogus wfe if the expected value
was hit in one iteration of a loop. Not the biggest problem on current
hardware where WFE only ever sleeps for 1-4 system cycles, but on future
hardware where WFE might actually sleep for longer then this could have
been an issue.
2024-01-17 10:41:16 -08:00
Ryan Houdek 2af7e997f4 Spinlocks: Fix assembly
Need to have a source be +r so it doesn't get overwritten.
2024-01-17 10:19:38 -08:00
Ryan Houdek ab6c00bbcf FEXCore/Utils: Rename FutexSpinWait to SpinWaitLock 2024-01-17 10:19:38 -08:00
Ryan Houdek e18453cb57 Jitarm64: Implements spin-loop futex for JIT blocks
This will ensure that multiple concurrent SIGBUS handlers in the same
code block doesn't modify the same code.
2024-01-17 10:19:38 -08:00
Ryan Houdek 39f49782da Arm64: Move ParanoidTSO checks up out of the non-paranoid code bath 2024-01-17 10:19:38 -08:00
Ryan Houdek 2c5dd20f3c FutexSpinWait: Implement spin-loop Unique mutex. 2024-01-17 10:19:38 -08:00
Ryan Houdek 136fa78825 FEXCore: Implements an efficient spin-loop API
This will only be used internally inside of FEXCore for efficient shared
codecach backpatch spin-loops.
2024-01-17 10:19:38 -08:00
Tony Wasserka 190d8020ff FEXLinuxTests/Thunks: Add tests for assisted struct repacking 2024-01-15 20:40:13 +01:00
Tony Wasserka 6eaeb48fac Thunks/gen: Implement assisted struct repacking
This can be used to allow automatically handling structures that require
special behavior for one member but are automatically repackable otherwise.
The feature is enabled using the new custom_repack annotation and requires
additional repacking functions to be defined in the host file for each
customized member.
2024-01-15 20:40:13 +01:00
Ryan Houdek f956f008ea Merge pull request #3372 from alyssarosenzweig/opt/cmpxchg-review
Optimize GPR cmpxchg
2024-01-15 05:11:12 -08:00
Ryan Houdek be4d1a8860 Merge pull request #3365 from Sonicadvance1/pidof
Tools: Adds new FEXpidof tool
2024-01-15 05:10:56 -08:00
Ryan Houdek 8ff4b525fc Merge pull request #3367 from bylaws/prepwow
Windows: Commonise WOW64 logic that can be shared with ARM64EC
2024-01-15 05:09:50 -08:00
Tony Wasserka 832072362b Merge pull request #3371 from neobrain/feature_thunk_automatic_repacking
Library Forwarding: Implement automatic argument repacking
2024-01-15 11:23:42 +01:00
Ryan Houdek 1f7a619c79 OpcodeDispatcher: Fixes syscall rcx/r11 generation
Noticed this while writing #3342.

Fixes #3343

The syscall instruction is defined in the documentation that it will set
RCX to the next instruction's RIP and R11 to be RFLAGS. We entirely
skipped this which I noticed while writing unit tests.

Adds unittests to test both 32-bit and 64-bit behaviour because our
helper shares code with both.

I don't know if anything actually relied on this behaviour but we should
definitely support it.
2024-01-12 19:14:30 -08:00
Ryan Houdek ddfd393789 Tools: Adds new FEXpidof tool
This behaves exactly like pidof but only searches for FEX applications.
This fixes a long standing annoyance of mine that pidof doesn't work for
FEX. This behaves exactly like pidof but knows how to decode the command
line options to pull out the program data.

If the Linux kernel ever accepts the patches for binfmt_misc to change
how the interpreter is handled then this will become redundant, but
until that happens here is a utility that I want.
2024-01-12 19:05:37 -08:00
Alyssa Rosenzweig d9985eb877 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-12 12:23:34 -04:00
Alyssa Rosenzweig 58127bd0e8 OpcodeDispatcher: optimize trivial cmpxchgs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-12 12:23:34 -04:00
Alyssa Rosenzweig e8945dfb6d OpcodeDispatcher: optimize gpr cmpxchg
NZCV stuff.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-12 12:03:28 -04:00
Alyssa Rosenzweig 0a6e875329 InstCountCI: add nontrivial gpr cmpxchg cases
cmpxchg handles rax specially, so cmpxchg with dest=rax is a special case. test
also the general case.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-01-12 11:50:43 -04:00
Tony Wasserka cd4578de27 FEXLinuxTests/thunks: Add tests for automatic repacking 2024-01-12 14:47:10 +01:00
Tony Wasserka 295fe03eef unittests/ThunkLibs: Add simple struct repacking tests
These just ensure that two specific but common patterns continue to compile.
2024-01-12 14:47:10 +01:00
Tony Wasserka 4cbaad298a Thunks/gen: Repack out-parameters on exit
To avoid performance traps, several conditions must hold for exit repacking
to apply:
* Argument must be a pointer
* The pointee type must have differing data layout between guest and host
* The pointee type must be non-const

Arguments that don't meet the first two conditions are safe *not* to repack
on exit, since they're either passed by copy or have consistent data layout.
The third condition is a heuristic: In principle, an API function could modify
data through nested pointers even if the argument pointer is const. However,
automatic repacking is not supported for such types anyway, so this is a safe
heuristic to use.
2024-01-12 14:47:10 +01:00
Tony Wasserka dc477c3bd7 Thunks/gen: Implement automatic struct (entry-)repacking 2024-01-12 14:47:10 +01:00
Tony Wasserka 52419d9911 Thunks/gen: Emit metadata to check if a type has consistent data layout 2024-01-12 14:47:10 +01:00
Ryan Houdek 8c3163096b Merge pull request #3363 from Sonicadvance1/fix_label_allocations
ArmEmitter: Support single use forward labels
2024-01-12 00:26:31 -08:00
Billy Laws 9f311cd97e CommonWindows: Split out code invalidation logic from WOW64
This will also be used by FEX ARM64EC module.
2024-01-12 00:12:26 +00:00
Billy Laws 1115ce4a95 CommonWindows: Introduce, and split out CPU feature logic from WOW64
This can also be used by the FEX ARM64EC module.
2024-01-12 00:12:26 +00:00
Billy Laws 03b802cf8e Windows: Use MinGW importlib functions when possible for ntdll
Switch to only defining imports that are not part of MinGW in the FEX importlib
definitions, this is necessary to avoid linker problems with ARM64EC.
2024-01-12 00:12:26 +00:00
Ryan Houdek 615cfe0246 Merge pull request #3361 from Sonicadvance1/decompose_std_function
FEXCore: Decompose some std::function usage to regular pointers
2024-01-10 16:55:29 -08:00
Ryan Houdek dae16aaf18 Merge pull request #3362 from Sonicadvance1/glibc_alloc
Fixes some new glibc allocations that cropped up
2024-01-10 02:05:24 -08:00
Ryan Houdek 3d5f876585 Fixes some new glibc allocations that cropped up
I guess this was handled by brk things before.
2024-01-09 13:55:04 -08:00
Ryan Houdek 37102400b5 Arm64: Switches uses of forward label over to SingleUse if possible
Primary goal for this is to ensure that the delinker doesn't need to
allocate any memory. This delinker can end up getting hit heavily with
JIT code so we don't want it to be allocating memory.
2024-01-08 22:18:20 -08:00
Ryan Houdek c01e6283ae CodeEmitter: Support a single use forward label
Currently all uses of the forward label calls in to jemalloc to allocate
memory. This allows a forward label that doesn't require any memory
allocation, which is the common case in FEX.
2024-01-08 22:18:20 -08:00
Ryan Houdek 248dc97993 FEXCore: Decompose some std::function usage to regular pointers
The delinker step of the JIT was using std::function with capture
lambdas that required memory allocation when unnecessary.
Because the compiler can't see through our std::function usage it could
never decompose these by itself.

By passing the Thread's frame and record to the function as arguments
then we can have the signature be a raw function pointer.

This fixes an area of concern from:
https://github.com/FEX-Emu/FEX/blob/main/docs/ProgrammingConcerns.md#stdfunction-and-lambdas
2024-01-06 19:39:54 -08:00
Ryan Houdek d488592eda Merge pull request #3339 from Sonicadvance1/pass_thread_unaligned_fault_handler
FEXCore: Pass thread object to HandleUnalignedAccess
2024-01-04 18:20:37 -08:00
Ryan Houdek 743df8dfae Merge pull request #3327 from Sonicadvance1/remove_syscall_indirection
Arm64: Removes a vtable indirection in syscalls
2024-01-04 18:19:40 -08:00
Ryan Houdek 4b3792196f Merge pull request #3303 from Sonicadvance1/initial_runtime_longmode_switch
OpcodeDispatcher: Initial support for runtime long-mode switch
2024-01-04 18:17:54 -08:00
Ryan Houdek c333aac4f9 Merge pull request #3354 from Sonicadvance1/uprev_kernel_2
Linux uprev to v6.6
2024-01-03 14:25:13 -08:00
Ryan Houdek db7d7a6bd7 Merge pull request #3349 from Sonicadvance1/revert_frontend_ownership
Revert "FEXLoader: Moves thread management to the frontend"
2024-01-03 14:25:04 -08:00
Alyssa Rosenzweig 04a88ed3ab Merge pull request #3353 from Sonicadvance1/public_interface_cleaning
FEXCore interface cleaning
2024-01-03 15:14:54 -04:00
Alyssa Rosenzweig 9da08b40bd Merge pull request #3344 from Sonicadvance1/xbyak_upstream
Externals: Update xbyak to v7.02 and switch away from fork
2024-01-03 15:13:58 -04:00
Alyssa Rosenzweig 5467c3e478 Merge pull request #3357 from Sonicadvance1/remove_non_sra
FEXCore: Removes SRA option, it's now permanently enabled
2024-01-03 15:10:04 -04:00
Ryan Houdek 9841983955 Merge pull request #3358 from wannacu/main
JIT: Fixes broken register in VTBX1
2023-12-30 00:44:09 -08:00
wannacu 4e7bab849c JIT: Fixes broken register in VTBX1
If the Dst register is allocated as VectorIndices or VectorTable,
using Dst as an operand to perform the tbx operation will result in an error.
For example:
%131(FPR0) i128 = LoadNamedVectorIndexedConstant u8:Tmp:RegisterSize, #0x6, #0xaa0
%132(FPR0) i128 = VTBX1 u8:Tmp:RegisterSize, %129(FPRFixed6) i32v4, %126(FPRFixed10) i16v8, %131(FPR0) i128
Since the tbx instruction's destination register is also the original operand,
this is consistent with the semantics of VTBX1. Therefore,
directly using VectorSrcDst as the destination operand for the tbx instruction is safe.
2023-12-29 16:18:40 +08:00
Ryan Houdek d098545c20 FEXCore: Removes SRA option, it's now permanently enabled 2023-12-28 18:28:02 -08:00
Ryan Houdek 5358af7794 Revert "FEXLoader: Moves thread management to the frontend"
This reverts commit 58f2693954.
2023-12-27 04:33:50 -08:00
Ryan Houdek eea2e7bb57 Merge pull request #3351 from Sonicadvance1/remove_context_lock
FEXCore: Removes context wide and map lookup
2023-12-27 00:59:22 -08:00
Ryan Houdek 25bcddf3a5 FEXCore: Removes context wide and map lookup
While locking a shared_lock and doing an empty table lookup is fairly
fast, just remove them from the hot path entirely if no custom IR
handlers are installed.

This is only used for our IRLoader, which is losing its importance
significantly and should probably be removed anyway.
2023-12-26 11:11:44 -08:00
Ryan Houdek f785b38e4d Merge pull request #3352 from Sonicadvance1/remove_irloader
Removes IRLoader, unittests, and public interface
2023-12-26 11:08:26 -08:00
Ryan Houdek 0071c1bda2 Merge pull request #3350 from Sonicadvance1/minor_opt
FEXCore: Optimize HostFeatures and CPUID feature calculation
2023-12-26 09:01:58 -08:00
Ryan Houdek 058e691ef1 Merge pull request #3348 from Sonicadvance1/disable_sha
Support disabling SHA in CPUID
2023-12-26 08:24:59 -08:00
Tony Wasserka d806db53ec Merge pull request #3318 from neobrain/feature_thunk_layout_wrappers
Library Forwarding: Emit layout helpers to allow repacking struct data
2023-12-26 16:25:53 +01:00
Tony Wasserka bcd0efa724 Thunks: Build fix for clang versions older than 15
Returning incomplete types from deleted functions is valid C++, but clang
did not support it until version 15:
https://github.com/llvm/llvm-project/issues/52802
2023-12-26 16:02:08 +01:00
Tony Wasserka 48c2e0689a Thunks/gen: Specialize layout wrappers for pointer types
Pointer types inherently cause data layout compatibility issues, so they're
worth special-casing here. The wrappers will type-pun pointers to 32-bit or
64-bit integers (matching the guest architecture) to avoid direct host-side
use of guest pointers without consideration.
2023-12-26 16:02:08 +01:00
Tony Wasserka 4b09a6bee1 Thunks/gen: Implement host_layout->guest_layout conversion
This enables use of guest_layout for return values.
2023-12-26 16:02:08 +01:00
Tony Wasserka af645cb750 Thunks/gen: Fill guest_layout stub and implement conversion to host_layout
The guest_layout wrapper provides an architecture-agnostic representation of
the guest data layout of each struct used in a thunked library. A constructor
is added to host_layout to allow conversion of the data to the host layout.

For types that are already fully compatible, both layout wrappers are simple
type aliases to minimize overhead.
2023-12-26 16:02:07 +01:00
Ryan Houdek b8b9b2f1a7 Linux: Update syscalls for v6.6 2023-12-25 08:40:46 -08:00
Ryan Houdek f02bd5c9a6 Linux/drm: Update to v6.6 2023-12-25 08:28:45 -08:00
Ryan Houdek 2b39813581 Linux: Implement v6.5 syscall that was added
Fairly easy since the arguments can just be passed through.

The rest of the v6.5 changes don't seem to be anything worth caring
about.
2023-12-25 08:17:43 -08:00
Ryan Houdek 923c53c4ea Linux: New socket options 2023-12-25 08:11:15 -08:00
Ryan Houdek 66658d82c6 Linux: Uprev kernel to 6.4
Added a new prctl that we need to emulate.
Some ioctls have changed that we don't really care about or will
naturally work.
2023-12-25 07:59:03 -08:00
Ryan Houdek b115c144fb FEXCore: Removes NetStream from public API
Only used by GDBServer.
NFC.
2023-12-25 07:07:17 -08:00
Ryan Houdek d8f20751fe FEXCore: Moves IREmitter from the public API to backend
No functional change
2023-12-25 07:00:29 -08:00
Ryan Houdek 1977747fc2 Removes IRLoader, unittests, and public interface
This unit test hasn't really served any purpose for a while now and
mostly just causes pain when reworking things in the IR.

Just remove the IRLoader, its unit tests, the github action steps and
the public FEXCore interface to it. Since it isn't used by anything
other than Thunks.

Also moves some IR definitions from the public API to the backend.
2023-12-25 07:00:29 -08:00
Ryan Houdek 257016bf12 FEXCore: Moves BucketList out of public API
NFC
2023-12-25 06:58:22 -08:00
Ryan Houdek 69d65fba4a FEXCore: Removes unused SyscallVisitor
This was expected to be part of the syscall optimizations we did but
ended up getting manifested in a different way. Remove it.
2023-12-25 06:42:11 -08:00
Ryan Houdek bce694ebb5 FEXCore: Moves BitUtils to FHU
No functional change
2023-12-25 06:38:51 -08:00
Ryan Houdek 5d37d5db1a FEXCore: Optimize HostFeatures and CPUID feature calculation
Need #3348 merged first.

As I was casually thinking, this code made me realize that it was quite
branch heavy and could likely be optimized to logic.

The previous code generated some fairly nasty branch heavy code. This
can be optimized to be branchless and take roughly five instructions
per flag. Using a bitfield for each feature would turn each calculation
in to 3-4 instructions but that seems overkill.

Very minor thing.
2023-12-25 04:58:15 -08:00
Ryan Houdek 4d109c9ce0 Config: Fixes parsing strenum inside of json files
This wasn't wired up before.
2023-12-23 22:32:59 -08:00
Ryan Houdek db9b326534 FEXCore: Support disabling CPUID features based on config
Need to be able to disable sha by config.
2023-12-23 22:32:29 -08:00
Ryan Houdek 21ae8336cd Update issue templates
Fex doesn't support x86 anymore
2023-12-22 18:03:11 -08:00
Tony Wasserka c6f8901c16 Thunks/gen: For repackable/compatible structs, use a richer host_layout/guest_layout interface 2023-12-22 20:55:15 +01:00
Tony Wasserka a797699d62 Thunks/gen: Sort types by dependency before emitting data
This allows future changes to emit interdependent helper structures in the
same order.

The sort algorithm was chosen for simplicity rather than performance. It's
fast enough in practice even for APIs as large as Vulkan.
2023-12-22 20:55:15 +01:00
Tony Wasserka 36c524a021 Thunks: Remove callback_guest and fex_guest_function_ptr
These are superseeded by ptr_passthrough and the guest_layout wrapper.
2023-12-22 20:55:15 +01:00
Tony Wasserka dbf33523bd unittests/thunks: Fix build on newer libclang versions 2023-12-22 20:55:15 +01:00
Ryan Houdek de2cd469a2 Merge pull request #3346 from Sonicadvance1/remove_32bit_allocator
FEX: Removes legacy kernel 32-bit allocator
2023-12-22 10:35:49 -08:00
Ryan Houdek 1c34b25538 FEX: Removes legacy kernel 32-bit allocator
We only used this so that our Xavier CI system which were running old
kernels could run unit tests. We have now removed the Xaviers from CI
and this is no longer necessary.

Stop pretending that we support kernels older than 5.0 and allowing this
fallback.

The 32-bit allocator is still used for the MAP_32BIT mmap flag, so the
load bearing code can't be fully removed. Just remove the config and the
frontend things using it.
2023-12-21 06:21:01 -08:00
Ryan Houdek 38ad3f0e05 FEXCore: Pass thread object to HandleUnalignedAccess
Currently no functional change but public API breaks should come early.

The thread state object will be used for looking up thread specific
codebuffers in the future when we support MDWE with code mirrors.
2023-12-21 01:55:25 -08:00
Ryan Houdek 266f7feecb Arm64: Removes a vtable indirection in syscalls
We can safely call virtual functions through the JIT with a little bit
of work.

FEX's JIT has quite a few steps before it gets to a syscall handler.

Before this commit:
JIT->static HandleSyscall->SyscallHandler::HandleSyscall->SyscallHandler

After this commit:
JIT->SyscallHandler::HandleSyscall->SyscallHandler

A bit hard to notice this when this interface can spin at 67-million
calls per second though.
2023-12-21 01:55:02 -08:00
Ryan Houdek f9902142f7 Utils: Add ability to get VTable entries to PMF helper
This will be useful to remove an indirection.
2023-12-21 01:55:02 -08:00
Ryan Houdek 175d407c9f InstCountCI: Update for segment changes 2023-12-21 01:54:19 -08:00
Ryan Houdek ae2f98e017 Mingw: Update for GDT 2023-12-21 01:54:19 -08:00
Ryan Houdek 9e5d7aa5fe OpcodeDispatcher: Initial support for runtime long-mode switch
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.

Adds some asserts since currently it is unexpected if the configuration
changes at runtime.

This is fairly straightforward for an initial setup but isn't fully
fleshed out.

Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.

Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
2023-12-21 01:54:19 -08:00
Ryan Houdek 8648fb1485 FEXCore: Accurately store segment descriptors
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.

In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.

Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
2023-12-21 01:54:18 -08:00
Ryan Houdek 8b24f7fc26 Externals: Update xbyak to v7.02 and switch away from fork
The last few patches we need have been upstreamed so we shouldn't need
our downstream fork anymore.
2023-12-21 01:52:05 -08:00
Ryan Houdek b1c3737aef Merge pull request #3342 from Sonicadvance1/fix_proton_3
Fixes Proton again
2023-12-20 23:15:10 -08:00
Ryan Houdek 00669a1c89 Merge pull request #3336 from Sonicadvance1/warn_on_mdwe
FEXCore: Warn if MDWE is set
2023-12-20 13:27:43 -08:00
Ryan Houdek 1cedc3d85a FEXCore: Warn if MDWE is set
This will result in FEX not being able to allocate executable memory.
We can use shared memory in the future to work around this but for now
we don't support that as a fix.
2023-12-20 13:19:28 -08:00
Mai 5e26b77a7c Merge pull request #3326 from Sonicadvance1/thread_frontend_ownership_part3
FEXLoader: Moves thread management to the frontend
2023-12-19 21:05:18 -05:00
Ryan Houdek 58f2693954 FEXLoader: Moves thread management to the frontend
Lots going on here.

This moves OS thread object lifetime management and internal thread
state lifetime management to the frontend. This causes a bunch of thread
handling to move from the FEXCore Context to the frontend.

Looking at `FEXCore/include/FEXCore/Core/Context.h` really shows how
much of the API has moved to the frontend that FEXCore no longer needs
to manage. Primarily this makes FEXCore itself no longer need to care
about most of the management of the emulation state.

A large amount of the behaviour moved wholesale from Core.cpp to
LinuxEmulation's ThreadManager.cpp. Which this manages the lifetimes of
both the OS threads and the FEXCore thread state objects.

One feature lost was the instruction capability, but this was already
buggy and is going to be rewritten/fixed when gdbserver work continues.

Now that all of this management is moved to the frontend, the gdbserver
can start improving since it can start managing all thread state
directly.
2023-12-19 17:43:04 -08:00
Mai b4b8e81f24 Merge pull request #3321 from Sonicadvance1/thread_frontend_ownership_take2
FEXCore: Changes ParentThread ownership from the CTX to the frontend, take 2
2023-12-19 20:37:59 -05:00
Mai 82ce76b1d3 Merge pull request #3325 from Sonicadvance1/move_pmf
Dispatcher: Convert GetCompileBlockPtr to using PMF helper
2023-12-19 20:33:53 -05:00
Ryan Houdek a8f797d36b Dispatcher: Convert GetCompileBlockPtr to using PMF helper
This was older code that was written before the PMF helper was
available.
Switch it over.
2023-12-19 17:20:37 -08:00
Ryan Houdek 93ec676ce8 Merge pull request #3340 from Sonicadvance1/exitfunctionlink_data
FEXCore: Describe exit function linking object with a structure
2023-12-19 16:17:21 -08:00
Ryan Houdek 362cccc210 Syscalls: Fixes invalid register invalidation
FEX was incorrectly invalidating RBX and RBP when a new
thread/clone/fork was created. Only RAX should be set to zero to
signify that it is the child process.

I tried to find some reference as to why this was being done in FEX, and
I seem to recall some code somewhere that was setting up this state.

Sadly I can't seem to find any reference to this being required now by
the linux kernel when it creates new processes, so maybe I was just
gaslighting myself.

Fixes crash in bubblewrap where it was expecting a valid RBP to exist
after it does a clone syscall.

I have unit tests coming for this but it is uncovering some more bugs in
FEX's thread handling. Some of which are getting fixed with the current
PRs that are moving the thread handling to the frontend.

But since proton is broken now, this needs to get in.
2023-12-19 16:01:21 -08:00
Ryan Houdek c45dc6dd48 Syscalls: Fixes StackSize argument usage
This was left over and was using the wrong argument.
This would cause the stack for the child thread/process to be at the
start instead of the end.

Fixes a crash in bubblewrap.
2023-12-19 15:56:43 -08:00
Mai 3d2cbc5d08 Merge pull request #3317 from Sonicadvance1/fix_imul_flags2
OpcodeDispatcher: Fixes flags generation in imul
2023-12-19 11:43:21 -05:00
Mai 81c85d73b2 Merge pull request #3330 from Sonicadvance1/optimize_sib_addr_calc
OpcodeDispatcher: Optimize SIB addr calculation
2023-12-19 11:41:42 -05:00
Mai 5b4e9c6907 Merge pull request #3323 from Sonicadvance1/remove_unused_check
Dispatcher: Removes unused asserting CompileBlock function
2023-12-19 11:38:57 -05:00
Ryan Houdek aa2e8704bc FEXCore: Changes ParentThread ownership from the CTX to the frontend, take 2
Similar to #3284 but works around some of the bugs that one introduced.

This is the minimal amount of changes to move the ownership from FEXCore
to the frontend. Since the frontends don't yet have a full thread state
tracking, there is an opaque pointer that needs to be managed.

In the followup commits this will be changed to have the syscall handler
to be the thread object manager.
2023-12-18 14:54:07 -08:00
Ryan Houdek cf86ae6b65 FEXCore: Describe exit function linking object with a structure
Instead of just poking raw uint64_t data values, describe it with a
struct.

This will be a read-only in the future.
2023-12-18 13:31:53 -08:00
Ryan Houdek 68d6cf5f14 Merge pull request #3332 from Sonicadvance1/move_testharness_runner
TestHarnessRunner: Move to its own tool folder
2023-12-18 12:30:31 -08:00
Ryan Houdek c7bc15785b FEXLoader: Moves mingw check up one level
Now that only FEXLoader is under the FEXLoader folder, the mingw build
check can be moved up one level.

No functional change, just moving the check and re-aligning. Might be
good to view without whitespace changes.
2023-12-18 12:16:11 -08:00
Ryan Houdek c3d2d01c1f TestHarnessRunner: Move to its own tool folder
Needs #3331 merged first

No functional change, just code movement.
2023-12-18 12:16:11 -08:00
Ryan Houdek 9dda960529 Merge pull request #3331 from Sonicadvance1/move_irloader
Tools: Moves IRLoader to independent folder
2023-12-18 12:15:15 -08:00
Ryan Houdek f2bcc14cda Tools: Moves IRLoader to independent folder
Which requires moving LinuxEmulation to its own independent folder as
well. Since both IRLoader and FEXLoader rely on it.

No functional change, just moves the the code around.
2023-12-18 12:04:28 -08:00
Ryan Houdek 86654907bf Merge pull request #3334 from Sonicadvance1/remove_old_x86jit_references
FEXCore: Removes stale references to x86 JIT
2023-12-18 04:15:29 -08:00
Ryan Houdek 4333261639 Merge pull request #3329 from Sonicadvance1/vdso_fix_symbols
VDSO: Fixes symbol fetching for a few symbols
2023-12-18 04:15:10 -08:00
Ryan Houdek 12b72f908b Merge pull request #3335 from Sonicadvance1/remove_internalthreadstate_header
FEXCore: Removes old InternalThreadState header
2023-12-18 04:14:57 -08:00
Ryan Houdek f5997a084e Merge pull request #3333 from Sonicadvance1/remove_cpuid_init
CPUID: Removes Init and just uses constructor
2023-12-18 04:14:36 -08:00
Ryan Houdek 6c8a54ff84 FEXCore: Removes old InternalThreadState header
This was a temporary header to help with when this header was migrated
to our public API headers.

It's temporary nature is no longer necessary, just get rid of it.
2023-12-15 18:51:25 -08:00
Ryan Houdek bcc2901d7f FEXCore: Removes stale references to x86 JIT
It doesn't exist anymore.
2023-12-15 18:46:57 -08:00
Ryan Houdek 358bbb51ff CPUID: Removes Init and just uses constructor
No need to wait for initialization on for this anymore.
Ever since Init was refactored to do basically no work, this hasn't been
necessary.

CPUID does need to still be initialized after HostFeatures though, so
need to ensure correct member ordering there.
2023-12-15 18:43:23 -08:00
Ryan Houdek f2da70c7e7 InstCountCI: Update for SIB optimization
Only LEA uses SIB scale*index+base.
2023-12-15 13:11:42 -08:00
Ryan Houdek 1a2f41922c OpcodeDispatcher: Optimize SIB addr calculation
When the address calculation for SIB has both index and base then we can
optimize this to an add with a shifted register. This will convert a
three instruction sequence in to one instruction in most cases.
2023-12-15 13:08:46 -08:00
Ryan Houdek 0ede707e0b IR: Adds support for AddShift IR op
This matches x86 SIB's operation of `scale * index + base`
2023-12-15 13:08:06 -08:00
Ryan Houdek c4c31d4e80 VDSO: Fixes symbol fetching for a few symbols
When originally written I was testing symbol loading on an x86 host
assuming that the vdso symbol names were the same between x86 and
aarch64.

Turns out this is not true and aarch64 uses `__kernel_` prefix instead
of x86's `__vdso_`.

AArch64 VDSO only provides us with four symbols right now
- __kernel_rt_sigreturn
- __kernel_gettimeofday
- __kernel_clock_gettime
- __kernel_clock_getres

This means that currently AArch64 is missing getcpu and time.
time is unlikely to ever be implemented so the vdso scanning for that is
unlikely to ever get a pointer and will need to go down the glibc path.

getcpu is likely to get an implementation at some point. There's even an
old patch series that implemented it[1]. Currently the patch has been
dropped on the floor for whatever reason but hopefully they will come
back to it.

Benchmarks on Cortex-X1C:
- Native AArch64 clock_gettime bench
   - VDSO: 33.3ns per call
   - glibc (which uses vdso): 33.3ns per call
   - Roughly equivalent
- Emulated x86_64 clock_gettime bench
   - Prior to fix:
      - 49ns per call for both VDSO and glibc
   - After fix:
      - VDSO: ~46ns per call
      - glibc: ~49ns per call
- Emulated i386 clock_gettime bench
   - Only tested after fix
   - VDSO: 50ns per call
   - glibc: 65ns per call

The only way to improve this further would be to optimize thunks to push
certain function signature arguments in registers instead of stack.
Which is a future optimization as time goes on.

[1] https://patchwork.kernel.org/project/linux-kselftest/list/?series=335203
2023-12-15 10:56:58 -08:00
Ryan Houdek 12923ba1b7 Merge pull request #3322 from Sonicadvance1/remove_unused_exithandler
PassManager: Removes unused exit handler
2023-12-14 01:55:06 -08:00
Ryan Houdek bd13052708 Merge pull request #3315 from Sonicadvance1/clone_fix_allocation
Linux: Adjust when clone allocates stack memory
2023-12-12 07:31:04 -08:00
Ryan Houdek ec89a00a92 Merge pull request #3320 from Sonicadvance1/consteval_x86_tables
X86Tables: Converts tables to be mostly consteval
2023-12-11 23:03:50 -08:00
Ryan Houdek e657a27607 Dispatcher: Removes unused asserting CompileBlock function
While we were calling this function, its asserting nature hasn't been
used for a long time.

This used to trigger more frequently when CompileBlock would fail to
compile code, either due to not being able to decode an instruction or
hitting an instruction that FEX doesn't understand.

When these cases are hit today we still generate code blocks which
generate SIGILL. This means that this code was actually never hit.

Completely remove this function and have the JIT's dispatcher call the
CompileBlock function directly. Signature is slightly different since we
need to set x3 to be 0.
2023-12-11 17:20:51 -08:00
Ryan Houdek 8bb5462554 PassManager: Removes unused exit handler
git blame shows that 718b3e6b4c added this
handler.

It doesn't explain why this was desired but it was never wired up to
anything. Just remove it.
2023-12-11 17:14:41 -08:00
Ryan Houdek 98f21a2a28 X86Tables: Converts tables to be mostly consteval
Reduces the ELF's VM size from 9.8MB down to 9.37MB and should reduce
initialization time a smidge.

Slammed this out while waiting for other PRs to get reviewed.
2023-12-11 10:03:52 -08:00
Ryan Houdek 0a4e064da4 Merge pull request #3296 from Sonicadvance1/mov_thread_creation
FEXCore: Moves OS thread creation to the frontend
2023-12-11 06:35:26 -08:00
Ryan Houdek 5660065eea FEXCore: Moves OS thread creation to the frontend
Fairly lightweight since it is almost 1:1 transplanting the code from
FEXCore in to the SyscallHandler's thread creation code.

Minor changes:
- ExecutionThreadHandler gets freed before executing the thread
   - Saves 16-bytes of memory per thread
- Start all threads paused by default
   - Since I moved the code to the frontend, I noticed we needed to do
     some post thread-creation setup.
   - Without the pause we were racing code execution with TLS setup and
     a few other things.
2023-12-11 06:22:50 -08:00
Ryan Houdek 131bf48f4b Linux: Adjust when clone allocates stack memory
clone needs to allocate memory before the fork locks are held, otherwise
they will hang forever waiting on a locked mutex.

This wasn't previously seen since nothing was using fork with clone3.
2023-12-11 06:21:14 -08:00
Ryan Houdek a1cf14f2d9 Merge pull request #3319 from Sonicadvance1/fix_gdbserver_crash
GdbServer: Fixes crash on gdb detach
2023-12-11 06:19:40 -08:00
Ryan Houdek 499a40ecb5 GdbServer: Fixes crash on gdb detach
GdbServer object was getting deleted before the socket thread was
shutdown which was causing a crash on detach.

Now on destructor it will wait for the thread to thread to exit.
2023-12-08 06:03:58 -08:00
Ryan Houdek 7524029a06 Merge pull request #3294 from Sonicadvance1/mov_xid_check
FEXCore: Moves XID check to the frontend
2023-12-07 01:17:45 -08:00
Ryan Houdek fe3df001e2 InstCountCI: Update for instruction changes. 2023-12-07 01:10:09 -08:00
Ryan Houdek 2053c712a5 unittests/ASM: Adds test to test imul flag calculation
This was incorrect before.
2023-12-07 01:09:39 -08:00
Ryan Houdek 5c6f229e76 OpcodeDispatcher: Fixes flags generation imul
On overflow with 32-bit we weren't setting the flags correctly.
2023-12-07 01:08:02 -08:00
Ryan Houdek acdb4c7061 IR: Adds support for {S,U}Mull
Lets us do a 32-bit multiply returning a 64-bit result, signed and
unsigned.
2023-12-07 01:06:58 -08:00
Ryan Houdek b613576ac4 Merge pull request #3314 from Sonicadvance1/more_installscript
Scripts: More changes to InstallFEX script
2023-12-06 23:57:19 -08:00
Ryan Houdek 8f52d8e719 Scripts: More changes to InstallFEX script
A dictionary lookup was only ever returning the .ero result since we
were adding multiple values to the same key. Adds them as a list now and
walks them.

Additionally if we forget to add a new version again, have a fallback to
try checking using our file format layout.

Fixes an issue where a user running 23.10 was faulting the python
script.
2023-12-06 23:50:43 -08:00
Alyssa Rosenzweig 0fe5e3d1e7 Merge pull request #3313 from alyssarosenzweig/opt/andn-mask
rm andn masking
2023-12-06 12:21:24 -04:00
Ryan Houdek 26c9d5d427 Merge pull request #3310 from Sonicadvance1/disable_clear_sighand
FEXLoader: Temporarily disable CLONE_CLEAR_SIGHAND
2023-12-06 00:04:16 -08:00
Ryan Houdek 60b0852cde Merge pull request #3309 from neobrain/fix_thunks_32bit_addresses
Thunks: Add workarounds for pointers not readable by 32-bit guests
2023-12-05 23:52:59 -08:00
Ryan Houdek 3a5ac39bb1 Merge pull request #3312 from neobrain/fix_format_strings
FEXLoader: Fix incorrect format strings
2023-12-05 23:51:00 -08:00
Ryan Houdek 301310329b FEXLoader: Temporarily disable CLONE_CLEAR_SIGHAND
This flag breaks FEX heavily for now.
glibc 2.38 started using this flag as an optimization for posix_spawn.
It will fall back to a "non-optimized" implementation if the clone
syscall returns EINVAL. For now do this while we investigate a more
proper implementation.

Should be backported to 2312.1.
2023-12-05 23:46:48 -08:00
Ryan Houdek 86c6ca3049 Merge pull request #3307 from neobrain/fix_vulkan_query_through_instance
Thunks: Allow querying customized functions through vkGetInstanceProcAddr
2023-12-05 23:40:03 -08:00
Ryan Houdek eb5cf1a667 Merge pull request #3311 from Sonicadvance1/update_install_script
Scripts: Updates InstallFEX with supported Ubuntu versions
2023-12-05 19:35:06 -08:00
Alyssa Rosenzweig f515b1eeb1 Merge pull request #3308 from neobrain/fix_armops_warning
Arm64/BranchOps: Fix unused-variable warning
2023-12-05 09:07:07 -04:00
Alyssa Rosenzweig 924cfb5429 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-12-05 09:05:25 -04:00
Alyssa Rosenzweig 2a2c389be6 OpcodeDispatcher: rm masking in 32-bit andn
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-12-05 09:04:32 -04:00
Tony Wasserka aaef83b777 FEXLoader: Fix incorrect format strings 2023-12-05 12:04:16 +01:00
Ryan Houdek 29d82cda48 Scripts: Updates InstallFEX with supported Ubuntu versions
Forgot to update this as our supported Ubuntu distros changed.
2023-12-05 02:44:17 -08:00
Tony Wasserka 9f721e96db Thunks: Work around argument data on the host stack not being accessible by 32-bit guests 2023-12-05 11:32:09 +01:00
Tony Wasserka 281be3aeeb FEXLoader: Exhaust internal jemalloc allocations after 32-bit address space blocking 2023-12-05 11:27:39 +01:00
Tony Wasserka 5b87307452 FEXLoader: Cleanup 32-bit allocator setup 2023-12-05 11:24:51 +01:00
Tony Wasserka aebabf6ef2 Revert "FEXLinuxTests: Temporarily limit thunk test execution to 64-bit guests"
This reverts commit 21b6cccb4e.
2023-12-04 14:57:43 +01:00
Tony Wasserka 0d5fd5dbaf FEXLinuxTests/thunks: Disable tests involving pointers on 32-bit
Passing pointers across architecture boundaries is not yet supported.
Skipping these tests allows for the rest to run on 32-bit.
2023-12-04 14:57:43 +01:00
Tony Wasserka 3729b02255 Arm64/BranchOps: Fix unused-variable warning 2023-12-04 11:32:09 +01:00
Tony Wasserka d4e9ed8e61 Thunks/vulkan: Use string_view literals to express string comparisons
This reduces syntax noise as more functions will be added to this list.
2023-12-04 11:09:25 +01:00
Tony Wasserka c10402f4f9 Thunks/vulkan: Move custom impl matching to common function 2023-12-04 11:09:25 +01:00
Ryan Houdek afebf73e8f FEXCore: Moves XID check to the frontend
The XID signal handler is OS specific and needs to be delegated to being
handled in the frontend. This only happens when we create a thread with
pthread rather than clone/fork/vfork.

The first created pthread will reset the XID signal handler in glibc,
which we need to check the first time this occurs. A little bit disjoint
while the actual thread creation still happens in the backend but will
come together once the thread creation gets moved to the frontend.

The only game I am aware of that requires this to work is SOMA since its
one of its middleware libraries tries asking for root privileges.

Retested the game to make sure it still works as expected.
2023-11-28 19:19:06 -08:00
396 changed files with 27030 additions and 37649 deletions

No files matched your search

@@ -37,7 +37,6 @@ If applicable, add screenshots and video to help explain your problem.
**Additional context**
- Is this an x86 or x86-64 game: [x86/x86-64/Both]
- Does this reproduce on x86-64 host with FEX: [Yes/No/Untested]
- Does this reproduce on AArch64 with Radeon/Intel/Nvidia: [Yes/No/Untested]
- Is this a Vulkan game: [Yes/No/Unknown]
- If Yes, What is your Vulkan driver:
-13
View File
@@ -13,7 +13,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_FORCE32BITALLOCATOR: 1
FEX_ENABLEAVX: 1
jobs:
@@ -78,18 +77,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
-13
View File
@@ -20,7 +20,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_FORCE32BITALLOCATOR: 1
FEX_ENABLEAVX: 1
jobs:
@@ -85,18 +84,6 @@ jobs:
shell: bash
run: cmake --build . --config $BUILD_TYPE --target install
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: gcc target tests 64
working-directory: ${{runner.workspace}}/build
shell: bash
-12
View File
@@ -100,18 +100,6 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM128bit.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target ir_tests
- name: IR Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_IR.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
+6 -6
View File
@@ -15,19 +15,19 @@
path = External/tiny-json
url = https://github.com/Sonicadvance1/tiny-json.git
[submodule "External/xbyak"]
shallow = true
shallow = true
path = External/xbyak
url = https://github.com/FEX-Emu/xbyak.git
url = https://github.com/herumi/xbyak.git
[submodule "External/fex-posixtest-bins"]
shallow = true
shallow = true
path = External/fex-posixtest-bins
url = https://github.com/FEX-Emu/fex-posixtest-bins.git
[submodule "External/fex-gvisor-tests-bins"]
shallow = true
shallow = true
path = External/fex-gvisor-tests-bins
url = https://github.com/FEX-Emu/fex-gvisor-tests-bins.git
[submodule "External/fex-gcc-target-tests-bins"]
shallow = true
shallow = true
path = External/fex-gcc-target-tests-bins
url = https://github.com/FEX-Emu/fex-gcc-target-tests-bins.git
[submodule "External/jemalloc"]
@@ -41,7 +41,7 @@
url = https://github.com/FEX-Emu/drm-headers.git
[submodule "External/xxhash"]
path = External/xxhash
url = https://github.com/FEX-Emu/xxHash.git
url = https://github.com/Cyan4973/xxHash.git
[submodule "External/Catch2"]
path = External/Catch2
url = https://github.com/catchorg/Catch2.git
+9 -2
View File
@@ -122,6 +122,11 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
add_definitions(-D_M_ARM_64=1)
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^arm64ec")
set(_M_ARM_64EC 1)
add_definitions(-D_M_ARM_64EC=1)
endif()
if (ENABLE_CCACHE)
find_program(CCACHE_PROGRAM ccache)
if(CCACHE_PROGRAM)
@@ -232,8 +237,10 @@ endif()
find_package(PkgConfig REQUIRED)
find_package(Python 3.0 REQUIRED COMPONENTS Interpreter)
add_subdirectory(External/xxhash/)
include_directories(External/xxhash/)
set(XXHASH_BUNDLED_MODE TRUE)
set(XXHASH_BUILD_XXHSUM FALSE)
set(BUILD_SHARED_LIBS OFF)
add_subdirectory(External/xxhash/cmake_unofficial/)
add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
-5
View File
@@ -1,5 +0,0 @@
{
"Config": {
"AdditionalArguments": "--no-sandbox"
}
}
+5 -4
View File
@@ -3,11 +3,12 @@ FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-10 llvm-10 nasm ninja-build \
libcap-dev libglfw3-dev libepoxy-dev python3-dev \
python3 linux-headers-generic
clang-10 llvm-10 nasm ninja-build pkg-config \
libcap-dev libglfw3-dev libepoxy-dev python3-dev libsdl2-dev \
python3 linux-headers-generic \
git
COPY . /opt/FEX
RUN git clone --recurse-submodules https://github.com/FEX-Emu/FEX.git
CMD [ "mkdir /opt/FEX/build" ]
+1 -1
+1 -1
+1 -1
+3 -2
View File
@@ -30,10 +30,11 @@ set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
void Testy() {
int Testy(int a, int b, int c, int d, int e, int f) {
return a + b + c + d + e + f;
}
int main() {
return 0;
return Testy(0, 1, 2, 3, 4, 5);
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
+21
View File
@@ -441,6 +441,24 @@ def print_parse_envloader_options(options):
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def print_parse_jsonloader_options(options):
output_argloader.write("#ifdef JSONLOADER\n")
output_argloader.write("#undef JSONLOADER\n")
output_argloader.write("if (false) {}\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (KeyName == \"{0}\") {{\n".format(op_key))
output_argloader.write("Set(KeyOption, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View));\n".format(op_key, op_key, op_key))
output_argloader.write("}\n")
output_argloader.write("else {{\n".format(op_key))
output_argloader.write("Set(KeyOption, ConfigString);\n")
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def print_parse_enum_options(options):
output_argloader.write("#ifdef ENUMDEFINES\n")
output_argloader.write("#undef ENUMDEFINES\n")
@@ -556,6 +574,9 @@ print_parse_argloader_options(options);
# Generate environment loader code
print_parse_envloader_options(options);
# Generate json loader code
print_parse_jsonloader_options(options);
# Generate enum variable options
print_parse_enum_options(options);
+2 -2
View File
@@ -7,6 +7,7 @@ set (FEXCORE_BASE_SRCS
Utils/FileLoading.cpp
Utils/ForcedAssert.cpp
Utils/LogManager.cpp
Utils/SpinWaitLock.cpp
)
if (NOT MINGW_BUILD)
@@ -149,7 +150,6 @@ set (SRCS
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/InlineCallOptimization.cpp
Utils/NetStream.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
@@ -194,7 +194,7 @@ endif()
# Some defines for the softfloat library
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
set (LIBS fmt::fmt vixl xxhash FEXHeaderUtils)
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils)
if (NOT MINGW_BUILD)
list (APPEND LIBS dl)
+1 -1
View File
@@ -1,10 +1,10 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/sstream.h>
#include <FEXCore/fextl/string.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <cmath>
#include <cstring>
+25 -19
View File
@@ -77,7 +77,9 @@
"ENABLECRYPTO": "enablecrypto",
"DISABLECRYPTO": "disablecrypto",
"ENABLERPRES": "enablerpres",
"DISABLERPRES": "disablerpres"
"DISABLERPRES": "disablerpres",
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi"
},
"Desc": [
"Allows controlling of the CPU features in the JIT.",
@@ -97,7 +99,28 @@
"\t{enable,disable}flagm: Will force enable or disable flagm even if the host doesn't support it",
"\t{enable,disable}flagm2: Will force enable or disable flagm2 even if the host doesn't support it",
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it"
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it"
]
},
"CPUID": {
"Type": "strenum",
"Default": "FEXCore::Config::CPUID::OFF",
"Enums": {
"ENABLESHA": "enablesha",
"DISABLESHA": "disablesha"
},
"Desc": [
"Allows controlling of the CPU features are exposed in CPUID.",
"\toff: Default CPU features queried from CPU features",
"\t{enable,disable}sha: Will force enable or disable sha even if the host doesn't support it"
]
},
"SmallTSCScale": {
"Type": "bool",
"Default": "true",
"Desc": [
"Scales the cycle counter on systems that have low frequencies."
]
}
},
@@ -247,23 +270,6 @@
"Disables optimizations passes for debugging."
]
},
"SRA": {
"Type": "bool",
"Default": "true",
"Desc": [
"Set to false to disable Static Register Allocation"
]
},
"Force32BitAllocator": {
"Type": "bool",
"Default": "false",
"Desc": [
"Forces use of the 32-bit allocator on 32-bit applications",
"Used to work around ulimit problems of CI runner",
"Potentially useful for debugging memory problems",
"32-bit allocator is always used if your host kernel is older than 4.17"
]
},
"GlobalJITNaming": {
"Type": "bool",
"Default": "false",
@@ -12,10 +12,6 @@
#include <string.h>
#include <utility>
namespace FEXCore::HLE {
class SyscallVisitor;
}
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
X86Tables::InitializeInfoTables(Mode);
@@ -34,10 +30,6 @@ namespace FEXCore::Context {
return CustomExitHandler;
}
void FEXCore::Context::ContextImpl::Stop() {
Stop(false);
}
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
@@ -46,10 +38,6 @@ namespace FEXCore::Context {
CompileBlock(Thread->CurrentFrame, GuestRIP, MaxInst);
}
bool FEXCore::Context::ContextImpl::IsDone() const {
return IsPaused();
}
void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) {
CustomCPUFactory = std::move(Factory);
}
+30 -80
View File
@@ -44,7 +44,6 @@ namespace CodeSerialize {
namespace CPU {
class Arm64JITCore;
class X86JITCore;
class Dispatcher;
}
namespace HLE {
@@ -69,28 +68,30 @@ namespace FEXCore::Context {
MODE_SINGLESTEP = 1,
};
struct ExitFunctionLinkData {
uint64_t HostBranch;
uint64_t GuestRIP;
};
using BlockDelinkerFunc = void(*)(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record);
constexpr uint32_t TSC_SCALE = 128;
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class ContextImpl final : public FEXCore::Context::Context {
public:
// Context base class implementation.
FEXCore::Core::InternalThreadState* InitCore(uint64_t InitialRIP, uint64_t StackPointer) override;
bool InitCore() override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
void Pause() override;
void Run() override;
void Stop() override;
void Step() override;
ExitReason RunUntilExit() override;
ExitReason RunUntilExit(FEXCore::Core::InternalThreadState *Thread) override;
void ExecuteThread(FEXCore::Core::InternalThreadState *Thread) override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
bool IsDone() const override;
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
HostFeatures GetHostFeatures() const override;
@@ -117,7 +118,8 @@ namespace FEXCore::Context {
* - CTX->RunUntilExit(Thread);
* OS thread Creation:
* - Thread = CreateThread(0, 0, NewState, PPID);
* - InitializeThread(Thread);
* - Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
* - ThreadHandler calls `CTX->ExecutionThread(Thread)`
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(0, 0, CopyOfThreadState, PPID);
@@ -132,27 +134,13 @@ namespace FEXCore::Context {
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Initializes the OS thread object and prepares to start executing on that new OS thread
*
* @param Thread The internal FEX thread state object
*
* The OS thread will wait until RunThread is executed
*/
void InitializeThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Starts the OS thread object to start executing guest code
*
* @param Thread The internal FEX thread state object
*/
void RunThread(FEXCore::Core::InternalThreadState *Thread) override;
void StopThread(FEXCore::Core::InternalThreadState *Thread) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread) override;
void DestroyThread(FEXCore::Core::InternalThreadState *Thread, bool NeedsTLSUninstall) override;
#ifndef _WIN32
void LockBeforeFork(FEXCore::Core::InternalThreadState *Thread) override;
@@ -184,13 +172,19 @@ namespace FEXCore::Context {
void WriteFilesWithCode(AOTIRCodeFileWriterFn Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, CodeRangeInvalidationFn callback) override;
void MarkMemoryShared() override;
FEXCore::ForkableSharedMutex &GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
void MarkMemoryShared(FEXCore::Core::InternalThreadState *Thread) override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) override;
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void *Creator = nullptr, void *Data = nullptr) override;
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void *Creator = nullptr, void *Data = nullptr);
void AppendThunkDefinitions(fextl::vector<FEXCore::IR::ThunkDefinition> const& Definitions) override;
@@ -199,9 +193,6 @@ namespace FEXCore::Context {
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
#ifdef JIT_X86_64
friend class FEXCore::CPU::X86JITCore;
#endif
friend class FEXCore::IR::Validation::IRValidation;
@@ -232,7 +223,6 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(ThunkHostLibsPath32, THUNKHOSTLIBS32);
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
@@ -242,25 +232,16 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
} Config;
FEXCore::HostFeatures HostFeatures;
std::mutex ThreadCreationMutex;
FEXCore::Core::InternalThreadState* ParentThread{};
fextl::vector<FEXCore::Core::InternalThreadState*> Threads;
std::atomic_bool CoreShuttingDown{false};
bool NeedToCheckXID{true};
std::mutex IdleWaitMutex;
std::condition_variable IdleWaitCV;
std::atomic<uint32_t> IdleWaitRefCount{};
Event PauseWait;
bool Running{};
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::HostFeatures HostFeatures;
// CPUID depends on HostFeatures so needs to be initialized after that.
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
@@ -280,21 +261,15 @@ namespace FEXCore::Context {
ContextImpl();
~ContextImpl();
bool IsPaused() const { return !Running; }
void WaitForThreadsToRun() override;
void Stop(bool IgnoreCurrentThread);
void WaitForIdle() override;
void SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event);
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData *HostLink, const BlockDelinkerFunc &delinker);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, ExitFunctionLinkData *Record) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
return Fn(Frame, record);
return Fn(Frame, Record);
}
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
@@ -332,9 +307,6 @@ namespace FEXCore::Context {
[[nodiscard]] CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
// same as CompileBlock, but aborts on failure
void CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP);
// Used for thread creation from syscalls
/**
* @brief Initializes TID, PID and TLS data for a thread
@@ -359,10 +331,6 @@ namespace FEXCore::Context {
}
}
void IncrementIdleRefCount() override {
++IdleWaitRefCount;
}
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
@@ -392,18 +360,9 @@ namespace FEXCore::Context {
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
ThreadsState GetThreads() override {
return ThreadsState {
.ParentThread = ParentThread,
.Threads = &Threads,
};
}
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread);
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
// If the hardware supports TSO then we don't need to emulate it through atomics.
@@ -425,15 +384,8 @@ namespace FEXCore::Context {
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void WaitForIdleWithTimeout();
void NotifyPause();
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
// Entry Cache
std::mutex ExitMutex;
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
@@ -445,9 +397,7 @@ namespace FEXCore::Context {
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::atomic<bool> HasCustomIRHandlers{};
fextl::unordered_map<uint64_t, std::tuple<CustomIREntrypointHandler, void *, void *>> CustomIRHandlers;
FEXCore::CPU::DispatcherConfig DispatcherConfig;
};
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args);
}
@@ -9,10 +9,11 @@
#include "Interface/HLE/Thunks/Thunks.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <aarch64/cpu-aarch64.h>
#include <aarch64/instructions-aarch64.h>
#include <cpu-features.h>
@@ -28,6 +29,7 @@ namespace FEXCore::CPU {
// TODO: Allow x18 register allocation on Linux in the future to gain one more register.
namespace x64 {
#ifndef _M_ARM_64EC
// All but x19 and x29 are caller saved
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
@@ -81,6 +83,54 @@ namespace x64 {
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
};
#else
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r0,
FEXCore::ARMEmitter::Reg::r1, FEXCore::ARMEmitter::Reg::r27,
// SP's register location isn't specified by the ARM64EC ABI, we choose to use r23
FEXCore::ARMEmitter::Reg::r23, FEXCore::ARMEmitter::Reg::r29,
FEXCore::ARMEmitter::Reg::r25, FEXCore::ARMEmitter::Reg::r26,
FEXCore::ARMEmitter::Reg::r2, FEXCore::ARMEmitter::Reg::r3,
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r19, FEXCore::ARMEmitter::Reg::r20,
FEXCore::ARMEmitter::Reg::r21, FEXCore::ARMEmitter::Reg::r22,
REG_PF, REG_AF,
};
constexpr std::array<FEXCore::ARMEmitter::Register, 7> RA = {
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r14,FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r30,
};
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 3> RAPair = {{
{FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7},
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17},
}};
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
};
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21,
FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31
};
#endif
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
@@ -340,7 +390,7 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr
: Emitter(static_cast<uint8_t*>(EmissionPtr), size)
, EmitterCTX {ctx}
#ifdef VIXL_SIMULATOR
, Simulator {&SimDecoder}
, Simulator {&SimDecoder, stdout, vixl::aarch64::SimStack(SimulatorStackSize).Allocate()}
#endif
{
#ifdef VIXL_SIMULATOR
@@ -353,8 +403,10 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr
// Only setup the disassembler if enabled.
// vixl's decoder is expensive to setup.
if (Disassemble()) {
DisasmBuffer.resize(DISASM_BUFFER_SIZE);
Disasm = fextl::make_unique<vixl::aarch64::Disassembler>(DisasmBuffer.data(), DISASM_BUFFER_SIZE);
DisasmDecoder = fextl::make_unique<vixl::aarch64::Decoder>();
DisasmDecoder->AppendVisitor(&Disasm);
DisasmDecoder->AppendVisitor(Disasm.get());
}
#endif
@@ -367,6 +419,9 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr
GeneralPairRegisters = x64::RAPair;
StaticFPRegisters = x64::SRAFPR;
GeneralFPRegisters = x64::RAFPR;
#ifdef _M_ARM_64EC
ConfiguredDynamicRegisterBase = std::span(x64::RA.begin(), 7);
#endif
}
else {
ConfiguredDynamicRegisterBase = std::span(x32::RA.begin() + 6, 8);
@@ -611,10 +666,6 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
mrs(TmpReg, ARMEmitter::SystemRegister::NZCV);
str(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
if (!StaticRegisterAllocation()) {
return;
}
// PF/AF are special, remove them from the mask
uint32_t PFAFMask = ((1u << REG_PF.Idx()) | ((1u << REG_AF.Idx())));
unsigned PFAFSpillMask = GPRSpillMask & PFAFMask;
@@ -728,10 +779,6 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
if (!StaticRegisterAllocation()) {
return;
}
if (FPRs) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
@@ -936,7 +983,9 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
// Push the general registers.
PushGeneralRegisters(TmpReg, ConfiguredDynamicRegisterBase);
#ifndef _M_ARM_64EC
str(ARMEmitter::XReg::lr, TmpReg, 0);
#endif
}
void Arm64Emitter::PopDynamicRegsAndLR() {
@@ -948,7 +997,9 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
// Pop GPRs second
PopGeneralRegisters(ConfiguredDynamicRegisterBase);
#ifndef _M_ARM_64EC
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
#endif
}
void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs) {
@@ -21,6 +21,7 @@
#endif
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <cstddef>
@@ -36,16 +37,37 @@ namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
#ifndef _M_ARM_64EC
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x0;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x1;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x2;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
constexpr bool TMP_ABIARGS = true;
// We pin r26/r27 as PF/AF respectively, this is internal FEX ABI.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r26;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r27;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
#else
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x10;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x11;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x12;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x13;
constexpr bool TMP_ABIARGS = false;
// We pin r11/r12 as PF/AF respectively for arm64ec, as r26/r27 are used for SRA.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r9;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r24;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v16;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v17;
#endif
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
@@ -53,9 +75,6 @@ constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_16B = FEXCore::ARMEmitter::PReg::p6;
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PReg::p7;
// We pin r26/r27 as PF/AF respectively, this is internal FEX ABI.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r26;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r27;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
@@ -230,15 +249,17 @@ protected:
#ifdef VIXL_SIMULATOR
vixl::aarch64::Decoder SimDecoder;
vixl::aarch64::Simulator Simulator;
constexpr static size_t SimulatorStackSize = 8 * 1024 * 1024;
#endif
#ifdef VIXL_DISASSEMBLER
vixl::aarch64::Disassembler Disasm;
fextl::vector<char> DisasmBuffer;
constexpr static int DISASM_BUFFER_SIZE {256};
fextl::unique_ptr<vixl::aarch64::Disassembler> Disasm;
fextl::unique_ptr<vixl::aarch64::Decoder> DisasmDecoder;
FEX_CONFIG_OPT(Disassemble, DISASSEMBLE);
#endif
FEX_CONFIG_OPT(StaticRegisterAllocation, SRA);
};
}
@@ -35,8 +35,10 @@ public:
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adr(FEXCore::ARMEmitter::Register rd, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::ADR });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void adr(FEXCore::ARMEmitter::Register rd, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::ADR });
constexpr uint32_t Op = 0b0001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
@@ -62,8 +64,10 @@ public:
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, Imm);
}
void adrp(FEXCore::ARMEmitter::Register rd, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::ADRP });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void adrp(FEXCore::ARMEmitter::Register rd, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::ADRP });
constexpr uint32_t Op = 0b1001'0000 << 24;
DataProcessing_PCRel_Imm(Op, rd, 0);
}
@@ -105,7 +109,7 @@ public:
}
}
void LongAddressGen(FEXCore::ARMEmitter::Register rd, ForwardLabel* Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::LONG_ADDRESS_GEN });
Label->Insts.emplace_back(SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN });
// Emit a register index and a nop. These will be backpatched.
dc32(rd.Idx());
nop();
@@ -18,8 +18,10 @@ public:
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm >> 2);
}
void b(FEXCore::ARMEmitter::Condition Cond, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
@@ -45,8 +47,10 @@ public:
Branch_Conditional(Op, 0, 1, Cond, Imm >> 2);
}
void bc(FEXCore::ARMEmitter::Condition Cond, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bc(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
}
@@ -102,8 +106,10 @@ public:
UnconditionalBranch(Op, Imm >> 2);
}
void b(ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::B });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::B });
constexpr uint32_t Op = 0b0001'01 << 26;
UnconditionalBranch(Op, 0);
@@ -131,8 +137,10 @@ public:
UnconditionalBranch(Op, Imm >> 2);
}
void bl(ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::B });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bl(LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::B });
constexpr uint32_t Op = 0b1001'01 << 26;
UnconditionalBranch(Op, 0);
@@ -163,8 +171,10 @@ public:
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0100 << 24;
@@ -195,8 +205,10 @@ public:
CompareAndBranch(Op, s, rt, Imm >> 2);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::BC });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0101 << 24;
@@ -226,8 +238,11 @@ public:
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::TEST_BRANCH });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0110 << 24;
@@ -256,8 +271,11 @@ public:
TestAndBranch(Op, rt, Bit, Imm >> 2);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::TEST_BRANCH });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
@@ -4,13 +4,14 @@
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <aarch64/assembler-aarch64.h>
#include <array>
@@ -537,27 +538,29 @@ namespace FEXCore::ARMEmitter {
uint8_t *Location{};
};
/* This `ForwardLabel` struct used for retaining a location for PC-Relative instructions.
/* This `SingleUseForwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `above` an instruction that uses it.
* Which means that a branch would jump forwards.
*
* This can be bound to multiple instructions, so it needs a vector for each bind instruction type.
* The `ForwardLabel` struct can be bound to multiple instructions, so it needs a vector for each bind instruction type.
*/
struct ForwardLabel {
struct Instructions {
enum class InstType {
ADR,
ADRP,
B,
BC,
TEST_BRANCH,
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
uint8_t *Location{};
InstType Type;
struct SingleUseForwardLabel {
enum class InstType {
UNKNOWN,
ADR,
ADRP,
B,
BC,
TEST_BRANCH,
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
fextl::vector<Instructions> Insts{};
uint8_t *Location{};
InstType Type = InstType::UNKNOWN;
};
struct ForwardLabel {
fextl::vector<SingleUseForwardLabel> Insts{};
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
@@ -569,6 +572,15 @@ namespace FEXCore::ARMEmitter {
ForwardLabel Forward;
};
static inline void AddLocationToLabel(SingleUseForwardLabel *Label, SingleUseForwardLabel&& Location) {
LOGMAN_THROW_A_FMT(Label->Type == SingleUseForwardLabel::InstType::UNKNOWN, "Trying to bind a SingleUseForwardLabel to multiple locations. Use ForwardLabel instead.");
*Label = std::move(Location);
}
static inline void AddLocationToLabel(ForwardLabel *Label, SingleUseForwardLabel&& Location) {
Label->Insts.emplace_back(std::move(Location));
}
// Some FCMA ASIMD instructions support a rotation argument.
enum class Rotation : uint32_t {
ROTATE_0 = 0b00,
@@ -628,6 +640,121 @@ namespace FEXCore::ARMEmitter {
Label->Location = GetCursorAddress<uint8_t*>();
}
void Bind(const SingleUseForwardLabel *Label) {
uint8_t *CurrentAddress = GetCursorAddress<uint8_t*>();
// Patch up the instructions
switch (Label->Type) {
case SingleUseForwardLabel::InstType::ADR: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::ADRP: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::B: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= Offset;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::TEST_BRANCH: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::BC:
case SingleUseForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t *Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
}
else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
}
else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
}
else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
break;
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
}
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
@@ -636,119 +763,8 @@ namespace FEXCore::ARMEmitter {
if constexpr (WarnAboutEmpty) {
LOGMAN_THROW_A_FMT(Label->Insts.empty() == false, "Binding forward label that didn't have any instructions using it");
}
uint8_t *CurrentAddress = GetCursorAddress<uint8_t*>();
for (const auto &Inst : Label->Insts) {
// Patch up the instructions
switch (Inst.Type) {
case ForwardLabel::Instructions::InstType::ADR: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::ADRP: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::B: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= Offset;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::TEST_BRANCH: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::BC:
case ForwardLabel::Instructions::InstType::RELATIVE_LOAD: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case ForwardLabel::Instructions::InstType::LONG_ADDRESS_GEN: {
uint32_t *Instructions = reinterpret_cast<uint32_t*>(Inst.Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
}
else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
}
else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
}
else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
break;
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
for (auto &Inst : Label->Insts) {
Bind(&Inst);
}
}
@@ -2121,38 +2121,58 @@ public:
LoadStoreLiteral(Op, prfop, static_cast<uint32_t>(Imm >> 2) & 0x7'FFFF);
}
void ldr(FEXCore::ARMEmitter::WRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldr(FEXCore::ARMEmitter::WRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b0001'1000 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void ldr(FEXCore::ARMEmitter::SRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldr(FEXCore::ARMEmitter::SRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b0001'1100 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void ldr(FEXCore::ARMEmitter::XRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldr(FEXCore::ARMEmitter::XRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b0101'1000 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void ldr(FEXCore::ARMEmitter::DRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldr(FEXCore::ARMEmitter::DRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b0101'1100 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void ldrsw(FEXCore::ARMEmitter::XRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldrsw(FEXCore::ARMEmitter::XRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b1001'1000 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void ldr(FEXCore::ARMEmitter::QRegister rt, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void ldr(FEXCore::ARMEmitter::QRegister rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b1001'1100 << 24;
LoadStoreLiteral(Op, rt, 0);
}
void prfm(FEXCore::ARMEmitter::Prefetch prfop, ForwardLabel *Label) {
Label->Insts.emplace_back(ForwardLabel::Instructions{ .Location = GetCursorAddress<uint8_t*>(), .Type = ForwardLabel::Instructions::InstType::RELATIVE_LOAD });
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void prfm(FEXCore::ARMEmitter::Prefetch prfop, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::RELATIVE_LOAD });
constexpr uint32_t Op = 0b1101'1000 << 24;
LoadStoreLiteral(Op, prfop, 0);
}
@@ -5,6 +5,10 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
#ifndef _WIN32
#include <sys/prctl.h>
#endif
namespace FEXCore {
namespace CPU {
@@ -349,6 +353,31 @@ auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
//
// MDWE prevents applications from creating RWX memory mappings.
// This prevents FEX from doing anything JIT related, as FEX uses RWX for JIT memory mappings.
//
// A potential workaround to make FEX work with MDWE is to call mprotect every time we need to write or modify code.
// Alternatively, FEX could use a memory mirror where one half is mapped as RW and the other is RX.
//
// Once MDWE is enabled with the prctl, the feature is sealed and it can /NOT/ be turned off.
//
// Status of MDWE is queried through prctl using `PR_GET_MDWE`:
// -1: The kernel doesn't support MDWE
// 0: MDWE is supported but disabled
// >0: MDWE is enabled, hence prohibiting RWX mappings
#ifndef PR_GET_MDWE
#define PR_GET_MDWE 66
#endif
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
+35 -10
View File
@@ -98,7 +98,7 @@ constexpr uint32_t FAMILY_IDENTIFIER =
#endif
#ifdef _M_ARM_64
static uint32_t GetCycleCounterFrequency() {
uint32_t GetCycleCounterFrequency() {
uint64_t Result{};
__asm("mrs %[Res], CNTFRQ_EL0"
: [Res] "=r" (Result));
@@ -349,7 +349,7 @@ void CPUIDEmu::SetupHostHybridFlag() {
}
#else
static uint32_t GetCycleCounterFrequency() {
uint32_t GetCycleCounterFrequency() {
return 0;
}
@@ -358,6 +358,34 @@ void CPUIDEmu::SetupHostHybridFlag() {
#endif
void CPUIDEmu::SetupFeatures() {
// TODO: Enable once AVX is supported.
if (false && CTX->HostFeatures.SupportsAVX) {
XCR0 |= XCR0_AVX;
}
// Override features if the user has specifically called for it.
FEX_CONFIG_OPT(CPUIDFeatures, CPUID);
if (!CPUIDFeatures()) {
// Early exit if no features are overriden.
return;
}
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::DISABLE##enum_name) != 0; \
const bool Enable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features.FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features.FeatureName = Result; \
} while (0)
ENABLE_DISABLE_OPTION(SHA, SHA, SHA);
#undef ENABLE_DISABLE_OPTION
}
FEXCore::CPUID::FunctionResults CPUIDEmu::Function_0h(uint32_t Leaf) const {
FEXCore::CPUID::FunctionResults Res{};
@@ -639,7 +667,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 26) | // Reserved
(0 << 27) | // Reserved
(0 << 28) | // Reserved
(1 << 29) | // SHA instructions
(Features.SHA << 29) | // SHA instructions
(0 << 30) | // Reserved
(0 << 31); // Reserved
@@ -775,7 +803,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_15h(uint32_t Leaf) const {
uint32_t FrequencyHz = GetCycleCounterFrequency();
if (FrequencyHz) {
Res.eax = 1;
Res.ebx = 1;
Res.ebx = CTX->Config.SmallTSCScale() ? FEXCore::Context::TSC_SCALE : 1;
Res.ecx = FrequencyHz;
}
return Res;
@@ -1205,17 +1233,14 @@ FEXCore::CPUID::XCRResults CPUIDEmu::XCRFunction_0h() const {
return Res;
}
void CPUIDEmu::Init(FEXCore::Context::ContextImpl *ctx) {
CTX = ctx;
CPUIDEmu::CPUIDEmu(FEXCore::Context::ContextImpl const *ctx)
: CTX {ctx} {
Cores = FEXCore::CPUInfo::CalculateNumberOfCPUs();
// Setup some state tracking
SetupHostHybridFlag();
// TODO: Enable once AVX is supported.
if (false && CTX->HostFeatures.SupportsAVX) {
XCR0 |= XCR0_AVX;
}
SetupFeatures();
}
}
+16 -3
View File
@@ -14,6 +14,8 @@ namespace Context {
class ContextImpl;
}
uint32_t GetCycleCounterFrequency();
// Debugging define to switch what family of CPU we execute as.
// Might be useful if an application makes an assumption about a CPU.
// #define CPUID_AMD
@@ -28,12 +30,12 @@ private:
constexpr static uint32_t CPUID_VENDOR_AMD3 = 0x444D4163; // "cAMD"
public:
CPUIDEmu(FEXCore::Context::ContextImpl const *ctx);
// X86 cacheline size effectively has to be hardcoded to 64
// if we report anything differently then applications are likely to break
constexpr static uint64_t CACHELINE_SIZE = 64;
void Init(FEXCore::Context::ContextImpl *ctx);
FEXCore::CPUID::FunctionResults RunFunction(uint32_t Function, uint32_t Leaf) const {
if (Function < Primary.size()) {
const auto Handler = Primary[Function];
@@ -111,10 +113,11 @@ public:
}
private:
FEXCore::Context::ContextImpl *CTX;
FEXCore::Context::ContextImpl const *CTX;
bool Hybrid{};
uint32_t Cores{};
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
// XFEATURE_ENABLED_MASK
// Mask that configures what features are enabled on the CPU.
@@ -136,6 +139,15 @@ private:
constexpr static uint64_t XCR0_SSE = 1ULL << 1;
constexpr static uint64_t XCR0_AVX = 1ULL << 2;
struct FeaturesConfig {
uint64_t SHA : 1;
uint64_t _pad : 63;
};
FeaturesConfig Features {
.SHA = 1,
};
uint64_t XCR0 {
XCR0_X87 |
XCR0_SSE
@@ -189,6 +201,7 @@ private:
FEXCore::CPUID::XCRResults XCRFunction_0h() const;
void SetupHostHybridFlag();
void SetupFeatures();
static constexpr size_t PRIMARY_FUNCTION_COUNT = 27;
static constexpr size_t HYPERVISOR_FUNCTION_COUNT = 2;
static constexpr size_t EXTENDED_FUNCTION_COUNT = 32;
+52 -363
View File
@@ -20,6 +20,8 @@ $end_info$
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/HLE/Thunks/Thunks.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
@@ -38,7 +40,6 @@ $end_info$
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Allocator.h>
@@ -78,7 +79,8 @@ $end_info$
namespace FEXCore::Context {
ContextImpl::ContextImpl()
: IRCaptureCache {this} {
: CPUID {this}
, IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
#endif
@@ -97,32 +99,19 @@ namespace FEXCore::Context {
Symbols.InitFile();
}
if (FEXCore::GetCycleCounterFrequency() >= FEXCore::Context::TSC_SCALE_MAXIMUM) {
Config.SmallTSCScale = false;
}
// Track atomic TSO emulation configuration.
UpdateAtomicTSOEmulationConfig();
CPUID.Init(this);
}
ContextImpl::~ContextImpl() {
if (ParentThread) {
DestroyThread(ParentThread);
}
{
if (CodeObjectCacheService) {
CodeObjectCacheService->Shutdown();
}
for (auto &Thread : Threads) {
if (Thread->ExecutionThread->joinable()) {
Thread->ExecutionThread->join(nullptr);
}
}
for (auto &Thread : Threads) {
delete Thread;
}
Threads.clear();
}
}
@@ -269,7 +258,7 @@ namespace FEXCore::Context {
Frame->State.flags[X86State::RFLAG_IF_LOC] = 1;
}
FEXCore::Core::InternalThreadState* ContextImpl::InitCore(uint64_t InitialRIP, uint64_t StackPointer) {
bool ContextImpl::InitCore() {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
case FEXCore::Config::CONFIG_IRJIT:
@@ -279,21 +268,20 @@ namespace FEXCore::Context {
// Do nothing
break;
default:
ERROR_AND_DIE_FMT("Unknown core configuration");
break;
LogMan::Msg::EFmt("Unknown core configuration");
return false;
}
DispatcherConfig.StaticRegisterAllocation = Config.StaticRegisterAllocation && BackendFeatures.SupportsStaticRegisterAllocation;
Dispatcher = FEXCore::CPU::Dispatcher::Create(this, DispatcherConfig);
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
// Set up the SignalDelegator config since core is initialized.
FEXCore::SignalDelegator::SignalDelegatorConfig SignalConfig {
.StaticRegisterAllocation = DispatcherConfig.StaticRegisterAllocation,
.SupportsAVX = HostFeatures.SupportsAVX,
.DispatcherBegin = Dispatcher->Start,
.DispatcherEnd = Dispatcher->End,
.AbsoluteLoopTopAddress = Dispatcher->AbsoluteLoopTopAddress,
.AbsoluteLoopTopAddressFillSRA = Dispatcher->AbsoluteLoopTopAddressFillSRA,
.SignalHandlerReturnAddress = Dispatcher->SignalHandlerReturnAddress,
.SignalHandlerReturnAddressRT = Dispatcher->SignalHandlerReturnAddressRT,
@@ -331,174 +319,17 @@ namespace FEXCore::Context {
StartPaused = true;
}
FEXCore::Core::InternalThreadState *Thread = CreateThread(InitialRIP, StackPointer, nullptr, 0);
// We are the parent thread
ParentThread = Thread;
return Thread;
return true;
}
void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) {
static_cast<ContextImpl*>(Thread->CTX)->Dispatcher->ExecuteJITCallback(Thread->CurrentFrame, RIP);
}
void ContextImpl::WaitForIdle() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
IdleWaitCV.wait(lk, [this] {
return IdleWaitRefCount.load() == 0;
});
Running = false;
}
void ContextImpl::WaitForIdleWithTimeout() {
std::unique_lock<std::mutex> lk(IdleWaitMutex);
bool WaitResult = IdleWaitCV.wait_for(lk, std::chrono::milliseconds(1500),
[this] {
return IdleWaitRefCount.load() == 0;
});
if (!WaitResult) {
// The wait failed, this will occur if we stepped in to a syscall
// That's okay, we just need to pause the threads manually
NotifyPause();
}
// We have sent every thread a pause signal
// Now wait again because they /will/ be going to sleep
WaitForIdle();
}
void ContextImpl::NotifyPause() {
// Tell all the threads that they should pause
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
SignalDelegation->SignalThread(Thread, FEXCore::Core::SignalEvent::Pause);
}
}
void ContextImpl::Pause() {
// If we aren't running, WaitForIdle will never compete.
if (Running) {
NotifyPause();
WaitForIdle();
}
}
void ContextImpl::Run() {
// Spin up all the threads
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->SignalReason.store(FEXCore::Core::SignalEvent::Return);
}
for (auto &Thread : Threads) {
Thread->StartRunning.NotifyAll();
}
}
void ContextImpl::WaitForThreadsToRun() {
size_t NumThreads{};
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
NumThreads = Threads.size();
}
// Spin while waiting for the threads to start up
std::unique_lock<std::mutex> lk(IdleWaitMutex);
IdleWaitCV.wait(lk, [this, NumThreads] {
return IdleWaitRefCount.load() >= NumThreads;
});
Running = true;
}
void ContextImpl::Step() {
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
// Walk the threads and tell them to clear their caches
// Useful when our block size is set to a large number and we need to step a single instruction
for (auto &Thread : Threads) {
ClearCodeCache(Thread);
}
}
CoreRunningMode PreviousRunningMode = this->Config.RunningMode;
int64_t PreviousMaxIntPerBlock = this->Config.MaxInstPerBlock;
this->Config.RunningMode = FEXCore::Context::CoreRunningMode::MODE_SINGLESTEP;
this->Config.MaxInstPerBlock = 1;
Run();
WaitForThreadsToRun();
WaitForIdle();
this->Config.RunningMode = PreviousRunningMode;
this->Config.MaxInstPerBlock = PreviousMaxIntPerBlock;
}
void ContextImpl::Stop(bool IgnoreCurrentThread) {
pid_t tid = FHU::Syscalls::gettid();
FEXCore::Core::InternalThreadState* CurrentThread{};
// Tell all the threads that they should stop
{
std::lock_guard<std::mutex> lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
if (IgnoreCurrentThread &&
Thread->ThreadManager.TID == tid) {
// If we are callign stop from the current thread then we can ignore sending signals to this thread
// This means that this thread is already gone
continue;
}
else if (Thread->ThreadManager.TID == tid) {
// We need to save the current thread for last to ensure all threads receive their stop signals
CurrentThread = Thread;
continue;
}
if (Thread->RunningEvents.Running.load()) {
StopThread(Thread);
}
// If the thread is waiting to start but immediately killed then there can be a hang
// This occurs in the case of gdb attach with immediate kill
if (Thread->RunningEvents.WaitingToStart.load()) {
Thread->RunningEvents.EarlyExit = true;
Thread->StartRunning.NotifyAll();
}
}
}
// Stop the current thread now if we aren't ignoring it
if (CurrentThread) {
StopThread(CurrentThread);
}
}
void ContextImpl::StopThread(FEXCore::Core::InternalThreadState *Thread) {
if (Thread->RunningEvents.Running.exchange(false)) {
SignalDelegation->SignalThread(Thread, FEXCore::Core::SignalEvent::Stop);
}
}
void ContextImpl::SignalThread(FEXCore::Core::InternalThreadState *Thread, FEXCore::Core::SignalEvent Event) {
if (Thread->RunningEvents.Running.load()) {
SignalDelegation->SignalThread(Thread, Event);
}
}
FEXCore::Context::ExitReason ContextImpl::RunUntilExit() {
if(!StartPaused) {
// We will only have one thread at this point, but just in case run notify everything
std::lock_guard lk(ThreadCreationMutex);
for (auto &Thread : Threads) {
Thread->StartRunning.NotifyAll();
}
}
ExecutionThread(ParentThread);
FEXCore::Context::ExitReason ContextImpl::RunUntilExit(FEXCore::Core::InternalThreadState *Thread) {
ExecutionThread(Thread);
while(true) {
this->WaitForIdle();
auto reason = ParentThread->ExitReason;
auto reason = Thread->ExitReason;
// Don't return if a custom exit handling the exit
if (!CustomExitHandler || reason == ExitReason::EXIT_SHUTDOWN) {
@@ -511,61 +342,18 @@ namespace FEXCore::Context {
Dispatcher->ExecuteDispatch(Thread->CurrentFrame);
}
struct ExecutionThreadHandler {
ContextImpl *This;
FEXCore::Core::InternalThreadState *Thread;
};
static void *ThreadHandler(void* Data) {
ExecutionThreadHandler *Handler = reinterpret_cast<ExecutionThreadHandler*>(Data);
Handler->This->ExecutionThread(Handler->Thread);
FEXCore::Allocator::free(Handler);
return nullptr;
}
void ContextImpl::InitializeThread(FEXCore::Core::InternalThreadState *Thread) {
// This will create the execution thread but it won't actually start executing
ExecutionThreadHandler *Arg = reinterpret_cast<ExecutionThreadHandler*>(FEXCore::Allocator::malloc(sizeof(ExecutionThreadHandler)));
Arg->This = this;
Arg->Thread = Thread;
Thread->StartPaused = NeedToCheckXID;
Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
// Wait for the thread to have started
Thread->ThreadWaiting.Wait();
if (NeedToCheckXID) {
// The first time an application creates a thread, GLIBC installs their SETXID signal handler.
// FEX needs to capture all signals and defer them to the guest.
// Once FEX creates its first guest thread, overwrite the GLIBC SETXID handler *again* to ensure
// FEX maintains control of the signal handler on this signal.
NeedToCheckXID = false;
SignalDelegation->CheckXIDHandler();
Thread->StartRunning.NotifyAll();
}
}
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
if (Config.BlockJITNaming() ||
Config.GlobalJITNaming() ||
Config.LibraryJITNaming()) {
// Allocate a TLS JIT symbol buffer only if enabled.
Thread->SymbolBuffer = JITSymbols::AllocateBuffer();
}
SignalDelegation->RegisterTLSState(Thread);
if (ThunkHandler) {
ThunkHandler->RegisterTLSState(Thread);
}
}
void ContextImpl::RunThread(FEXCore::Core::InternalThreadState *Thread) {
// Tell the thread to start executing
Thread->StartRunning.NotifyAll();
#ifndef _WIN32
Alloc::OSAllocator::RegisterTLSData(Thread);
#endif
}
void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread) {
@@ -574,9 +362,6 @@ namespace FEXCore::Context {
Thread->LookupCache = fextl::make_unique<FEXCore::LookupCache>(this);
Thread->FrontendDecoder = fextl::make_unique<FEXCore::Frontend::Decoder>(this);
Thread->PassManager = fextl::make_unique<FEXCore::IR::PassManager>();
Thread->PassManager->RegisterExitHandler([this]() {
Stop(false /* Ignore current thread */);
});
Thread->CurrentFrame->Pointers.Common.L1Pointer = Thread->LookupCache->GetL1Pointer();
Thread->CurrentFrame->Pointers.Common.L2Pointer = Thread->LookupCache->GetPagePointer();
@@ -585,9 +370,7 @@ namespace FEXCore::Context {
Thread->CTX = this;
bool DoSRA = DispatcherConfig.StaticRegisterAllocation;
Thread->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT, DoSRA);
Thread->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT);
Thread->PassManager->AddDefaultValidationPasses();
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
@@ -595,7 +378,7 @@ namespace FEXCore::Context {
// Create CPU backend
switch (Config.Core) {
case FEXCore::Config::CONFIG_IRJIT:
Thread->PassManager->InsertRegisterAllocationPass(DoSRA, HostFeatures.SupportsAVX);
Thread->PassManager->InsertRegisterAllocationPass(HostFeatures.SupportsAVX);
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
break;
case FEXCore::Config::CONFIG_CUSTOM:
@@ -629,30 +412,21 @@ namespace FEXCore::Context {
Thread->CurrentFrame->State.DeferredSignalRefCount.Store(0);
Thread->CurrentFrame->State.DeferredSignalFaultAddress = reinterpret_cast<Core::NonAtomicRefCounter<uint64_t>*>(FEXCore::Allocator::VirtualAlloc(4096));
// Insert after the Thread object has been fully initialized
{
std::lock_guard lk(ThreadCreationMutex);
Threads.push_back(Thread);
if (Config.BlockJITNaming() ||
Config.GlobalJITNaming() ||
Config.LibraryJITNaming()) {
// Allocate a JIT symbol buffer only if enabled.
Thread->SymbolBuffer = JITSymbols::AllocateBuffer();
}
return Thread;
}
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState *Thread) {
// remove new thread object
{
std::lock_guard lk(ThreadCreationMutex);
auto It = std::find(Threads.begin(), Threads.end(), Thread);
LOGMAN_THROW_A_FMT(It != Threads.end(), "Thread wasn't in Threads");
Threads.erase(It);
}
if (Thread->ExecutionThread &&
Thread->ExecutionThread->IsSelf()) {
// To be able to delete a thread from itself, we need to detached the std::thread object
Thread->ExecutionThread->detach();
void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState *Thread, bool NeedsTLSUninstall) {
if (NeedsTLSUninstall) {
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
}
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096);
@@ -670,43 +444,6 @@ namespace FEXCore::Context {
CodeInvalidationMutex.unlock();
return;
}
// This function is called after fork
// We need to cleanup some of the thread data that is dead
for (auto &DeadThread : Threads) {
if (DeadThread == LiveThread) {
continue;
}
// Setting running to false ensures that when they are shutdown we won't send signals to kill them
DeadThread->RunningEvents.Running = false;
// Despite what google searches may susgest, glibc actually has special code to handle forks
// with multiple active threads.
// It cleans up the stacks of dead threads and marks them as terminated.
// It also cleans up a bunch of internal mutexes.
// FIXME: TLS is probally still alive. Investigate
// Deconstructing the Interneal thread state should clean up most of the state.
// But if anything on the now deleted stack is holding a refrence to the heap, it will be leaked
delete DeadThread;
// FIXME: Make sure sure nothing gets leaked via the heap. Ideas:
// * Make sure nothing is allocated on the heap without ref in InternalThreadState
// * Surround any code that heap allocates with a per-thread mutex.
// Before forking, the the forking thread can lock all thread mutexes.
}
// Remove all threads but the live thread from Threads
Threads.clear();
Threads.push_back(LiveThread);
// We now only have one thread
IdleWaitRefCount = 1;
// Clean up dead stacks
FEXCore::Threads::Thread::CleanupAfterFork();
}
void ContextImpl::LockBeforeFork(FEXCore::Core::InternalThreadState *Thread) {
@@ -775,17 +512,20 @@ namespace FEXCore::Context {
uint64_t TotalInstructions {0};
uint64_t TotalInstructionsLength {0};
bool HasCustomIR{};
std::shared_lock lk(CustomIRMutex);
if (HasCustomIRHandlers.load(std::memory_order_relaxed)) {
std::shared_lock lk(CustomIRMutex);
auto Handler = CustomIRHandlers.find(GuestRIP);
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
TotalInstructionsLength = 1;
std::get<0>(Handler->second)(GuestRIP, Thread->OpDispatcher.get());
HasCustomIR = true;
}
}
auto Handler = CustomIRHandlers.find(GuestRIP);
if (Handler != CustomIRHandlers.end()) {
TotalInstructions = 1;
TotalInstructionsLength = 1;
std::get<0>(Handler->second)(GuestRIP, Thread->OpDispatcher.get());
lk.unlock();
} else {
lk.unlock();
if (!HasCustomIR) {
uint8_t const *GuestCode{};
GuestCode = reinterpret_cast<uint8_t const*>(GuestRIP);
@@ -1027,17 +767,6 @@ namespace FEXCore::Context {
};
}
void ContextImpl::CompileBlockJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto NewBlock = CompileBlock(Frame, GuestRIP);
if (NewBlock == 0) {
LogMan::Msg::EFmt("CompileBlockJit: Failed to compile code {:X} - aborting process", GuestRIP);
// Return similar behaviour of SIGILL abort
Frame->Thread->StatusCode = 128 + SIGILL;
Stop(false /* Ignore current thread */);
}
}
uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP, uint64_t MaxInst) {
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
@@ -1142,16 +871,11 @@ namespace FEXCore::Context {
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_WAITING;
InitializeThreadTLSData(Thread);
#ifndef _WIN32
Alloc::OSAllocator::RegisterTLSData(Thread);
#endif
++IdleWaitRefCount;
// Now notify the thread that we are initialized
Thread->ThreadWaiting.NotifyAll();
if (Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread || StartPaused || Thread->StartPaused) {
if (StartPaused || Thread->StartPaused) {
// Parent thread doesn't need to wait to run
Thread->StartRunning.Wait();
}
@@ -1186,18 +910,9 @@ namespace FEXCore::Context {
}
}
--IdleWaitRefCount;
IdleWaitCV.notify_all();
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
SignalDelegation->UninstallTLSState(Thread);
// If the parent thread is waiting to join, then we can't destroy our thread object
if (!Thread->DestroyedByParent && Thread != static_cast<ContextImpl*>(Thread->CTX)->ParentThread) {
Thread->CTX->DestroyThread(Thread);
}
}
static void InvalidateGuestThreadCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
@@ -1214,44 +929,21 @@ namespace FEXCore::Context {
}
}
static void InvalidateGuestCodeRangeInternal(ContextImpl *CTX, uint64_t Start, uint64_t Length) {
std::lock_guard lk(static_cast<ContextImpl*>(CTX)->ThreadCreationMutex);
for (auto &Thread : static_cast<ContextImpl*>(CTX)->Threads) {
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto lk = GuardSignalDeferringSectionWithFallback(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(this, Start, Length);
InvalidateGuestThreadCodeRange(Thread, Start, Length);
}
void ContextImpl::InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, CodeRangeInvalidationFn CallAfter) {
// Potential deferred since Thread might not be valid.
// Thread object isn't valid very early in frontend's initialization.
// To be more optimal the frontend should provide this code with a valid Thread object earlier.
auto lk = GuardSignalDeferringSectionWithFallback(CodeInvalidationMutex, Thread);
InvalidateGuestCodeRangeInternal(this, Start, Length);
InvalidateGuestThreadCodeRange(Thread, Start, Length);
CallAfter(Start, Length);
}
void ContextImpl::MarkMemoryShared() {
void ContextImpl::MarkMemoryShared(FEXCore::Core::InternalThreadState *Thread) {
if (!IsMemoryShared) {
IsMemoryShared = true;
UpdateAtomicTSOEmulationConfig();
if (Config.TSOAutoMigration) {
std::lock_guard<std::mutex> lkThreads(ThreadCreationMutex);
LogMan::Throw::AFmt(Threads.size() == 1, "First MarkMemoryShared called must be before creating any threads");
auto Thread = Threads[0];
// Only the lookup cache is cleared here, so that old code can keep running until next compilation
std::lock_guard<std::recursive_mutex> lkLookupCache(Thread->LookupCache->WriteLock);
Thread->LookupCache->ClearCache();
@@ -1262,7 +954,7 @@ namespace FEXCore::Context {
}
}
void ContextImpl::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
void ContextImpl::ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData *HostLink, const FEXCore::Context::BlockDelinkerFunc &delinker) {
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
Thread->LookupCache->AddBlockLink(GuestDestination, HostLink, delinker);
@@ -1274,7 +966,7 @@ namespace FEXCore::Context {
std::lock_guard<std::recursive_mutex> lk(Thread->LookupCache->WriteLock);
Thread->DebugStore.erase(GuestRIP);
Thread->LookupCache->Erase(GuestRIP);
Thread->LookupCache->Erase(Thread->CurrentFrame, GuestRIP);
}
CustomIRResult ContextImpl::AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void *Creator, void *Data) {
@@ -1283,6 +975,7 @@ namespace FEXCore::Context {
std::unique_lock lk(CustomIRMutex);
auto InsertedIterator = CustomIRHandlers.emplace(Entrypoint, std::tuple(Handler, Creator, Data));
HasCustomIRHandlers = true;
if (!InsertedIterator.second) {
const auto &[fn, Creator, Data] = InsertedIterator.first->second;
@@ -1301,12 +994,8 @@ namespace FEXCore::Context {
InvalidateGuestCodeRange(nullptr, Entrypoint, 1, [this](uint64_t Entrypoint, uint64_t) {
CustomIRHandlers.erase(Entrypoint);
});
}
uint64_t HandleSyscall(FEXCore::HLE::SyscallHandler *Handler, FEXCore::Core::CpuStateFrame *Frame, FEXCore::HLE::SyscallArguments *Args) {
uint64_t Result{};
Result = Handler->HandleSyscall(Frame, Args);
return Result;
HasCustomIRHandlers = !CustomIRHandlers.empty();
}
IR::AOTIRCacheEntry *ContextImpl::LoadAOTIRCacheEntry(const fextl::string &filename) {
@@ -5,12 +5,14 @@
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -23,42 +25,15 @@
namespace FEXCore::CPU {
void Dispatcher::SleepThread(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::CpuStateFrame *Frame) {
auto Thread = Frame->Thread;
--ctx->IdleWaitRefCount;
ctx->IdleWaitCV.notify_all();
Thread->RunningEvents.ThreadSleeping = true;
// Go to sleep
Thread->StartRunning.Wait();
Thread->RunningEvents.Running = true;
++ctx->IdleWaitRefCount;
Thread->RunningEvents.ThreadSleeping = false;
ctx->IdleWaitCV.notify_all();
}
uint64_t Dispatcher::GetCompileBlockPtr() {
using ClassPtrType = void (FEXCore::Context::ContextImpl::*)(FEXCore::Core::CpuStateFrame *, uint64_t);
union PtrCast {
ClassPtrType ClassPtr;
uintptr_t Data;
};
PtrCast CompileBlockPtr;
CompileBlockPtr.ClassPtr = &FEXCore::Context::ContextImpl::CompileBlockJit;
return CompileBlockPtr.Data;
static void SleepThread(FEXCore::Context::ContextImpl *CTX, FEXCore::Core::CpuStateFrame *Frame) {
CTX->SyscallHandler->SleepThread(CTX, Frame);
}
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
Dispatcher::Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &config)
Dispatcher::Dispatcher(FEXCore::Context::ContextImpl *ctx)
: Arm64Emitter(ctx, FEXCore::Allocator::VirtualAlloc(MAX_DISPATCHER_CODE_SIZE, true), MAX_DISPATCHER_CODE_SIZE)
, CTX {ctx}
, config {config} {
, CTX {ctx} {
EmitDispatcher();
}
@@ -85,8 +60,8 @@ void Dispatcher::EmitDispatcher() {
// }
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::ForwardLabel l_Sleep;
ARMEmitter::ForwardLabel l_CompileBlock;
ARMEmitter::SingleUseForwardLabel l_Sleep;
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
// Push all the register we need to save
PushCalleeSavedRegisters();
@@ -103,9 +78,7 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation) {
FillStaticRegs();
}
FillStaticRegs();
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
@@ -117,87 +90,86 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
// Don't modify x2 since it contains our RIP once the block doesn't exist
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
auto RipReg = ARMEmitter::XReg::x2;
auto RipReg = TMP3;
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
// L1 Cache
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, ARMEmitter::Reg::r0, ARMEmitter::Reg::r3, ARMEmitter::ShiftType::LSL , 4);
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, 0);
sub(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, RipReg);
cbnz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, &FullLookup);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL , 4);
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP1, TMP1, 0);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(ARMEmitter::Reg::r3);
br(TMP4);
// L1C check failed, do a full lookup
Bind(&FullLookup);
// This is the block cache lookup routine
// It matches what is going on it LookupCache.h::FindBlock
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L2Pointer));
// Mask the address by the virtual address size so we can check for aliases
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), VirtualMemorySize - 1);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), VirtualMemorySize - 1);
}
else {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg.R(), ARMEmitter::Reg::r3);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), TMP4);
}
ARMEmitter::ForwardLabel NoBlock;
{
// Offset the address and add to our page pointer
lsr(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::r3, 12);
lsr(ARMEmitter::Size::i64Bit, TMP2, TMP4, 12);
// Load the pointer from the offset
ldr(ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, ARMEmitter::Reg::r1, ARMEmitter::ExtendedType::LSL_64, 3);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 3);
// If page pointer is zero then we have no block
cbz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, &NoBlock);
cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, ARMEmitter::Reg::r3, 0x0FFF);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
// Shift the offset by the size of the block cache entry
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::ShiftType::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, (int)log2(sizeof(FEXCore::LookupCache::LookupCacheEntry)));
// The the full LookupCacheEntry with a single LDP.
// Check the guest address first to ensure it maps to the address we are currently at.
// This fixes aliasing problems
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x1, ARMEmitter::Reg::r0, 0);
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP2, TMP1, 0);
// If the guest address doesn't match, Compile the block.
sub(ARMEmitter::XReg::x1, ARMEmitter::XReg::x1, RipReg);
cbnz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, &NoBlock);
sub(TMP2, TMP2, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP2, &NoBlock);
// Check the host address to see if it matches, else compile the block.
cbz(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, &NoBlock);
cbz(ARMEmitter::Size::i64Bit, TMP4, &NoBlock);
// If we've made it here then we have a real compiled block
{
// update L1 cache
ldr(ARMEmitter::XReg::x0, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x3, ARMEmitter::XReg::x2, ARMEmitter::Reg::r0);
and_(ARMEmitter::Size::i64Bit, TMP2, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP3, TMP1);
// Jump to the block
br(ARMEmitter::Reg::r3);
br(TMP4);
}
}
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
ThreadStopHandlerAddress = GetCursorAddress<uint64_t>();
@@ -210,8 +182,7 @@ void Dispatcher::EmitDispatcher() {
{
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
@@ -228,26 +199,32 @@ void Dispatcher::EmitDispatcher() {
blr(ARMEmitter::Reg::r2);
}
if (config.StaticRegisterAllocation)
FillStaticRegs();
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x1, ARMEmitter::XReg::x1, 1);
str(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
FillStaticRegs();
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(ARMEmitter::XReg::x1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, ARMEmitter::XReg::x1, 0);
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP2, 0);
br(ARMEmitter::Reg::r0);
br(TMP1);
}
// Need to create the block
{
Bind(&NoBlock);
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, TMP3);
}
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
@@ -255,22 +232,22 @@ void Dispatcher::EmitDispatcher() {
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x3, &l_CompileBlock);
// x2 contains guest RIP
mov(ARMEmitter::XReg::x3, 0);
ldr(ARMEmitter::XReg::x4, &l_CompileBlock);
// X2 contains our guest RIP
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void *, uint64_t, void *>(ARMEmitter::Reg::r3);
GenerateIndirectRuntimeCall<uintptr_t, void *, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
}
else {
blr(ARMEmitter::Reg::r3); // { CTX, Frame, RIP}
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP, MaxInst }
}
if (config.StaticRegisterAllocation)
FillStaticRegs();
FillStaticRegs();
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
@@ -300,8 +277,7 @@ void Dispatcher::EmitDispatcher() {
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGILL = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
hlt(0);
}
@@ -311,8 +287,7 @@ void Dispatcher::EmitDispatcher() {
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
brk(0);
}
@@ -322,8 +297,7 @@ void Dispatcher::EmitDispatcher() {
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGSEGV = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
// hlt/udf = SIGILL
// brk = SIGTRAP
@@ -343,8 +317,7 @@ void Dispatcher::EmitDispatcher() {
{
ThreadPauseHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
if (config.StaticRegisterAllocation)
SpillStaticRegs(TMP1);
SpillStaticRegs(TMP1);
ThreadPauseHandlerAddress = GetCursorAddress<uint64_t>();
// We are pausing, this means the frontend should be waiting for this thread to idle
@@ -411,112 +384,59 @@ void Dispatcher::EmitDispatcher() {
str(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, State.rip));
// load static regs
if (config.StaticRegisterAllocation)
FillStaticRegs();
FillStaticRegs();
// Now go back to the regular dispatcher loop
b(&LoopTop);
}
{
LUDIVHandlerAddress = GetCursorAddress<uint64_t>();
auto EmitLongALUOpHandler = [&](auto R, auto Offset) {
auto Address = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs(ARMEmitter::Reg::r3);
PushDynamicRegsAndLR(TMP4);
SpillStaticRegs(TMP4);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x0, TMP1);
mov(ARMEmitter::XReg::x1, TMP2);
mov(ARMEmitter::XReg::x2, TMP3);
}
ldr(ARMEmitter::XReg::x3, R, Offset);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
// Result is now in x0
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
return Address;
};
{
LDIVHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LUREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
{
LREMHandlerAddress = GetCursorAddress<uint64_t>();
PushDynamicRegsAndLR(ARMEmitter::Reg::r3);
SpillStaticRegs(ARMEmitter::Reg::r3);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
FillStaticRegs();
// Result is now in x0
// Fix the stack and any values that were stepped on
PopDynamicRegsAndLR();
// Go back to our code block
ret();
}
LUDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LUDIV));
LDIVHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LDIV));
LUREMHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LUREM));
LREMHandlerAddress = EmitLongALUOpHandler(STATE_PTR(CpuStateFrame, Pointers.AArch64.LREM));
Bind(&l_CTX);
dc64(reinterpret_cast<uintptr_t>(CTX));
Bind(&l_Sleep);
dc64(reinterpret_cast<uint64_t>(SleepThread));
Bind(&l_CompileBlock);
dc64(GetCompileBlockPtr());
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::Context::ContextImpl::CompileBlock);
dc64(PMF.GetConvertedPointer());
Start = reinterpret_cast<uint64_t>(DispatchPtr);
End = GetCursorAddress<uint64_t>();
@@ -535,7 +455,7 @@ void Dispatcher::EmitDispatcher() {
const auto DisasmEnd = GetCursorAddress<const vixl::aarch64::Instruction*>();
for (auto PCToDecode = DisasmBegin; PCToDecode < DisasmEnd; PCToDecode += 4) {
DisasmDecoder->Decode(PCToDecode);
auto Output = Disasm.GetOutput();
auto Output = Disasm->GetOutput();
LogMan::Msg::IFmt("{}", Output);
}
}
@@ -580,8 +500,8 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread)
}
}
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config) {
return fextl::make_unique<Dispatcher>(CTX, Config);
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl *CTX) {
return fextl::make_unique<Dispatcher>(CTX);
}
}
@@ -31,18 +31,14 @@ class ContextImpl;
namespace FEXCore::CPU {
struct DispatcherConfig {
bool StaticRegisterAllocation = false;
};
#define STATE_PTR(STATE_TYPE, FIELD) \
STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
class Dispatcher final : public Arm64Emitter {
public:
static fextl::unique_ptr<Dispatcher> Create(FEXCore::Context::ContextImpl *CTX, const DispatcherConfig &Config);
static fextl::unique_ptr<Dispatcher> Create(FEXCore::Context::ContextImpl *CTX);
Dispatcher(FEXCore::Context::ContextImpl *ctx, const DispatcherConfig &Config);
Dispatcher(FEXCore::Context::ContextImpl *ctx);
~Dispatcher();
/**
@@ -106,15 +102,8 @@ public:
}
}
const DispatcherConfig& GetConfig() const { return config; }
protected:
FEXCore::Context::ContextImpl *CTX;
DispatcherConfig config;
static void SleepThread(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::CpuStateFrame *Frame);
static uint64_t GetCompileBlockPtr();
using AsmDispatch = void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
@@ -284,7 +284,6 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModR
if (DisplacementSize == 1) {
Literal = static_cast<int8_t>(Literal);
}
Displacement = DisplacementSize;
Operand->Type = DecodedOperand::OpType::GPRIndirect;
Operand->Data.GPRIndirect.GPR = MapModRMToReg(DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B ? 1 : 0, ModRM.rm, false, false, false, false);
+30 -115
View File
@@ -9,14 +9,6 @@
#ifdef _M_X86_64
#define XBYAK64
#define XBYAK_CUSTOM_ALLOC
#define XBYAK_CUSTOM_MALLOC FEXCore::Allocator::malloc
#define XBYAK_CUSTOM_FREE FEXCore::Allocator::free
#define XBYAK_CUSTOM_SETS
#define XBYAK_STD_UNORDERED_SET fextl::unordered_set
#define XBYAK_STD_UNORDERED_MAP fextl::unordered_map
#define XBYAK_STD_UNORDERED_MULTIMAP fextl::unordered_multimap
#define XBYAK_STD_LIST fextl::list
#define XBYAK_NO_EXCEPTION
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/unordered_map.h>
@@ -69,114 +61,42 @@ static void OverrideFeatures(HostFeatures *Features) {
return;
}
#define ENABLE_DISABLE_OPTION(name, enum_name) \
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features->FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features->FeatureName = Result; \
} while (0)
#define GET_SINGLE_OPTION(name, enum_name) \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive");
ENABLE_DISABLE_OPTION(AVX, AVX);
ENABLE_DISABLE_OPTION(AVX2, AVX2);
ENABLE_DISABLE_OPTION(SVE, SVE);
ENABLE_DISABLE_OPTION(AFP, AFP);
ENABLE_DISABLE_OPTION(LRCPC, LRCPC);
ENABLE_DISABLE_OPTION(LRCPC2, LRCPC2);
ENABLE_DISABLE_OPTION(CSSC, CSSC);
ENABLE_DISABLE_OPTION(PMULL128, PMULL128);
ENABLE_DISABLE_OPTION(RNG, RNG);
ENABLE_DISABLE_OPTION(CLZERO, CLZERO);
ENABLE_DISABLE_OPTION(Atomics, ATOMICS);
ENABLE_DISABLE_OPTION(FCMA, FCMA);
ENABLE_DISABLE_OPTION(FlagM, FLAGM);
ENABLE_DISABLE_OPTION(FlagM2, FLAGM2);
ENABLE_DISABLE_OPTION(Crypto, CRYPTO);
ENABLE_DISABLE_OPTION(RPRES, RPRES);
ENABLE_DISABLE_OPTION(SupportsAVX, AVX, AVX);
ENABLE_DISABLE_OPTION(SupportsAVX2, AVX2, AVX2);
ENABLE_DISABLE_OPTION(SupportsSVE, SVE, SVE);
ENABLE_DISABLE_OPTION(SupportsAFP, AFP, AFP);
ENABLE_DISABLE_OPTION(SupportsRCPC, LRCPC, LRCPC);
ENABLE_DISABLE_OPTION(SupportsTSOImm9, LRCPC2, LRCPC2);
ENABLE_DISABLE_OPTION(SupportsCSSC, CSSC, CSSC);
ENABLE_DISABLE_OPTION(SupportsPMULL_128Bit, PMULL128, PMULL128);
ENABLE_DISABLE_OPTION(SupportsRAND, RNG, RNG);
ENABLE_DISABLE_OPTION(SupportsCLZERO, CLZERO, CLZERO);
ENABLE_DISABLE_OPTION(SupportsAtomics, Atomics, ATOMICS);
ENABLE_DISABLE_OPTION(SupportsFCMA, FCMA, FCMA);
ENABLE_DISABLE_OPTION(SupportsFlagM, FlagM, FLAGM);
ENABLE_DISABLE_OPTION(SupportsFlagM2, FlagM2, FLAGM2);
ENABLE_DISABLE_OPTION(SupportsRPRES, RPRES, RPRES);
ENABLE_DISABLE_OPTION(SupportsPreserveAllABI, PRESERVEALLABI, PRESERVEALLABI);
GET_SINGLE_OPTION(Crypto, CRYPTO);
#undef ENABLE_DISABLE_OPTION
#undef GET_SINGLE_OPTION
if (EnableAVX) {
Features->SupportsAVX = true;
}
else if (DisableAVX) {
Features->SupportsAVX = false;
}
if (EnableAVX2) {
Features->SupportsAVX2 = true;
}
else if (DisableAVX2) {
Features->SupportsAVX2 = false;
}
if (EnableSVE) {
Features->SupportsSVE = true;
}
else if (DisableSVE) {
Features->SupportsSVE = false;
}
if (EnableAFP) {
Features->SupportsAFP = true;
}
else if (DisableAFP) {
Features->SupportsAFP = false;
}
if (EnableLRCPC) {
Features->SupportsRCPC = true;
}
else if (DisableLRCPC) {
Features->SupportsRCPC = false;
}
if (EnableLRCPC2) {
Features->SupportsTSOImm9 = true;
}
else if (DisableLRCPC2) {
Features->SupportsTSOImm9 = false;
}
if (EnableCSSC) {
Features->SupportsCSSC = true;
}
else if (DisableCSSC) {
Features->SupportsCSSC = false;
}
if (EnablePMULL128) {
Features->SupportsPMULL_128Bit = true;
}
else if (DisablePMULL128) {
Features->SupportsPMULL_128Bit = false;
}
if (EnableRNG) {
Features->SupportsRAND = true;
}
else if (DisableRNG) {
Features->SupportsRAND = false;
}
if (EnableCLZERO) {
Features->SupportsCLZERO = true;
}
else if (DisableCLZERO) {
Features->SupportsCLZERO = false;
}
if (EnableAtomics) {
Features->SupportsAtomics = true;
}
else if (DisableAtomics) {
Features->SupportsAtomics = false;
}
if (EnableFCMA) {
Features->SupportsFCMA = true;
}
else if (DisableFCMA) {
Features->SupportsFCMA = false;
}
if (EnableFlagM) {
Features->SupportsFlagM = true;
}
else if (DisableFlagM) {
Features->SupportsFlagM = false;
}
if (EnableFlagM2) {
Features->SupportsFlagM2 = true;
}
else if (DisableFlagM2) {
Features->SupportsFlagM2 = false;
}
if (EnableCrypto) {
Features->SupportsAES = true;
Features->SupportsCRC = true;
@@ -189,12 +109,6 @@ static void OverrideFeatures(HostFeatures *Features) {
Features->SupportsSHA = false;
Features->SupportsPMULL_128Bit = false;
}
if (EnableRPRES) {
Features->SupportsRPRES = true;
}
else if (DisableRPRES) {
Features->SupportsRPRES = false;
}
}
HostFeatures::HostFeatures() {
@@ -339,6 +253,7 @@ HostFeatures::HostFeatures() {
SupportsFloatExceptions = true;
#endif
#endif
SupportsPreserveAllABI = FEXCORE_HAS_PRESERVE_ALL_ATTR;
OverrideFeatures(this);
}
}
@@ -1,2 +0,0 @@
// SPDX-License-Identifier: MIT
#include <FEXCore/Debug/InternalThreadState.h>
@@ -85,7 +85,7 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_VPCMPISTRX] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle);
}
bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info) {
bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, IR::IROp_Header const *IROp, FallbackInfo *Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80CVTTO: {
@@ -93,11 +93,11 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
switch (Op->SrcSize) {
case 4: {
*Info = {FABI_F80_I16_F32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F80_I16_F32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F80_I16_F64, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F80_I16_F64, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -107,11 +107,11 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = {FABI_F32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -124,28 +124,28 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
switch (OpSize) {
case 2: {
if (Op->Truncate) {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2, SupportsPreserveAllABI};
}
return true;
}
case 4: {
if (Op->Truncate) {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4, SupportsPreserveAllABI};
}
return true;
}
case 8: {
if (Op->Truncate) {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8, SupportsPreserveAllABI};
}
return true;
}
@@ -167,7 +167,7 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags), FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags), SupportsPreserveAllABI};
return true;
}
@@ -176,11 +176,11 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
switch (Op->SrcSize) {
case 2: {
*Info = {FABI_F80_I16_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F80_I16_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2, SupportsPreserveAllABI};
return true;
}
case 4: {
*Info = {FABI_F80_I16_I32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_F80_I16_I32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
@@ -190,13 +190,13 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, FEXCORE_HAS_PRESERVE_ALL_ATTR}; \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, FEXCORE_HAS_PRESERVE_ALL_ATTR}; \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
@@ -244,10 +244,10 @@ bool InterpreterOps::GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInf
// SSE4.2 Fallbacks
case IR::OP_VPCMPESTRX:
*Info = {FABI_I32_I64_I64_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I32_I64_I64_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX, SupportsPreserveAllABI};
return true;
case IR::OP_VPCMPISTRX:
*Info = {FABI_I32_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, FEXCORE_HAS_PRESERVE_ALL_ATTR};
*Info = {FABI_I32_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, SupportsPreserveAllABI};
return true;
default:
@@ -45,6 +45,6 @@ namespace FEXCore::CPU {
class InterpreterOps {
public:
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(IR::IROp_Header const *IROp, FallbackInfo *Info);
static bool GetFallbackHandler(bool SupportsPreserveAllABI, IR::IROp_Header const *IROp, FallbackInfo *Info);
};
} // namespace FEXCore::CPU
+257 -103
View File
@@ -85,18 +85,51 @@ DEF_OP(Add) {
}
}
DEF_OP(AddNZCV) {
auto Op = IROp->C<IR::IROp_AddNZCV>();
const auto OpSize = IROp->Size;
DEF_OP(AddWithFlags) {
auto Op = IROp->C<IR::IROp_AddWithFlags>();
const uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
cmn(EmitSize, GetReg(Op->Src1.ID()), Const);
adds(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), Const);
} else {
cmn(EmitSize, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
adds(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
}
}
DEF_OP(AddShift) {
auto Op = IROp->C<IR::IROp_AddShift>();
const uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
add(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
DEF_OP(AddNZCV) {
auto Op = IROp->C<IR::IROp_AddNZCV>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
auto Src1 = GetReg(Op->Src1.ID());
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
LOGMAN_THROW_AA_FMT(OpSize >= 4, "Constant not allowed here");
cmn(EmitSize, Src1, Const);
} else {
unsigned Shift = OpSize < 4 ? (32 - (8 * OpSize)) : 0;
if (OpSize < 4) {
lsl(ARMEmitter::Size::i32Bit, TMP1, Src1, Shift);
cmn(EmitSize, TMP1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
} else {
cmn(EmitSize, Src1, GetReg(Op->Src2.ID()));
}
}
}
@@ -110,6 +143,36 @@ DEF_OP(AdcNZCV) {
adcs(EmitSize, ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
}
DEF_OP(AdcWithFlags) {
auto Op = IROp->C<IR::IROp_AdcWithFlags>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
adcs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
}
DEF_OP(Adc) {
auto Op = IROp->C<IR::IROp_Adc>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
adc(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
}
DEF_OP(SbbWithFlags) {
auto Op = IROp->C<IR::IROp_SbbWithFlags>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
sbcs(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
}
DEF_OP(SbbNZCV) {
auto Op = IROp->C<IR::IROp_SbbNZCV>();
const auto OpSize = IROp->Size;
@@ -120,6 +183,16 @@ DEF_OP(SbbNZCV) {
sbcs(EmitSize, ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
}
DEF_OP(Sbb) {
auto Op = IROp->C<IR::IROp_Sbb>();
const auto OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
sbc(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
}
DEF_OP(TestNZ) {
auto Op = IROp->C<IR::IROp_TestNZ>();
const uint8_t OpSize = IROp->Size;
@@ -168,7 +241,7 @@ DEF_OP(Sub) {
if (IsInlineConstant(Op->Src2, &Const)) {
sub(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), Const);
} else {
sub(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
sub(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
}
}
@@ -182,21 +255,47 @@ DEF_OP(SubShift) {
sub(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
}
DEF_OP(SubNZCV) {
auto Op = IROp->C<IR::IROp_SubNZCV>();
const auto OpSize = IROp->Size;
DEF_OP(SubWithFlags) {
auto Op = IROp->C<IR::IROp_SubWithFlags>();
const uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
cmp(EmitSize, GetReg(Op->Src1.ID()), Const);
} else if (IsInlineConstant(Op->Src1, &Const)) {
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
cmp(EmitSize, ARMEmitter::Reg::zr, GetReg(Op->Src2.ID()));
subs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), Const);
} else {
cmp(EmitSize, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
subs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
}
}
DEF_OP(SubNZCV) {
auto Op = IROp->C<IR::IROp_SubNZCV>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
uint64_t Const;
if (IsInlineConstant(Op->Src2, &Const)) {
LOGMAN_THROW_AA_FMT(OpSize >= 4, "Constant not allowed here");
cmp(EmitSize, GetReg(Op->Src1.ID()), Const);
} else {
unsigned Shift = OpSize < 4 ? (32 - (8 * OpSize)) : 0;
ARMEmitter::Register ShiftedSrc1 = GetZeroableReg(Op->Src1);
// Shift to fix flags for <32-bit ops.
// Any shift of zero is still zero so optimize out silly zero shifts.
if (OpSize < 4 && ShiftedSrc1 != ARMEmitter::Reg::zr) {
lsl(ARMEmitter::Size::i32Bit, TMP1, ShiftedSrc1, Shift);
ShiftedSrc1 = TMP1;
}
if (OpSize < 4) {
cmp(EmitSize, ShiftedSrc1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
} else {
cmp(EmitSize, ShiftedSrc1, GetReg(Op->Src2.ID()));
}
}
}
@@ -212,6 +311,20 @@ DEF_OP(RmifNZCV) {
rmif(GetReg(Op->Src.ID()).X(), Op->Rotate, Op->Mask);
}
DEF_OP(SetSmallNZV) {
auto Op = IROp->C<IR::IROp_SetSmallNZV>();
LOGMAN_THROW_A_FMT(CTX->HostFeatures.SupportsFlagM, "Unsupported flagm op");
const uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 1 || OpSize == 2, "Unsupported {} size: {}", __func__, OpSize);
if (OpSize == 1) {
setf8(GetReg(Op->Src.ID()).W());
} else {
setf16(GetReg(Op->Src.ID()).W());
}
}
DEF_OP(AXFlag) {
LOGMAN_THROW_A_FMT(CTX->HostFeatures.SupportsFlagM2, "Unsupported flagm2 op");
axflag();
@@ -254,9 +367,7 @@ DEF_OP(CondAddNZCV) {
ARMEmitter::StatusFlags Flags = (ARMEmitter::StatusFlags)Op->FalseNZCV;
uint64_t Const = 0;
auto Src1 = IsInlineConstant(Op->Src1, &Const) ? ARMEmitter::Reg::zr :
GetReg(Op->Src1.ID());
LOGMAN_THROW_A_FMT(Const == 0, "Unsupported inline constant");
auto Src1 = GetZeroableReg(Op->Src1);
if (IsInlineConstant(Op->Src2, &Const)) {
ccmn(EmitSize, Src1, Const, Flags, MapSelectCC(Op->Cond));
@@ -298,6 +409,16 @@ DEF_OP(UMul) {
mul(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
}
DEF_OP(UMull) {
auto Op = IROp->C<IR::IROp_UMull>();
umull(GetReg(Node).X(), GetReg(Op->Src1.ID()).W(), GetReg(Op->Src2.ID()).W());
}
DEF_OP(SMull) {
auto Op = IROp->C<IR::IROp_SMull>();
smull(GetReg(Node).X(), GetReg(Op->Src1.ID()).W(), GetReg(Op->Src2.ID()).W());
}
DEF_OP(Div) {
auto Op = IROp->C<IR::IROp_Div>();
@@ -543,6 +664,41 @@ DEF_OP(And) {
}
}
DEF_OP(AndWithFlags) {
auto Op = IROp->C<IR::IROp_AndWithFlags>();
const uint8_t OpSize = IROp->Size;
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
uint64_t Const;
const auto Dst = GetReg(Node);
auto Src1 = GetReg(Op->Src1.ID());
// See TestNZ
if (OpSize < 4) {
if (IsInlineConstant(Op->Src2, &Const)) {
and_(EmitSize, Dst, Src1, Const);
} else {
auto Src2 = GetReg(Op->Src2.ID());
if (Src1 != Src2) {
and_(EmitSize, Dst, Src1, Src2);
} else if (Dst != Src1) {
mov(ARMEmitter::Size::i64Bit, Dst, Src1);
}
}
unsigned Shift = 32 - (OpSize * 8);
cmn(EmitSize, ARMEmitter::Reg::zr, Dst, ARMEmitter::ShiftType::LSL, Shift);
} else {
if (IsInlineConstant(Op->Src2, &Const)) {
ands(EmitSize, Dst, Src1, Const);
} else {
const auto Src2 = GetReg(Op->Src2.ID());
ands(EmitSize, Dst, Src1, Src2);
}
}
}
DEF_OP(Andn) {
auto Op = IROp->C<IR::IROp_Andn>();
const uint8_t OpSize = IROp->Size;
@@ -693,64 +849,58 @@ DEF_OP(PDep) {
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Input = GetReg(Op->Input.ID());
const auto Mask = GetReg(Op->Mask.ID());
const auto Dest = GetReg(Node);
const auto ShiftedBitReg = TMP1.R();
const auto BitReg = TMP2.R();
const auto SubMaskReg = TMP3.R();
const auto IndexReg = TMP4.R();
const auto ZeroReg = ARMEmitter::Reg::zr;
// PDep implementation follows the ideas from
// http://0x80.pl/articles/pdep-soft-emu.html ... Basically, iterate the *set*
// bits only, which will be faster than the naive implementation as long as
// there are enough holes in the mask.
//
// The specific arm64 assembly used is based on the sequence that clang
// generates for the C code, giving context to the scheduling yielding better
// ILP than I would do by hand. The registers are allocated by hand however,
// to fit within the tight constraints we have here withot spilling. Also, we
// use cbz/cbnz for conditional branching to avoid clobbering NZCV.
const auto InputReg = StaticRegisters[0];
const auto MaskReg = StaticRegisters[1];
const auto DestReg = StaticRegisters[2];
// We can't clobber these
const auto OrigInput = GetReg(Op->Input.ID());
const auto OrigMask = GetReg(Op->Mask.ID());
const auto SpillCode = 1U << InputReg.Idx() |
1U << MaskReg.Idx() |
1U << DestReg.Idx();
// So we have shadow as temporaries
const auto Input = TMP1.R();
const auto Mask = TMP2.R();
// these get used variously as scratch
const auto T0 = TMP3.R();
const auto T1 = TMP4.R();
ARMEmitter::ForwardLabel EarlyExit;
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::ForwardLabel Done;
cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, IndexReg, ZeroReg);
ARMEmitter::SingleUseForwardLabel Done;
// We sadly need to spill regs for this for the time being
// TODO: Remove when scratch registers can be allocated
// explicitly.
SpillStaticRegs(TMP1, false, SpillCode);
// First, copy the input/mask, since we'll be clobbering. Copy as 64-bit to
// make this 0-uop on Firestorm.
mov(ARMEmitter::Size::i64Bit, Input, OrigInput);
mov(ARMEmitter::Size::i64Bit, Mask, OrigMask);
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
cbz(EmitSize, OrigMask, &Done);
mov(EmitSize, InputReg, Input);
mov(EmitSize, MaskReg, Mask);
mov(EmitSize, DestReg, ZeroReg);
// Setup for first iteration
neg(EmitSize, T0, Mask);
and_(EmitSize, T0, T0, Mask);
// Main loop
Bind(&NextBit);
rbit(EmitSize, ShiftedBitReg, MaskReg);
clz(EmitSize, ShiftedBitReg, ShiftedBitReg);
lsrv(EmitSize, BitReg, InputReg, IndexReg);
and_(EmitSize, BitReg, BitReg, 1);
sub(EmitSize, SubMaskReg, MaskReg, 1);
add(EmitSize, IndexReg, IndexReg, 1);
ands(EmitSize, MaskReg, MaskReg, SubMaskReg);
lslv(EmitSize, ShiftedBitReg, BitReg, ShiftedBitReg);
orr(EmitSize, DestReg, DestReg, ShiftedBitReg);
b(ARMEmitter::Condition::CC_NE, &NextBit);
// Store result in a temp so it doesn't get clobbered.
// and restore it after the re-fill below.
mov(EmitSize, IndexReg, DestReg);
// Restore our registers before leaving
// TODO: Also remove along with above TODO.
FillStaticRegs(false, SpillCode);
mov(EmitSize, Dest, IndexReg);
b(&Done);
// Early exit
Bind(&EarlyExit);
mov(EmitSize, Dest, ZeroReg);
sbfx(EmitSize, T1, Input, 0, 1);
eor(EmitSize, Mask, Mask, T0);
and_(EmitSize, T0, T1, T0);
neg(EmitSize, T1, Mask);
orr(EmitSize, Dest, Dest, T0);
lsr(EmitSize, Input, Input, 1);
and_(EmitSize, T0, Mask, T1);
cbnz(EmitSize, T0, &NextBit);
// All done with nothing to do.
Bind(&Done);
@@ -772,9 +922,9 @@ DEF_OP(PExt) {
const auto BitReg = TMP2;
const auto ValueReg = TMP3;
ARMEmitter::ForwardLabel EarlyExit;
ARMEmitter::SingleUseForwardLabel EarlyExit;
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::ForwardLabel Done;
ARMEmitter::SingleUseForwardLabel Done;
cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, MaskReg, Mask);
@@ -828,8 +978,8 @@ DEF_OP(LDiv) {
break;
}
case 8: {
ARMEmitter::ForwardLabel Only64Bit{};
ARMEmitter::ForwardLabel LongDIVRet{};
ARMEmitter::SingleUseForwardLabel Only64Bit{};
ARMEmitter::SingleUseForwardLabel LongDIVRet{};
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
@@ -841,18 +991,18 @@ DEF_OP(LDiv) {
// Long divide
{
mov(EmitSize, ARMEmitter::Reg::r0, Upper);
mov(EmitSize, ARMEmitter::Reg::r1, Lower);
mov(EmitSize, ARMEmitter::Reg::r2, Divisor);
mov(EmitSize, TMP1, Upper);
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LDIVHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(ARMEmitter::Reg::r3);
blr(TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
// Move result to its destination register
mov(EmitSize, Dst, ARMEmitter::Reg::r0);
mov(EmitSize, Dst, TMP1);
// Skip 64-bit path
b(&LongDIVRet);
@@ -900,8 +1050,8 @@ DEF_OP(LUDiv) {
break;
}
case 8: {
ARMEmitter::ForwardLabel Only64Bit{};
ARMEmitter::ForwardLabel LongDIVRet{};
ARMEmitter::SingleUseForwardLabel Only64Bit{};
ARMEmitter::SingleUseForwardLabel LongDIVRet{};
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
@@ -909,18 +1059,18 @@ DEF_OP(LUDiv) {
// Long divide
{
mov(EmitSize, ARMEmitter::Reg::r0, Upper);
mov(EmitSize, ARMEmitter::Reg::r1, Lower);
mov(EmitSize, ARMEmitter::Reg::r2, Divisor);
mov(EmitSize, TMP1, Upper);
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUDIVHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(ARMEmitter::Reg::r3);
blr(TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
// Move result to its destination register
mov(EmitSize, Dst, ARMEmitter::Reg::r0);
mov(EmitSize, Dst, TMP1);
// Skip 64-bit path
b(&LongDIVRet);
@@ -972,8 +1122,8 @@ DEF_OP(LRem) {
break;
}
case 8: {
ARMEmitter::ForwardLabel Only64Bit{};
ARMEmitter::ForwardLabel LongDIVRet{};
ARMEmitter::SingleUseForwardLabel Only64Bit{};
ARMEmitter::SingleUseForwardLabel LongDIVRet{};
// Check if the upper bits match the top bit of the lower 64-bits
// Sign extend the top bit of lower bits
@@ -985,18 +1135,18 @@ DEF_OP(LRem) {
// Long divide
{
mov(EmitSize, ARMEmitter::Reg::r0, Upper);
mov(EmitSize, ARMEmitter::Reg::r1, Lower);
mov(EmitSize, ARMEmitter::Reg::r2, Divisor);
mov(EmitSize, TMP1, Upper);
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREMHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LREMHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(ARMEmitter::Reg::r3);
blr(TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
// Move result to its destination register
mov(EmitSize, Dst, ARMEmitter::Reg::r0);
mov(EmitSize, Dst, TMP1);
// Skip 64-bit path
b(&LongDIVRet);
@@ -1046,8 +1196,8 @@ DEF_OP(LURem) {
break;
}
case 8: {
ARMEmitter::ForwardLabel Only64Bit{};
ARMEmitter::ForwardLabel LongDIVRet{};
ARMEmitter::SingleUseForwardLabel Only64Bit{};
ARMEmitter::SingleUseForwardLabel LongDIVRet{};
// Check the upper bits for zero
// If the upper bits are zero then we can do a 64-bit divide
@@ -1055,18 +1205,18 @@ DEF_OP(LURem) {
// Long divide
{
mov(EmitSize, ARMEmitter::Reg::r0, Upper);
mov(EmitSize, ARMEmitter::Reg::r1, Lower);
mov(EmitSize, ARMEmitter::Reg::r2, Divisor);
mov(EmitSize, TMP1, Upper);
mov(EmitSize, TMP2, Lower);
mov(EmitSize, TMP3, Divisor);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREMHandler));
ldr(TMP4, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.AArch64.LUREMHandler));
str<ARMEmitter::IndexType::PRE>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, -16);
blr(ARMEmitter::Reg::r3);
blr(TMP4);
ldr<ARMEmitter::IndexType::POST>(ARMEmitter::XReg::lr, ARMEmitter::Reg::rsp, 16);
// Move result to its destination register
mov(EmitSize, Dst, ARMEmitter::Reg::r0);
mov(EmitSize, Dst, TMP1);
// Skip 64-bit path
b(&LongDIVRet);
@@ -1199,6 +1349,11 @@ DEF_OP(FindTrailingZeroes) {
rbit(EmitSize, Dst, Src);
if (OpSize == 2) {
// This orr does two things. First, if the (masked) source is zero, it
// reverses to zero in the top so it forces clz to return 16. Second, it
// ensures garbage in the upper bits of the source don't affect clz, because
// they'll rbit to garbage in the bottom below the 0x8000 and be ignored by
// the clz. So we handle Src upper garbage without explicitly masking.
orr(EmitSize, Dst, Dst, 0x8000);
}
@@ -1216,6 +1371,8 @@ DEF_OP(CountLeadingZeroes) {
const auto Src = GetReg(Op->Src.ID());
if (OpSize == 2) {
// Expressing as lsl+orr+clz clears away any garbage in the upper bits
// (alternatively could do uxth+clz+sub.. equal cost in total).
lsl(EmitSize, Dst, Src, 16);
orr(EmitSize, Dst, Dst, 0x8000);
clz(EmitSize, Dst, Dst);
@@ -1414,11 +1571,8 @@ DEF_OP(NZCVSelect) {
csetm(EmitSize, Dst, cc);
else
cset(EmitSize, Dst, cc);
} else if (is_const_false) {
LOGMAN_THROW_A_FMT(const_false == 0, "NZCVSelect: unsupported constant");
csel(EmitSize, Dst, GetReg(Op->TrueVal.ID()), ARMEmitter::Reg::zr, cc);
} else {
csel(EmitSize, Dst, GetReg(Op->TrueVal.ID()), GetReg(Op->FalseVal.ID()), cc);
csel(EmitSize, Dst, GetReg(Op->TrueVal.ID()), GetZeroableReg(Op->FalseVal), cc);
}
}
@@ -32,8 +32,8 @@ DEF_OP(CASPair) {
}
else {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
Bind(&LoopTop);
ldaxp(EmitSize, TMP2, TMP3, MemSrc);
@@ -82,8 +82,8 @@ DEF_OP(CAS) {
}
else {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::ForwardLabel LoopNotExpected;
ARMEmitter::ForwardLabel LoopExpected;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
if (OpSize == 1) {
@@ -310,8 +310,7 @@ DEF_OP(AtomicSwap) {
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mov(EmitSize, TMP2, Src);
ldswpal(SubEmitSize, TMP2, GetReg(Node), MemSrc);
ldswpal(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
ARMEmitter::BackwardLabel LoopTop;
@@ -11,9 +11,9 @@ $end_info$
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/MathUtils.h>
#include <Interface/HLE/Thunks/Thunks.h>
@@ -53,33 +53,30 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
ARMEmitter::ForwardLabel l_BranchHost;
ARMEmitter::ForwardLabel l_BranchGuest;
ARMEmitter::SingleUseForwardLabel l_BranchHost;
ldr(ARMEmitter::XReg::x0, &l_BranchHost);
blr(ARMEmitter::Reg::r0);
ldr(TMP1, &l_BranchHost);
blr(TMP1);
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
Bind(&l_BranchGuest);
dc64(NewRIP);
} else {
ARMEmitter::ForwardLabel FullLookup;
ARMEmitter::SingleUseForwardLabel FullLookup;
auto RipReg = GetReg(Op->NewRIP.ID());
// L1 Cache
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, RipReg, LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, ARMEmitter::XReg::x3, ARMEmitter::ShiftType::LSL, 4);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg, LookupCache::L1_ENTRIES_MASK);
add(TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
// Note: sub+cbnz used over cmp+br to preserve flags.
ldp<ARMEmitter::IndexType::OFFSET>(ARMEmitter::XReg::x1, ARMEmitter::XReg::x0, ARMEmitter::Reg::r0, 0);
sub(TMP1, ARMEmitter::XReg::x0, RipReg.X());
ldp<ARMEmitter::IndexType::OFFSET>(TMP2, TMP1, TMP1, 0);
sub(TMP1, TMP1, RipReg.X());
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
br(ARMEmitter::Reg::r1);
br(TMP2);
Bind(&FullLookup);
ldr(TMP1, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.DispatcherLoopTop));
@@ -131,8 +128,8 @@ DEF_OP(CondJump) {
if (Op->FromNZCV) {
b(MapBranchCC(Op->Cond), TrueTargetLabel);
} else {
uint64_t Const;
const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
const auto Size = Op->CompareSize == 4 ? ARMEmitter::Size::i32Bit : ARMEmitter::Size::i64Bit;
@@ -353,58 +350,58 @@ DEF_OP(ValidateCode) {
int idx = 0;
LoadConstant(ARMEmitter::Size::i64Bit, GetReg(Node), 0);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, Entry + Op->Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 1);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, Entry + Op->Offset);
LoadConstant(ARMEmitter::Size::i64Bit, TMP2, 1);
const auto Dst = GetReg(Node);
while (len >= 8)
{
ldr(ARMEmitter::XReg::x2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
ldr(ARMEmitter::XReg::x2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 8;
idx += 8;
}
while (len >= 4)
{
ldr(ARMEmitter::WReg::w2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
ldr(ARMEmitter::WReg::w2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 4;
idx += 4;
}
while (len >= 2)
{
ldrh(ARMEmitter::Reg::r2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint16_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
ldrh(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint16_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 2;
idx += 2;
}
while (len >= 1)
{
ldrb(ARMEmitter::Reg::r2, ARMEmitter::Reg::r0, idx);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r3, *(const uint8_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::r3);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, ARMEmitter::Reg::r1, ARMEmitter::Condition::CC_EQ);
ldrb(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint8_t *)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 1;
idx += 1;
}
}
DEF_OP(ThreadRemoveCodeEntry) {
PushDynamicRegsAndLR(TMP4);
SpillStaticRegs(TMP4);
// Arguments are passed as follows:
// X0: Thread
// X1: RIP
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, STATE.R());
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Entry);
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
@@ -423,16 +420,23 @@ DEF_OP(ThreadRemoveCodeEntry) {
DEF_OP(CPUID) {
auto Op = IROp->C<IR::IROp_CPUID>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
mov(ARMEmitter::Size::i64Bit, TMP2, GetReg(Op->Function.ID()));
mov(ARMEmitter::Size::i64Bit, TMP3, GetReg(Op->Leaf.ID()));
PushDynamicRegsAndLR(TMP4);
SpillStaticRegs(TMP4);
// x0 = CPUID Handler
// x1 = CPUID Function
// x2 = CPUID Leaf
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDFunction));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, GetReg(Op->Leaf.ID()));
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, TMP2);
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, TMP3);
}
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
@@ -440,6 +444,11 @@ DEF_OP(CPUID) {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, TMP1, ARMEmitter::Reg::r0);
mov(ARMEmitter::Size::i64Bit, TMP2, ARMEmitter::Reg::r1);
}
FillStaticRegs();
PopDynamicRegsAndLR();
@@ -447,21 +456,22 @@ DEF_OP(CPUID) {
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i64Bit, Dst.first, ARMEmitter::Reg::r0);
mov(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r1);
mov(ARMEmitter::Size::i64Bit, Dst.first, TMP1);
mov(ARMEmitter::Size::i64Bit, Dst.second, TMP2);
}
DEF_OP(XGetBV) {
auto Op = IROp->C<IR::IROp_XGetBV>();
PushDynamicRegsAndLR(TMP1);
SpillStaticRegs(TMP1);
PushDynamicRegsAndLR(TMP4);
SpillStaticRegs(TMP4);
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
// x0 = CPUID Handler
// x1 = XCR Function
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.CPUIDObj));
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction));
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, GetReg(Op->Function.ID()));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, void*, uint32_t>(ARMEmitter::Reg::r2);
}
@@ -469,6 +479,10 @@ DEF_OP(XGetBV) {
blr(ARMEmitter::Reg::r2);
}
if (!TMP_ABIARGS) {
mov(ARMEmitter::Size::i64Bit, TMP1, ARMEmitter::Reg::r0);
}
FillStaticRegs();
PopDynamicRegsAndLR();
@@ -476,8 +490,8 @@ DEF_OP(XGetBV) {
// Results are in x0
// Results want to be in a i32v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i32Bit, Dst.first, ARMEmitter::Reg::r0);
lsr(ARMEmitter::Size::i64Bit, Dst.second, ARMEmitter::Reg::r0, 32);
mov(ARMEmitter::Size::i32Bit, Dst.first, TMP1);
lsr(ARMEmitter::Size::i64Bit, Dst.second, TMP1, 32);
}
#undef DEF_OP
+127 -109
View File
@@ -18,17 +18,18 @@ $end_info$
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/Core/InternalThreadState.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Utils/MemberFunctionToPointer.h"
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include "Interface/Core/Interpreter/InterpreterOps.h"
@@ -78,11 +79,45 @@ namespace FEXCore::CPU {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(IROp, &Info)) {
if (!InterpreterOps::GetFallbackHandler(CTX->HostFeatures.SupportsPreserveAllABI, IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_MSG_A_FMT("Unhandled IR Op: {}", FEXCore::IR::GetName(IROp->Op));
#endif
} else {
auto FillF80Result = [&]() {
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
mov(TMP2, ARMEmitter::XReg::x1);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, TMP1);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, TMP2);
};
auto FillF64Result = [&]() {
if (!TMP_ABIARGS) {
mov(VTMP1.D(), ARMEmitter::DReg::d0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
mov(Dst.D(), VTMP1.D());
};
auto FillI32Result = [&]() {
if (!TMP_ABIARGS) {
mov(TMP1.W(), ARMEmitter::WReg::w0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(Dst.W(), TMP1.W());
};
switch(Info.ABI) {
case FABI_F80_I16_F32:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
@@ -98,12 +133,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r1);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, ARMEmitter::Reg::r0);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
FillF80Result();
}
break;
@@ -121,12 +151,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r1);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, ARMEmitter::Reg::r0);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
FillF80Result();
}
break;
@@ -135,13 +160,13 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
if (Info.ABI == FABI_F80_I16_I16) {
sxth(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
else {
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint32_t>(ARMEmitter::Reg::r2);
@@ -150,12 +175,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r2);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, ARMEmitter::Reg::r0);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
FillF80Result();
}
break;
@@ -176,10 +196,13 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
fmov(VTMP1.S(), ARMEmitter::SReg::s0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
fmov(Dst.S(), ARMEmitter::SReg::s0);
fmov(Dst.S(), VTMP1.S());
}
break;
@@ -200,10 +223,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r3);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
mov(Dst.D(), ARMEmitter::DReg::d0);
FillF64Result();
}
break;
@@ -222,21 +242,24 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r1);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
mov(Dst.D(), ARMEmitter::DReg::d0);
FillF64Result();
}
break;
case FABI_F64_I16_F64_F64: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
mov(ARMEmitter::DReg::d1, Src2.D());
mov(VTMP1.D(), Src1.D());
mov(VTMP2.D(), Src2.D());
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
if (!TMP_ABIARGS) {
mov(ARMEmitter::DReg::d0, VTMP1.D());
mov(ARMEmitter::DReg::d1, VTMP2.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
@@ -246,10 +269,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r1);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
mov(Dst.D(), ARMEmitter::DReg::d0);
FillF64Result();
}
break;
@@ -270,10 +290,13 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
sxth(ARMEmitter::Size::i64Bit, Dst, ARMEmitter::Reg::r0);
sxth(ARMEmitter::Size::i64Bit, Dst, TMP1);
}
break;
case FABI_I32_I16_F80:{
@@ -293,10 +316,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r3);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i32Bit, Dst, ARMEmitter::Reg::r0);
FillI32Result();
}
break;
case FABI_I64_I16_F80:{
@@ -315,10 +335,14 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
else {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, ARMEmitter::Reg::r0);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
}
break;
case FABI_I64_I16_F80_F80:{
@@ -341,10 +365,14 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
else {
blr(ARMEmitter::Reg::r5);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, ARMEmitter::Reg::r0);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
}
break;
case FABI_F80_I16_F80:{
@@ -364,12 +392,7 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r3);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, ARMEmitter::Reg::r0);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
FillF80Result();
}
break;
case FABI_F80_I16_F80_F80:{
@@ -393,27 +416,28 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r5);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
eor(Dst.Q(), Dst.Q(), Dst.Q());
ins(ARMEmitter::SubRegSize::i64Bit, Dst, 0, ARMEmitter::Reg::r0);
ins(ARMEmitter::SubRegSize::i16Bit, Dst, 4, ARMEmitter::Reg::r1);
FillF80Result();
}
break;
case FABI_I32_I64_I64_I128_I128_I16: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
mov(TMP1, SrcRAX.X());
mov(TMP2, SrcRDX.X());
SpillForABICall(Info.SupportsPreserveAllABI, TMP3, true);
const auto Control = Op->Control;
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
mov(ARMEmitter::XReg::x0, SrcRAX.X());
mov(ARMEmitter::XReg::x1, SrcRDX.X());
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x0, TMP1);
mov(ARMEmitter::XReg::x1, TMP2);
}
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src1, 1);
@@ -431,12 +455,9 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r7);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(Dst.W(), ARMEmitter::WReg::w0);
break;
FillI32Result();
}
break;
case FABI_I32_I128_I128_I16: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
@@ -462,12 +483,9 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
blr(ARMEmitter::Reg::r5);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(Dst.W(), ARMEmitter::WReg::w0);
break;
FillI32Result();
}
break;
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -480,9 +498,26 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, uint64_t *record) {
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
uintptr_t branch = (uintptr_t)(Record) - 8;
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
FEXCore::ARMEmitter::SingleUseForwardLabel l_BranchHost;
emit.ldr(TMP1, &l_BranchHost);
emit.blr(TMP1);
emit.Bind(&l_BranchHost);
emit.dc64(LinkerAddress);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 8);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Record->HostBranch = LinkerAddress;
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
auto Thread = Frame->Thread;
auto GuestRip = record[1];
auto GuestRip = Record->GuestRIP;
auto HostCode = Thread->LookupCache->FindBlock(GuestRip);
@@ -491,35 +526,24 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
return Frame->Pointers.Common.DispatcherLoopTop;
}
uintptr_t branch = (uintptr_t)(record) - 8;
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
uintptr_t branch = (uintptr_t)(Record) - 8;
auto offset = HostCode/4 - branch/4;
if (vixl::IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 24);
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
emit.b(offset);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 24);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 4);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [branch, LinkerAddress]{
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 24);
FEXCore::ARMEmitter::ForwardLabel l_BranchHost;
emit.ldr(FEXCore::ARMEmitter::XReg::x0, &l_BranchHost);
emit.blr(FEXCore::ARMEmitter::Reg::r0);
emit.Bind(&l_BranchHost);
emit.dc64(LinkerAddress);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 24);
});
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
} else {
// fallback case - do a soft-er link by patching the pointer
record[0] = HostCode;
Record->HostBranch = HostCode;
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, (uintptr_t)record, [record, LinkerAddress]{
record[0] = LinkerAddress;
});
Thread->LookupCache->AddBlockLink(GuestRip, Record, IndirectBlockDelinker);
}
return HostCode;
@@ -574,8 +598,11 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
Common.XCRFunction = PMF.GetConvertedPointer();
}
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = reinterpret_cast<uint64_t>(FEXCore::Context::HandleSyscall);
{
FEXCore::Utils::MemberFunctionToPointerCast PMF(&FEXCore::HLE::SyscallHandler::HandleSyscall);
Common.SyscallHandlerObj = reinterpret_cast<uint64_t>(CTX->SyscallHandler);
Common.SyscallHandlerFunc = PMF.GetVTableEntry(CTX->SyscallHandler);
}
Common.ExitFunctionLink = reinterpret_cast<uintptr_t>(&Context::ContextImpl::ThreadExitFunctionLink<Arm64JITCore_ExitFunctionLink>);
// Fill in the fallback handlers
@@ -594,15 +621,6 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
ClearCache();
// Setup dynamic dispatch.
if (CTX->Dispatcher->GetConfig().StaticRegisterAllocation) {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegisterSRA;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegisterSRA;
}
else {
RT_LoadRegister = &Arm64JITCore::Op_LoadRegister;
RT_StoreRegister = &Arm64JITCore::Op_StoreRegister;
}
if (ParanoidTSO()) {
RT_LoadMemTSO = &Arm64JITCore::Op_ParanoidLoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_ParanoidStoreMemTSO;
@@ -762,8 +780,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
if (vixl::aarch64::Assembler::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, ARMEmitter::XReg::x0, ARMEmitter::ExtendedType::LSL_64, 0);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
@@ -840,6 +858,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
// TODO: This needs to be a data RIP relocation once code caching works.
// Current relocation code doesn't support this feature yet.
JITBlockTail->RIP = Entry;
JITBlockTail->SpinLockFutex = 0;
{
// Store the RIP entries.
@@ -881,7 +900,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
LogMan::Msg::IFmt("Disassemble Begin");
for (auto PCToDecode = DisasmBegin; PCToDecode < DisasmEnd; PCToDecode += 4) {
DisasmDecoder->Decode(PCToDecode);
auto Output = Disasm.GetOutput();
auto Output = Disasm->GetOutput();
LogMan::Msg::IFmt("{}", Output);
}
LogMan::Msg::IFmt("Disassemble End");
@@ -909,8 +928,8 @@ void Arm64JITCore::ResetStack() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
// Too big to fit in a 12bit immediate
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, TotalSpillSlotsSize);
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, ARMEmitter::XReg::x0, ARMEmitter::ExtendedType::LSL_64, 0);
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::rsp, ARMEmitter::XReg::rsp, TMP1, ARMEmitter::ExtendedType::LSL_64, 0);
}
}
@@ -920,7 +939,6 @@ fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *
CPUBackendFeatures GetArm64JITBackendFeatures() {
return CPUBackendFeatures {
.SupportsStaticRegisterAllocation = true,
.SupportsFlags = true,
.SupportsSaturatingRoundingShifts = true,
.SupportsVTBL2 = true,
@@ -116,6 +116,16 @@ private:
return PhyReg;
}
[[nodiscard]] FEXCore::ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
uint64_t Const;
if (IsInlineConstant(Src, &Const)) {
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
return ARMEmitter::Reg::zr;
} else {
return GetReg(Src.ID());
}
}
// Converts IR-base shift type to ARMEmitter shift type.
// Will be a no-op, only a type conversion since the two definitions match.
[[nodiscard]] ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) const {
@@ -226,9 +236,6 @@ private:
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
// Runtime selection;
// Load and store register style.
OpType RT_LoadRegister;
OpType RT_StoreRegister;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
OpType RT_StoreMemTSO;
@@ -236,9 +243,6 @@ private:
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
// Dynamic Dispatcher supporting operations
DEF_OP(LoadRegisterSRA);
DEF_OP(StoreRegisterSRA);
DEF_OP(ParanoidLoadMemTSO);
DEF_OP(ParanoidStoreMemTSO);
@@ -131,170 +131,6 @@ DEF_OP(LoadRegister) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < StaticRegisters.size(), "out of range regId");
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
ldrb(GetReg(Node), STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldrh(GetReg(Node), STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
ldr(GetReg(Node).X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled LoadRegister GPR size: {}", OpSize);
break;
}
}
else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE256 ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE256, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < StaticFPRegisters.size(), "out of range regId");
const auto host = GetVReg(Node);
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrb(host, STATE, Op->Offset);
break;
}
case 2: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldrh(host, STATE, Op->Offset);
break;
}
case 4: {
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.S(), STATE, Op->Offset);
break;
}
case 8: {
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
ldr(host.D(), STATE, Op->Offset);
break;
}
case 16: {
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
ldr(host.Q(), STATE, Op->Offset);
break;
}
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
[[maybe_unused]] const auto regId = (Op->Offset / Core::CPUState::GPR_REG_SIZE) - 1;
const auto regOffs = Op->Offset & 7;
LOGMAN_THROW_A_FMT(regId < StaticFPRegisters.size(), "out of range regId");
const auto Src = GetReg(Op->Value.ID());
switch (OpSize) {
case 1:
LOGMAN_THROW_AA_FMT(regOffs == 0 || regOffs == 1, "unexpected regOffs");
strb(Src, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
strh(Src, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.W(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs");
str(Src.X(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister GPR size: {}", OpSize);
break;
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsSVE256 ? Core::CPUState::XMM_AVX_REG_SIZE
: Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regId = (Op->Offset - offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0])) / regSize;
LOGMAN_THROW_A_FMT(HostSupportsSVE256, "Unsupported code path!");
LOGMAN_THROW_A_FMT(regId < StaticFPRegisters.size(), "regId out of range");
const auto host = GetVReg(Op->Value.ID());
const auto regOffs = Op->Offset & 15;
switch (OpSize) {
case 1:
strb(host, STATE, Op->Offset);
break;
case 2:
LOGMAN_THROW_AA_FMT((regOffs & 1) == 0, "unexpected regOffs: {}", regOffs);
strh(host, STATE, Op->Offset);
break;
case 4:
LOGMAN_THROW_AA_FMT((regOffs & 3) == 0, "unexpected regOffs: {}", regOffs);
str(host.S(), STATE, Op->Offset);
break;
case 8:
LOGMAN_THROW_AA_FMT((regOffs & 7) == 0, "unexpected regOffs: {}", regOffs);
str(host.D(), STATE, Op->Offset);
break;
case 16:
LOGMAN_THROW_AA_FMT(regOffs == 0, "unexpected regOffs: {}", regOffs);
str(host.Q(), STATE, Op->Offset);
break;
default:
LOGMAN_MSG_A_FMT("Unhandled StoreRegister FPR size: {}", OpSize);
break;
}
} else {
LOGMAN_THROW_AA_FMT(false, "Unhandled Op->Class {}", Op->Class);
}
}
DEF_OP(LoadRegisterSRA) {
const auto Op = IROp->C<IR::IROp_LoadRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
const auto regId =
Op->Offset == offsetof(Core::CpuStateFrame, State.pf_raw) ? (StaticRegisters.size() - 2) :
@@ -339,7 +175,7 @@ DEF_OP(LoadRegisterSRA) {
if (HostSupportsSVE256) {
const auto regOffs = Op->Offset & 31;
ARMEmitter::ForwardLabel DataLocation;
ARMEmitter::SingleUseForwardLabel DataLocation;
const auto LoadPredicate = [this, &DataLocation] {
const auto Predicate = ARMEmitter::PReg::p0;
adr(TMP1, &DataLocation);
@@ -348,7 +184,7 @@ DEF_OP(LoadRegisterSRA) {
};
const auto EmitData = [this, &DataLocation](uint32_t Value) {
ARMEmitter::ForwardLabel PastConstant;
ARMEmitter::SingleUseForwardLabel PastConstant;
b(&PastConstant);
Bind(&DataLocation);
dc32(Value);
@@ -473,7 +309,7 @@ DEF_OP(LoadRegisterSRA) {
}
}
DEF_OP(StoreRegisterSRA) {
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
@@ -528,7 +364,7 @@ DEF_OP(StoreRegisterSRA) {
const auto regOffs = Op->Offset & 31;
// Compartmentalized setting up of the predicate for the cases that need it.
ARMEmitter::ForwardLabel DataLocation;
ARMEmitter::SingleUseForwardLabel DataLocation;
const auto LoadPredicate = [this, &DataLocation] {
const auto Predicate = ARMEmitter::PReg::p0;
adr(TMP1, &DataLocation);
@@ -541,7 +377,7 @@ DEF_OP(StoreRegisterSRA) {
// It's helpful to treat LoadPredicate and EmitData as a prologue and epilogue
// respectfully.
const auto EmitData = [this, &DataLocation](uint32_t Data) {
ARMEmitter::ForwardLabel PastConstant;
ARMEmitter::SingleUseForwardLabel PastConstant;
b(&PastConstant);
Bind(&DataLocation);
dc32(Data);
@@ -1879,8 +1715,8 @@ DEF_OP(MemSet) {
//
// Counter is decremented regardless.
ARMEmitter::ForwardLabel BackwardImpl{};
ARMEmitter::ForwardLabel Done{};
ARMEmitter::SingleUseForwardLabel BackwardImpl{};
ARMEmitter::SingleUseForwardLabel Done{};
mov(TMP1, Length.X());
if (Op->Prefix.IsInvalid()) {
@@ -1953,7 +1789,7 @@ DEF_OP(MemSet) {
const int32_t SizeDirection = Size * Direction;
ARMEmitter::BackwardLabel AgainInternal{};
ARMEmitter::ForwardLabel DoneInternal{};
ARMEmitter::SingleUseForwardLabel DoneInternal{};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
@@ -2059,8 +1895,8 @@ DEF_OP(MemCpy) {
//
// Counter is decremented regardless.
ARMEmitter::ForwardLabel BackwardImpl{};
ARMEmitter::ForwardLabel Done{};
ARMEmitter::SingleUseForwardLabel BackwardImpl{};
ARMEmitter::SingleUseForwardLabel Done{};
mov(TMP1, Length.X());
if (Op->PrefixDest.IsInvalid()) {
@@ -2214,7 +2050,7 @@ DEF_OP(MemCpy) {
const int32_t SizeDirection = Size * Direction;
ARMEmitter::BackwardLabel AgainInternal{};
ARMEmitter::ForwardLabel DoneInternal{};
ARMEmitter::SingleUseForwardLabel DoneInternal{};
// Early exit if zero count.
cbz(ARMEmitter::Size::i64Bit, TMP1, &DoneInternal);
@@ -1372,6 +1372,32 @@ DEF_OP(VUMinV) {
}
}
DEF_OP(VUMaxV) {
const auto Op = IROp->C<IR::IROp_VUMaxV>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = Op->Header.ElementSize;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8, "Invalid size");
const auto SubRegSize =
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i8Bit;
if (HostSupportsSVE256 && Is256Bit) {
const auto Pred = PRED_TMP_32B;
umaxv(SubRegSize, Dst, Pred, Vector.Z());
} else {
// Vector
umaxv(SubRegSize, Dst.Q(), Vector.Q());
}
}
DEF_OP(VURAvg) {
const auto Op = IROp->C<IR::IROp_VURAvg>();
const auto OpSize = IROp->Size;
@@ -5009,42 +5035,50 @@ DEF_OP(VTBX1) {
if (Dst != VectorSrcDst) {
switch (OpSize) {
case 8: {
mov(Dst.D(), VectorSrcDst.D());
mov(VTMP1.D(), VectorSrcDst.D());
tbx(VTMP1.D(), VectorTable.Q(), VectorIndices.D());
mov(Dst.D(), VTMP1.D());
break;
}
case 16: {
mov(Dst.Q(), VectorSrcDst.Q());
mov(VTMP1.Q(), VectorSrcDst.Q());
tbx(VTMP1.Q(), VectorTable.Q(), VectorIndices.Q());
mov(Dst.Q(), VTMP1.Q());
break;
}
case 32: {
mov(Dst.Z(), VectorSrcDst.Z());
LOGMAN_THROW_AA_FMT(HostSupportsSVE256,
"Host does not support SVE. Cannot perform 256-bit table lookup");
mov(VTMP1.Z(), VectorSrcDst.Z());
tbx(ARMEmitter::SubRegSize::i8Bit, VTMP1.Z(), VectorTable.Z(), VectorIndices.Z());
mov(Dst.Z(), VTMP1.Z());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown OpSize: {}", OpSize);
break;
}
}
} else {
switch (OpSize) {
case 8: {
tbx(VectorSrcDst.D(), VectorTable.Q(), VectorIndices.D());
break;
}
case 16: {
tbx(VectorSrcDst.Q(), VectorTable.Q(), VectorIndices.Q());
break;
}
case 32: {
LOGMAN_THROW_AA_FMT(HostSupportsSVE256,
"Host does not support SVE. Cannot perform 256-bit table lookup");
switch (OpSize) {
case 8: {
tbx(Dst.D(), VectorTable.Q(), VectorIndices.D());
break;
tbx(ARMEmitter::SubRegSize::i8Bit, VectorSrcDst.Z(), VectorTable.Z(), VectorIndices.Z());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown OpSize: {}", OpSize);
break;
}
case 16: {
tbx(Dst.Q(), VectorTable.Q(), VectorIndices.Q());
break;
}
case 32: {
LOGMAN_THROW_AA_FMT(HostSupportsSVE256,
"Host does not support SVE. Cannot perform 256-bit table lookup");
tbx(ARMEmitter::SubRegSize::i8Bit, Dst.Z(), VectorTable.Z(), VectorIndices.Z());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown OpSize: {}", OpSize);
break;
}
}
@@ -15,10 +15,6 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateX86JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
CPUBackendFeatures GetX86JITBackendFeatures();
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
CPUBackendFeatures GetArm64JITBackendFeatures();
+7 -8
View File
@@ -100,15 +100,15 @@ public:
L1Entry.HostCode = (uintptr_t)HostCode;
}
void Erase(uint64_t Address) {
void Erase(FEXCore::Core::CpuStateFrame *Frame, uint64_t Address) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, 0});
auto upper = BlockLinks->upper_bound({Address, UINTPTR_MAX});
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData *>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second();
it->second(Frame, it->first.HostLink);
}
// Remove from BlockList
@@ -141,8 +141,7 @@ public:
BlockPointers[PageOffset].HostCode = 0;
}
void AddBlockLink(uint64_t GuestDestination, uintptr_t HostLink, const std::function<void()> &delinker) {
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData * HostLink, const FEXCore::Context::BlockDelinkerFunc &delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
@@ -224,7 +223,7 @@ private:
struct BlockLinkTag {
uint64_t GuestDestination;
uintptr_t HostLink;
FEXCore::Context::ExitFunctionLinkData *HostLink;
bool operator <(const BlockLinkTag& other) const {
if (GuestDestination < other.GuestDestination)
@@ -243,7 +242,7 @@ private:
//
// This makes `BlockLinks` look like a raw pointer that could memory leak, but since it is backed by the MBR, it won't.
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, std::function<void()>>;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, FEXCore::Context::BlockDelinkerFunc>;
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType *BlockLinks;
@@ -32,9 +32,6 @@ namespace FEXCore::CodeSerialize {
// ABI local flag unsafe optimization
unsigned ABILocalFlags : 1;
// Static register allocation enabled
unsigned SRA : 1;
// Paranoid TSO mode enabled
unsigned ParanoidTSO : 1;
@@ -49,7 +46,7 @@ namespace FEXCore::CodeSerialize {
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 18;
unsigned _Pad : 19;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
@@ -59,7 +56,6 @@ namespace FEXCore::CodeSerialize {
HardwareTSOEnabled == other.HardwareTSOEnabled &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
SRA == other.SRA &&
ParanoidTSO == other.ParanoidTSO &&
Is64BitMode == other.Is64BitMode &&
SMCChecks == other.SMCChecks &&
@@ -75,7 +71,6 @@ namespace FEXCore::CodeSerialize {
Hash <<= 1; Hash |= other.HardwareTSOEnabled;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.SRA;
Hash <<= 1; Hash |= other.ParanoidTSO;
Hash <<= 1; Hash |= other.Is64BitMode;
Hash <<= 2; Hash |= other.SMCChecks;
@@ -18,7 +18,6 @@ namespace FEXCore::CodeSerialize {
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.SRA = ctx->Config.StaticRegisterAllocation;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
File diff suppressed because it is too large. Load diff
@@ -4,13 +4,13 @@
#include "Interface/Core/Frontend.h"
#include "Interface/Core/X86Tables/X86Tables.h"
#include "Interface/Context/Context.h"
#include "Interface/IR/IREmitter.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -78,10 +78,7 @@ friend class FEXCore::IR::PassManager;
public:
enum class FlagsGenerationType : uint8_t {
TYPE_NONE,
TYPE_ADC,
TYPE_SBB,
TYPE_SUB,
TYPE_ADD,
TYPE_MUL,
TYPE_UMUL,
TYPE_LOGICAL,
@@ -102,8 +99,7 @@ public:
TYPE_BLSR,
TYPE_POPCOUNT,
TYPE_BZHI,
TYPE_TZCNT,
TYPE_LZCNT,
TYPE_ZCNT,
TYPE_RDRAND,
};
@@ -261,6 +257,7 @@ public:
template<FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp>
void ALUOp(OpcodeArgs);
void INTOp(OpcodeArgs);
template<bool IsSyscallInst>
void SyscallOp(OpcodeArgs);
void ThunkOp(OpcodeArgs);
void LEAOp(OpcodeArgs);
@@ -271,8 +268,9 @@ public:
void SecondaryALUOp(OpcodeArgs);
template<uint32_t SrcIndex>
void ADCOp(OpcodeArgs);
template<uint32_t SrcIndex, bool SetFlags>
template<uint32_t SrcIndex>
void SBBOp(OpcodeArgs);
void SALCOp(OpcodeArgs);
void PUSHOp(OpcodeArgs);
void PUSHREGOp(OpcodeArgs);
void PUSHAOp(OpcodeArgs);
@@ -352,6 +350,11 @@ public:
void PUSHFOp(OpcodeArgs);
void POPFOp(OpcodeArgs);
struct CycleCounterPair {
OrderedNode *CounterLow;
OrderedNode *CounterHigh;
};
CycleCounterPair CycleCounter();
void RDTSCOp(OpcodeArgs);
void INCOp(OpcodeArgs);
void DECOp(OpcodeArgs);
@@ -867,6 +870,8 @@ public:
void PSADBW(OpcodeArgs);
OrderedNode *BitwiseAtLeastTwo(OrderedNode *A, OrderedNode *B, OrderedNode *C);
void SHA1NEXTEOp(OpcodeArgs);
void SHA1MSG1Op(OpcodeArgs);
void SHA1MSG2Op(OpcodeArgs);
@@ -995,7 +1000,7 @@ private:
// Used during new op bringup
bool ShouldDump{false};
void ALUOpImpl(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp);
void ALUOpImpl(OpcodeArgs, FEXCore::IR::IROps ALUIROp, FEXCore::IR::IROps AtomicFetchOp, unsigned SrcIdx);
// Opcode helpers for generalizing behavior across VEX and non-VEX variants.
@@ -1261,6 +1266,30 @@ private:
return NZCVMask;
}
// Set flag tracking to prepare for an operation that directly writes NZCV. If
// some bits are known to be zeroed, the PossiblySetNZCVBits mask can be
// passed. Otherwise, it defaults to assuming all bits may be set after
// (this is conservative).
void HandleNZCVWrite(uint32_t _PossiblySetNZCVBits = ~0) {
InvalidateDeferredFlags();
CachedNZCV = nullptr;
PossiblySetNZCVBits = _PossiblySetNZCVBits;
NZCVDirty = false;
}
// Set flag tracking to prepare for a read-modify-write operation on NZCV.
void HandleNZCV_RMW(uint32_t _PossiblySetNZCVBits = ~0) {
if (NZCVDirty && CachedNZCV)
_StoreNZCV(CachedNZCV);
HandleNZCVWrite(_PossiblySetNZCVBits);
}
// Special case of the above where we are known to zero C/V
void HandleNZ00Write() {
HandleNZCVWrite((1u << 31) | (1u << 30));
}
OrderedNode *GetNZCV() {
if (!CachedNZCV) {
CachedNZCV = _LoadNZCV();
@@ -1601,11 +1630,11 @@ private:
OrderedNode *Res{};
union {
// UMUL, BEXTR, BLSI, BLSMSK, POPCOUNT, TZCNT, LZCNT, RDRAND
// UMUL, BEXTR, BLSI, POPCOUNT, ZCNT, RDRAND
struct {
} NoSource;
// MUL, BLSR, BZHI
// MUL, BLSR, BLSMSKB, BZHI
struct {
OrderedNode *Src1;
} OneSource;
@@ -1616,13 +1645,6 @@ private:
OrderedNode *Src2;
} TwoSource;
// ADC, SBB
struct {
OrderedNode *Src1;
OrderedNode *Src2;
OrderedNode *Src3;
} ThreeSource;
// LSHLI, LSHRI, ASHRI, RORI, ROLI
struct {
OrderedNode *Src1;
@@ -1713,13 +1735,13 @@ private:
OrderedNode *LoadAF();
void FixupAF();
void CalculatePF(OrderedNode *Res);
void CalculateAF(OpSize OpSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
void CalculateAF(OrderedNode *Src1, OrderedNode *Src2);
void CalculateOF(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool Sub);
void CalculateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF);
void CalculateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
OrderedNode *CalculateFlags_ADC(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2);
OrderedNode *CalculateFlags_SBB(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2);
OrderedNode *CalculateFlags_SUB(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
OrderedNode *CalculateFlags_ADD(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true);
void CalculateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High);
void CalculateFlags_UMUL(OrderedNode *High);
void CalculateFlags_Logical(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2);
@@ -1737,12 +1759,11 @@ private:
void CalculateFlags_RotateLeftImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift);
void CalculateFlags_BEXTR(OrderedNode *Src);
void CalculateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src);
void CalculateFlags_BLSMSK(OrderedNode *Src);
void CalculateFlags_BLSMSK(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src);
void CalculateFlags_BLSR(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src);
void CalculateFlags_POPCOUNT(OrderedNode *Src);
void CalculateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src);
void CalculateFlags_TZCNT(OrderedNode *Src);
void CalculateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src);
void CalculateFlags_ZCNT(uint8_t SrcSize, OrderedNode *Result);
void CalculateFlags_RDRAND(OrderedNode *Src);
/** @} */
@@ -1751,37 +1772,7 @@ private:
*
* Depending on the operation it may force a RFLAGs calculation before storing the new deferred state.
* @{ */
void GenerateFlags_ADC(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADC,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SBB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SBB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.ThreeSource = {
.Src1 = Src1,
.Src2 = Src2,
.Src3 = CF,
},
},
};
}
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
void GenerateFlags_SUB(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
@@ -1789,26 +1780,6 @@ private:
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_SUB,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
.Src2 = Src2,
.UpdateCF = UpdateCF,
},
},
};
}
void GenerateFlags_ADD(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF = true) {
if (!UpdateCF) {
// If we aren't updating CF then we need to calculate flags. Invalidation mask would make this not required.
CalculateDeferredFlags();
}
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_ADD,
.SrcSize = GetSrcSize(Op),
.Res = Res,
.Sources = {
.TwoSrcImmediate = {
.Src1 = Src1,
@@ -2061,11 +2032,16 @@ private:
};
}
void GenerateFlags_BLSMSK(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
void GenerateFlags_BLSMSK(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Res, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_BLSMSK,
.SrcSize = GetSrcSize(Op),
.Res = Src,
.Res = Res,
.Sources = {
.OneSource = {
.Src1 = Src,
},
},
};
}
@@ -2103,17 +2079,9 @@ private:
};
}
void GenerateFlags_TZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
void GenerateFlags_ZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_TZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
}
void GenerateFlags_LZCNT(FEXCore::X86Tables::DecodedOp Op, OrderedNode *Src) {
CurrentDeferredFlags = DeferredFlagData {
.Type = FlagsGenerationType::TYPE_LZCNT,
.Type = FlagsGenerationType::TYPE_ZCNT,
.SrcSize = GetSrcSize(Op),
.Res = Src,
};
@@ -2127,6 +2095,16 @@ private:
};
}
OrderedNode *AndConst(FEXCore::IR::OpSize Size, OrderedNode *Node, uint64_t Const) {
uint64_t NodeConst;
if (IsValueConstant(WrapNode(Node), &NodeConst)) {
return _Constant(NodeConst & Const);
} else {
return _And(Size, Node, _Constant(Const));
}
}
/** @} */
/** @} */
@@ -8,7 +8,6 @@ $end_info$
#include "Interface/Core/X86Tables/X86Tables.h"
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include "Interface/Core/OpcodeDispatcher.h"
@@ -104,7 +103,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
const auto f2 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._And(OpSize::i32Bit, B, D)), Self._And(OpSize::i32Bit, C, D));
return Self.BitwiseAtLeastTwo(B, C, D);
};
const auto f3 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
@@ -129,9 +128,6 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
auto W0E = _VExtractToGPR(16, 4, Src, 3);
auto W1 = _VExtractToGPR(16, 4, Src, 2);
auto W2 = _VExtractToGPR(16, 4, Src, 1);
auto W3 = _VExtractToGPR(16, 4, Src, 0);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
@@ -150,8 +146,12 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C,
OrderedNode *D, OrderedNode *E, OrderedNode *W) -> RoundResult {
auto ANext = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), W), E), K);
OrderedNode *D, OrderedNode *E, OrderedNode *Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(16, 4, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
auto ANext = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(32, 2));
auto DNext = C;
@@ -161,9 +161,9 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
};
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, W1);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, W2);
auto Final = Round1To3(A3, B3, C3, D3, E3, W3);
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(16, 4, 2, Dest3, std::get<1>(Final));
@@ -230,18 +230,25 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
StoreResult(FPRClass, Op, D0, -1);
}
OrderedNode *OpDispatchBuilder::BitwiseAtLeastTwo(OrderedNode *A, OrderedNode *B, OrderedNode *C) {
// Returns whether at least 2/3 of A/B/C is true.
// Expressed as (A & (B | C)) | (B & C)
//
// Equivalent to expression in SHA calculations: (A & B) ^ (A & C) ^ (B & C)
auto And = _And(OpSize::i32Bit, B, C);
auto Or = _Or(OpSize::i32Bit, B, C);
return _Or(OpSize::i32Bit, _And(OpSize::i32Bit, A, Or), And);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
const auto Ch = [this](OrderedNode *E, OrderedNode *F, OrderedNode *G) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
};
const auto Major = [this](OrderedNode *A, OrderedNode *B, OrderedNode *C) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, A, B), _And(OpSize::i32Bit, A, C)), _And(OpSize::i32Bit, B, C));
};
const auto Sigma0 = [this](OrderedNode *A) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), _Ror(OpSize::i32Bit, A, _Constant(32, 13))), _Ror(OpSize::i32Bit, A, _Constant(32, 22)));
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), A, ShiftType::ROR, 13), A, ShiftType::ROR, 22);
};
const auto Sigma1 = [this](OrderedNode *E) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), _Ror(OpSize::i32Bit, E, _Constant(32, 11))), _Ror(OpSize::i32Bit, E, _Constant(32, 25)));
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), E, ShiftType::ROR, 11), E, ShiftType::ROR, 25);
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
@@ -249,42 +256,44 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
auto C0 = _VExtractToGPR(16, 4, Dest, 3);
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto E0 = _VExtractToGPR(16, 4, Src, 1);
auto F0 = _VExtractToGPR(16, 4, Src, 0);
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
auto H0 = _VExtractToGPR(16, 4, Dest, 0);
OrderedNode *Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
auto H0 = _VExtractToGPR(16, 4, Dest, 0);
Q0 = _Add(OpSize::i32Bit, Q0, H0);
auto A0 = _VExtractToGPR(16, 4, Src, 3);
auto B0 = _VExtractToGPR(16, 4, Src, 2);
auto C0 = _VExtractToGPR(16, 4, Dest, 3);
auto A1 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q0, BitwiseAtLeastTwo(A0, B0, C0)), Sigma0(A0));
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
OrderedNode * Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*,
OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
const auto Round = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C, OrderedNode *D,
OrderedNode *E, OrderedNode *F, OrderedNode *G, OrderedNode *H,
OrderedNode* WK) -> RoundResult {
auto ANext = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Ch(E, F, G), Sigma1(E)), WK), H), Major(A, B, C)), Sigma0(A));
auto BNext = A;
auto CNext = B;
auto DNext = C;
auto ENext = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Ch(E, F, G), Sigma1(E)), WK), H), D);
auto FNext = E;
auto GNext = F;
auto HNext = G;
// Rematerialize G0. Costs a move but saves spilling, coming out ahead.
G0 = _VExtractToGPR(16, 4, Dest, 1);
Q1 = _Add(OpSize::i32Bit, Q1, G0);
return {ANext, BNext, CNext, DNext, ENext, FNext, GNext, HNext};
};
auto A2 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Q1, BitwiseAtLeastTwo(A1, A0, B0)), Sigma0(A1));
// Rematerialize C0. As with G0.
C0 = _VExtractToGPR(16, 4, Dest, 3);
auto E2 = _Add(OpSize::i32Bit, Q1, C0);
auto [A1, B1, C1, D1, E1, F1, G1, H1] = Round(A0, B0, C0, D0, E0, F0, G0, H0, WK0);
auto Final = Round(A1, B1, C1, D1, E1, F1, G1, H1, WK1);
auto Res3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Res2 = _VInsGPR(16, 4, 2, Res3, std::get<1>(Final));
auto Res1 = _VInsGPR(16, 4, 1, Res2, std::get<4>(Final));
auto Res0 = _VInsGPR(16, 4, 0, Res1, std::get<5>(Final));
auto Res3 = _VInsGPR(16, 4, 3, Dest, A2);
auto Res2 = _VInsGPR(16, 4, 2, Res3, A1);
auto Res1 = _VInsGPR(16, 4, 1, Res2, E2);
auto Res0 = _VInsGPR(16, 4, 0, Res1, E1);
StoreResult(FPRClass, Op, Res0, -1);
}
@@ -282,19 +282,25 @@ void OpDispatchBuilder::CalculatePF(OrderedNode *Res) {
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(Res);
}
void OpDispatchBuilder::CalculateAF(OpSize OpSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2) {
void OpDispatchBuilder::CalculateAF(OrderedNode *Src1, OrderedNode *Src2) {
// We only care about bit 4 in the subsequent XOR. If we'll XOR with 0,
// there's no sense XOR'ing at all. This affects INC.
// there's no sense XOR'ing at all. If we'll XOR with 1, that's just
// inverting.
uint64_t Const;
if (IsValueConstant(WrapNode(Src2), &Const) && (Const & (1u << 4)) == 0) {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Src1);
if (IsValueConstant(WrapNode(Src2), &Const)) {
if (Const & (1u << 4)) {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(_Not(OpSize::i32Bit, Src1));
} else {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(Src1);
}
return;
}
// We store the XOR of the arguments. At read time, we XOR with the
// appropriate bit of the result (available as the PF flag) and extract the
// appropriate bit.
OrderedNode *XorRes = _Xor(OpSize, Src1, Src2);
OrderedNode *XorRes = _Xor(OpSize::i32Bit, Src1, Src2);
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
@@ -310,34 +316,9 @@ void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
}
switch (CurrentDeferredFlags.Type) {
case FlagsGenerationType::TYPE_ADC:
CalculateFlags_ADC(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SBB:
CalculateFlags_SBB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.ThreeSource.Src1,
CurrentDeferredFlags.Sources.ThreeSource.Src2,
CurrentDeferredFlags.Sources.ThreeSource.Src3);
break;
case FlagsGenerationType::TYPE_SUB:
CalculateFlags_SUB(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_ADD:
CalculateFlags_ADD(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2,
CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
@@ -444,7 +425,10 @@ void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_BLSMSK:
CalculateFlags_BLSMSK(CurrentDeferredFlags.Res);
CalculateFlags_BLSMSK(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_BLSR:
CalculateFlags_BLSR(
@@ -461,11 +445,8 @@ void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_TZCNT:
CalculateFlags_TZCNT(CurrentDeferredFlags.Res);
break;
case FlagsGenerationType::TYPE_LZCNT:
CalculateFlags_LZCNT(
case FlagsGenerationType::TYPE_ZCNT:
CalculateFlags_ZCNT(
CurrentDeferredFlags.SrcSize,
CurrentDeferredFlags.Res);
break;
@@ -486,150 +467,126 @@ void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
NZCVDirty = false;
}
void OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
OrderedNode *OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2) {
auto Zero = _Constant(0);
auto One = _Constant(1);
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
OrderedNode *Res;
CalculateAF(OpSize, Res, Src1, Src2);
CalculatePF(Res);
CalculateAF(Src1, Src2);
if (SrcSize >= 4) {
if (NZCVDirty && CachedNZCV)
_StoreNZCV(CachedNZCV);
CachedNZCV = nullptr;
_AdcNZCV(OpSize, Src1, Src2);
PossiblySetNZCVBits = ~0;
HandleNZCV_RMW();
Res = _AdcWithFlags(OpSize, Src1, Src2);
} else {
// SF/ZF
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
Res = _Adc(OpSize, Src1, Src2);
Res = _Bfe(OpSize, SrcSize * 8, 0, Res);
auto SelectOpLT = _Select(FEXCore::IR::COND_ULT, Res, Src2, One, Zero);
auto SelectOpLE = _Select(FEXCore::IR::COND_ULE, Res, Src2, One, Zero);
auto SelectCF = _Select(FEXCore::IR::COND_EQ, CF, One, SelectOpLE, SelectOpLT);
SetNZ_ZeroCV(SrcSize, Res);
// CF
// Unsigned
{
auto SelectOpLT = _Select(FEXCore::IR::COND_ULT, Res, Src2, One, Zero);
auto SelectOpLE = _Select(FEXCore::IR::COND_ULE, Res, Src2, One, Zero);
auto SelectCF = _Select(FEXCore::IR::COND_EQ, CF, One, SelectOpLE, SelectOpLT);
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(SelectCF);
}
// Signed
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(SelectCF);
CalculateOF(SrcSize, Res, Src1, Src2, false);
}
CalculatePF(Res);
return Res;
}
void OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, OrderedNode *CF) {
OrderedNode *OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2) {
auto Zero = _Constant(0);
auto One = _Constant(1);
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(OpSize, Res, Src1, Src2);
CalculatePF(Res);
CalculateAF(Src1, Src2);
OrderedNode *Res;
if (SrcSize >= 4) {
// Rectify input carry
CarryInvert();
if (NZCVDirty && CachedNZCV)
_StoreNZCV(CachedNZCV);
CachedNZCV = nullptr;
NZCVDirty = false;
_SbbNZCV(OpSize, Src1, Src2);
PossiblySetNZCVBits = ~0;
HandleNZCV_RMW();
Res = _SbbWithFlags(OpSize, Src1, Src2);
// Rectify output carry
CarryInvert();
} else {
// SF/ZF
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
Res = _Sub(OpSize, Src1, _Add(OpSize, Src2, CF));
Res = _Bfe(OpSize, SrcSize * 8, 0, Res);
auto SelectOpLT = _Select(FEXCore::IR::COND_UGT, Res, Src1, One, Zero);
auto SelectOpLE = _Select(FEXCore::IR::COND_UGE, Res, Src1, One, Zero);
auto SelectCF = _Select(FEXCore::IR::COND_EQ, CF, One, SelectOpLE, SelectOpLT);
SetNZ_ZeroCV(SrcSize, Res);
// CF
// Unsigned
{
auto SelectOpLT = _Select(FEXCore::IR::COND_UGT, Res, Src1, One, Zero);
auto SelectOpLE = _Select(FEXCore::IR::COND_UGE, Res, Src1, One, Zero);
auto SelectCF = _Select(FEXCore::IR::COND_EQ, CF, One, SelectOpLE, SelectOpLT);
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(SelectCF);
}
// Signed
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(SelectCF);
CalculateOF(SrcSize, Res, Src1, Src2, true);
}
CalculatePF(Res);
return Res;
}
void OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(OpSize, Res, Src1, Src2);
CalculatePF(Res);
OrderedNode *OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// Stash CF before stomping over it
auto OldCF = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
// TODO: Could do this path for small sources if we have FEAT_FlagM
HandleNZCVWrite();
CalculateAF(Src1, Src2);
OrderedNode *Res;
if (SrcSize >= 4) {
_SubNZCV(OpSize, Src1, Src2);
CachedNZCV = nullptr;
NZCVDirty = false;
PossiblySetNZCVBits = ~0;
// We only bother inverting CF if we're actually going to update CF.
if (UpdateCF)
CarryInvert();
Res = _SubWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
} else {
// SF/ZF
SetNZ_ZeroCV(SrcSize, Res);
// CF
if (UpdateCF) {
// Grab carry bit from unmasked output.
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Res, SrcSize * 8, true);
}
CalculateOF(SrcSize, Res, Src1, Src2, true);
_SubNZCV(IR::SizeToOpSize(SrcSize), Src1, Src2);
Res = _Sub(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
// If we're updating CF, we need to invert it for correctness. If we're not
// updating CF, we need to restore the CF since we stomped over it.
if (UpdateCF)
CarryInvert();
else
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(OldCF);
return Res;
}
OrderedNode *OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
// Stash CF before stomping over it
auto OldCF = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
HandleNZCVWrite();
CalculateAF(Src1, Src2);
OrderedNode *Res;
if (SrcSize >= 4) {
Res = _AddWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
} else {
_AddNZCV(IR::SizeToOpSize(SrcSize), Src1, Src2);
Res = _Add(OpSize::i32Bit, Src1, Src2);
}
CalculatePF(Res);
// We stomped over CF while calculation flags, restore it.
if (!UpdateCF)
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(OldCF);
}
void OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, OrderedNode *Src2, bool UpdateCF) {
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
CalculateAF(OpSize, Res, Src1, Src2);
CalculatePF(Res);
// Stash CF before stomping over it
auto OldCF = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
// TODO: Could do this path for small sources if we have FEAT_FlagM
if (SrcSize >= 4) {
_AddNZCV(OpSize, Src1, Src2);
CachedNZCV = nullptr;
NZCVDirty = false;
PossiblySetNZCVBits = ~0;
} else {
// SF/ZF
SetNZ_ZeroCV(SrcSize, Res);
// CF
if (UpdateCF) {
// Grab carry bit from unmasked output
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Res, SrcSize * 8, true);
}
CalculateOF(SrcSize, Res, Src1, Src2, false);
}
// We stomped over CF while calculation flags, restore it.
if (!UpdateCF)
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(OldCF);
return Res;
}
void OpDispatchBuilder::CalculateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, OrderedNode *High) {
HandleNZCVWrite();
// PF/AF/ZF/SF
// Undefined
{
@@ -648,13 +605,12 @@ void OpDispatchBuilder::CalculateFlags_MUL(uint8_t SrcSize, OrderedNode *Res, Or
// undefined, this does what we need.
auto Zero = _Constant(0);
_CondAddNZCV(OpSize::i64Bit, Zero, Zero, CondClassType{COND_EQ}, 0x3 /* nzCV */);
CachedNZCV = nullptr;
NZCVDirty = false;
PossiblySetNZCVBits = ~0;
}
}
void OpDispatchBuilder::CalculateFlags_UMUL(OrderedNode *High) {
HandleNZCVWrite();
auto Zero = _Constant(0);
OpSize Size = IR::SizeToOpSize(GetOpSize(High));
@@ -674,9 +630,6 @@ void OpDispatchBuilder::CalculateFlags_UMUL(OrderedNode *High) {
// If High = 0, then sets to nZcv. Else sets to nzCV. Since SF/ZF undefined,
// this does what we need.
_CondAddNZCV(Size, Zero, Zero, CondClassType{COND_EQ}, 0x3 /* nzCV */);
CachedNZCV = nullptr;
NZCVDirty = false;
PossiblySetNZCVBits = ~0;
}
}
@@ -817,8 +770,8 @@ void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize,
}
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
// Stash OF before overwriting it
auto OldOF = Shift != 1 ? GetRFLAG(FEXCore::X86State::RFLAG_OF_RAW_LOC) : NULL;
// Set SF and PF. Clobbers OF, but OF only defined for Shift = 1 where it is
// set below.
SetNZ_ZeroCV(SrcSize, Res);
// CF
@@ -832,11 +785,6 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize
// AF
// Undefined
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
// Preserve OF if it won't be written
if (Shift != 1) {
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(OldOF);
}
}
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode *Res, OrderedNode *Src1, uint64_t Shift) {
@@ -976,112 +924,67 @@ void OpDispatchBuilder::CalculateFlags_RotateLeftImmediate(uint8_t SrcSize, Orde
}
void OpDispatchBuilder::CalculateFlags_BEXTR(OrderedNode *Src) {
auto Zero = _Constant(0);
auto One = _Constant(1);
// ZF is set properly. CF and OF are defined as being set to zero. SF, PF, and
// AF are undefined.
SetNZ_ZeroCV(GetOpSize(Src), Src);
// Handle flag setting.
//
// All that matters primarily for this instruction is
// that we only set the ZF flag properly.
//
// CF and OF are defined as being set to zero
//
// Every other flag is considered undefined after a
// BEXTR instruction, but we opt to reliably clear them.
//
ZeroMultipleFlags(FullNZCVMask);
// PF/AF undefined
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) |
(1UL << X86State::RFLAG_AF_RAW_LOC));
// ZF
auto ZeroOp = _Select(IR::COND_EQ,
Src, Zero,
One, Zero);
SetRFLAG<X86State::RFLAG_ZF_RAW_LOC>(ZeroOp);
}
void OpDispatchBuilder::CalculateFlags_BLSI(uint8_t SrcSize, OrderedNode *Src) {
// Now for the flags:
void OpDispatchBuilder::CalculateFlags_BLSI(uint8_t SrcSize, OrderedNode *Result) {
// CF is cleared if Src is zero, otherwise it's set. However, Src is zero iff
// Result is zero, so we can test the result instead. So, CF is just the
// inverted ZF.
//
// Only CF, SF, ZF and OF are defined as being updated
// CF is cleared if Src is zero, otherwise it's set.
// SF is set to the value of the most significant operand bit of Result.
// OF is always cleared
// ZF is set, as usual, if Result is zero or not.
// ZF/SF/OF set as usual.
SetNZ_ZeroCV(SrcSize, Result);
auto Zero = _Constant(0);
auto One = _Constant(1);
SetNZ_ZeroCV(SrcSize, Src);
auto CFOp = GetRFLAG(X86State::RFLAG_ZF_RAW_LOC, true /* Invert */);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
// PF/AF undefined
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) |
(1UL << X86State::RFLAG_AF_RAW_LOC));
// CF
{
auto CFOp = _Select(IR::COND_NEQ, Src, Zero, One, Zero);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
}
}
void OpDispatchBuilder::CalculateFlags_BLSMSK(OrderedNode *Src) {
// Now for the flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
uint32_t FlagsMaskToZero =
(1U << X86State::RFLAG_ZF_RAW_LOC) |
(1U << X86State::RFLAG_OF_RAW_LOC);
ZeroMultipleFlags(FlagsMaskToZero);
void OpDispatchBuilder::CalculateFlags_BLSMSK(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// PF/AF undefined
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) |
(1UL << X86State::RFLAG_AF_RAW_LOC));
auto CFOp = _Select(IR::COND_NEQ, Src, Zero, One, Zero);
// CF set according to the Src
auto Zero = _Constant(0);
auto One = _Constant(1);
auto CFOp = _Select(IR::COND_EQ, Src, Zero, One, Zero);
// The output of BLSMSK is always nonzero, so TST will clear Z (along with C
// and O) while setting S.
SetNZ_ZeroCV(SrcSize, Result);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
}
void OpDispatchBuilder::CalculateFlags_BLSR(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
// Now for flags.
auto Zero = _Constant(0);
auto One = _Constant(1);
auto CFOp = _Select(IR::COND_EQ, Src, Zero, One, Zero);
SetNZ_ZeroCV(SrcSize, Result);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
// PF/AF undefined
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) |
(1UL << X86State::RFLAG_AF_RAW_LOC));
// CF
{
auto CFOp = _Select(IR::COND_NEQ, Src, Zero, One, Zero);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
}
}
void OpDispatchBuilder::CalculateFlags_POPCOUNT(OrderedNode *Src) {
// Set ZF
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
void OpDispatchBuilder::CalculateFlags_POPCOUNT(OrderedNode *Result) {
// We need to set ZF while clearing the rest of NZCV. The result of a popcount
// is in the range [0, 63]. In particular, it is always positive. So a
// combined NZ test will correctly zero SF/CF/OF while setting ZF.
SetNZ_ZeroCV(OpSize::i32Bit, Result);
// Set flags
uint32_t FlagsMaskToZero =
FullNZCVMask |
(1U << X86State::RFLAG_AF_RAW_LOC) |
(1U << X86State::RFLAG_PF_RAW_LOC);
ZeroMultipleFlags(FlagsMaskToZero);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(ZFResult);
ZeroMultipleFlags((1U << X86State::RFLAG_AF_RAW_LOC) |
(1U << X86State::RFLAG_PF_RAW_LOC));
}
void OpDispatchBuilder::CalculateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result, OrderedNode *Src) {
@@ -1093,32 +996,16 @@ void OpDispatchBuilder::CalculateFlags_BZHI(uint8_t SrcSize, OrderedNode *Result
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(Src);
}
void OpDispatchBuilder::CalculateFlags_TZCNT(OrderedNode *Src) {
void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, OrderedNode *Result) {
// OF, SF, AF, PF all undefined
ZeroNZCV();
// Test ZF of result, SF is undefined so this is ok.
SetNZ_ZeroCV(SrcSize, Result);
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(Src, 0, true);
}
void OpDispatchBuilder::CalculateFlags_LZCNT(uint8_t SrcSize, OrderedNode *Src) {
// OF, SF, AF, PF all undefined
ZeroNZCV();
auto Zero = _Constant(0);
auto ZFResult = _Select(FEXCore::IR::COND_EQ,
Src, Zero,
_Constant(1), Zero);
// Set flags
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(ZFResult);
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(Src, SrcSize * 8 - 1, true);
// Now set CF if the Result = SrcSize * 8. Since SrcSize is a power-of-two and
// Result is <= SrcSize * 8, we equivalently check if the log2(SrcSize * 8)
// bit is set. No masking is needed because no higher bits could be set.
unsigned CarryBit = FEXCore::ilog2(SrcSize * 8u);
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Result, CarryBit);
}
void OpDispatchBuilder::CalculateFlags_RDRAND(OrderedNode *Src) {
@@ -3422,15 +3422,12 @@ void OpDispatchBuilder::VPALIGNROp(OpcodeArgs) {
template<size_t ElementSize>
void OpDispatchBuilder::UCOMISxOp(OpcodeArgs) {
InvalidateDeferredFlags();
const auto SrcSize = Op->Src[0].IsGPR() ? GetGuestVectorLength() : GetSrcSize(Op);
OrderedNode *Src1 = LoadSource_WithOpSize(FPRClass, Op, Op->Dest, GetGuestVectorLength(), Op->Flags);
OrderedNode *Src2 = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], SrcSize, Op->Flags);
CachedNZCV = nullptr;
HandleNZCVWrite();
_FCmp(ElementSize, Src1, Src2);
PossiblySetNZCVBits = ~0;
ConvertNZCVToSSE();
// Zero AF. Note that the comparison sets the raw PF to 0/1 above, so PF[4] is
@@ -4580,13 +4577,9 @@ void OpDispatchBuilder::PTestOp(OpcodeArgs) {
OrderedNode *Test1 = _VAnd(Size, 1, Dest, Src);
OrderedNode *Test2 = _VBic(Size, 1, Src, Dest);
Test1 = _VPopcount(Size, 1, Test1);
Test2 = _VPopcount(Size, 1, Test2);
// Element size doesn't matter here
// x86-64 doesn't support a horizontal byte add though
Test1 = _VAddV(Size, 2, Test1);
Test2 = _VAddV(Size, 2, Test2);
// Element size must be less than 32-bit for the sign bit tricks.
Test1 = _VUMaxV(Size, 2, Test1);
Test2 = _VUMaxV(Size, 2, Test2);
Test1 = _VExtractToGPR(Size, 2, Test1, 0);
Test2 = _VExtractToGPR(Size, 2, Test2, 0);
@@ -4594,15 +4587,14 @@ void OpDispatchBuilder::PTestOp(OpcodeArgs) {
auto ZeroConst = _Constant(0);
auto OneConst = _Constant(1);
Test1 = _Select(FEXCore::IR::COND_EQ,
Test1, ZeroConst, OneConst, ZeroConst);
Test2 = _Select(FEXCore::IR::COND_EQ,
Test2, ZeroConst, OneConst, ZeroConst);
// Careful, these flags are different between {V,}PTEST and VTESTP{S,D}
ZeroNZCV();
SetRFLAG<FEXCore::X86State::RFLAG_ZF_RAW_LOC>(Test1);
// Set ZF according to Test1. SF will be zeroed since we do a 32-bit test on
// the results of a 16-bit value from the UMaxV, so the 32-bit sign bit is
// cleared even if the 16-bit scalars were negative.
SetNZ_ZeroCV(32, Test1);
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Test2);
uint32_t FlagsMaskToZero =
@@ -4630,33 +4622,24 @@ void OpDispatchBuilder::VTESTOpImpl(OpcodeArgs, size_t ElementSize) {
OrderedNode *MaskedAnd = _VAnd(SrcSize, 1, AndTest, Mask);
OrderedNode *MaskedAndNot = _VAnd(SrcSize, 1, AndNotTest, Mask);
OrderedNode *AndPopCount = _VPopcount(SrcSize, 1, MaskedAnd);
OrderedNode *AndNotPopCount = _VPopcount(SrcSize, 1, MaskedAndNot);
OrderedNode *MaxAnd = _VUMaxV(SrcSize, 2, MaskedAnd);
OrderedNode *MaxAndNot = _VUMaxV(SrcSize, 2, MaskedAndNot);
OrderedNode *SummedAnd = _VAddV(SrcSize, 2, AndPopCount);
OrderedNode *SummedAndNot = _VAddV(SrcSize, 2, AndNotPopCount);
OrderedNode *AndGPR = _VExtractToGPR(SrcSize, 2, SummedAnd, 0);
OrderedNode *AndNotGPR = _VExtractToGPR(SrcSize, 2, SummedAndNot, 0);
OrderedNode *AndGPR = _VExtractToGPR(SrcSize, 2, MaxAnd, 0);
OrderedNode *AndNotGPR = _VExtractToGPR(SrcSize, 2, MaxAndNot, 0);
OrderedNode *ZeroConst = _Constant(0);
OrderedNode *OneConst = _Constant(1);
OrderedNode *ZFResult = _Select(IR::COND_EQ, AndGPR, ZeroConst,
OneConst, ZeroConst);
OrderedNode *CFResult = _Select(IR::COND_EQ, AndNotGPR, ZeroConst,
OneConst, ZeroConst);
SetRFLAG<X86State::RFLAG_ZF_RAW_LOC>(ZFResult);
// As in PTest, this sets Z appropriately while zeroing the rest of NZCV.
SetNZ_ZeroCV(32, AndGPR);
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFResult);
uint32_t FlagsMaskToZero =
(1U << X86State::RFLAG_PF_RAW_LOC) |
(1U << X86State::RFLAG_AF_RAW_LOC) |
(1U << X86State::RFLAG_SF_RAW_LOC) |
(1U << X86State::RFLAG_OF_RAW_LOC);
ZeroMultipleFlags(FlagsMaskToZero);
ZeroMultipleFlags((1U << X86State::RFLAG_PF_RAW_LOC) |
(1U << X86State::RFLAG_AF_RAW_LOC));
}
template <size_t ElementSize>
@@ -14,7 +14,6 @@ $end_info$
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/FPState.h>
#include <FEXCore/IR/IREmitter.h>
#include <stddef.h>
#include <stdint.h>
@@ -13,7 +13,6 @@ $end_info$
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IREmitter.h>
#include <stddef.h>
#include <stdint.h>
@@ -12,51 +12,16 @@ $end_info$
namespace FEXCore::X86Tables {
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps{};
std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps{};
std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps{};
std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps{};
std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps{};
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps{};
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps{};
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps{};
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops{};
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps{};
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps{};
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps{};
std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps{};
std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps{};
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps{};
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps{};
std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> EVEXTableOps{};
void InitializeBaseTables(Context::OperatingMode Mode);
void InitializeSecondaryTables(Context::OperatingMode Mode);
void InitializePrimaryGroupTables(Context::OperatingMode Mode);
void InitializeSecondaryGroupTables();
void InitializeSecondaryModRMTables();
void InitializeX87Tables();
void InitializeDDDTables();
void InitializeH0F38Tables();
void InitializeH0F3ATables(Context::OperatingMode Mode);
void InitializeVEXTables();
void InitializeXOPTables();
void InitializeEVEXTables();
void InitializeInfoTables(Context::OperatingMode Mode) {
InitializeBaseTables(Mode);
InitializeSecondaryTables(Mode);
InitializePrimaryGroupTables(Mode);
InitializeSecondaryGroupTables();
InitializeSecondaryModRMTables();
InitializeX87Tables();
InitializeDDDTables();
InitializeH0F38Tables();
InitializeH0F3ATables(Mode);
InitializeVEXTables();
InitializeXOPTables();
InitializeEVEXTables();
}
}
@@ -14,8 +14,10 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeBaseTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct BaseOpTable[] = {
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct BaseOpTable[] = {
// Prefixes
// Operand size overide
{0x66, 1, X86InstInfo{"", TYPE_PREFIX, FLAGS_NONE, 0, nullptr}},
@@ -233,6 +235,12 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xC4, 2, X86InstInfo{"", TYPE_VEX_TABLE_PREFIX, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(&Table.at(0), BaseOpTable, std::size(BaseOpTable));
return Table;
}();
void InitializeBaseTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct BaseOpTable_64[] = {
{0x06, 2, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x0E, 1, X86InstInfo{"[INV]", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -291,8 +299,6 @@ void InitializeBaseTables(Context::OperatingMode Mode) {
{0xEA, 1, X86InstInfo{"JMPF", TYPE_INST, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(&BaseOps.at(0), BaseOpTable, std::size(BaseOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(&BaseOps.at(0), BaseOpTable_64, std::size(BaseOpTable_64));
}
@@ -12,8 +12,9 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeDDDTables() {
static constexpr U8U8InfoStruct DDDNowOpTable[] = {
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> DDDNowOps = []() consteval {
std::array<X86InstInfo, MAX_3DNOW_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct DDDNowOpTable[] = {
{0x0C, 1, X86InstInfo{"PI2FW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x0D, 1, X86InstInfo{"PI2FD", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{0x1C, 1, X86InstInfo{"PF2IW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -52,6 +53,8 @@ void InitializeDDDTables() {
{0xBF, 1, X86InstInfo{"PAVGUSB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
};
GenerateTable(&DDDNowOps.at(0), DDDNowOpTable, std::size(DDDNowOpTable));
}
GenerateTable(&Table.at(0), DDDNowOpTable, std::size(DDDNowOpTable));
return Table;
}();
}
@@ -11,9 +11,9 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeEVEXTables() {
static constexpr U16U8InfoStruct EVEXTable[] = {
std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> EVEXTableOps = []() consteval {
std::array<X86InstInfo, MAX_EVEX_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct EVEXTable[] = {
{0x10, 1, X86InstInfo{"VMOVUPS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x11, 1, X86InstInfo{"VMOVUPS", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
{0x18, 1, X86InstInfo{"VBROADCASTSS", TYPE_INST, FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -29,6 +29,9 @@ void InitializeEVEXTables() {
{0xE7, 1, X86InstInfo{"VMOVNTDQ", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
};
GenerateTable(&EVEXTableOps.at(0), EVEXTable, std::size(EVEXTable));
}
GenerateTable(&Table.at(0), EVEXTable, std::size(EVEXTable));
return Table;
}();
}
@@ -12,15 +12,16 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> H0F38TableOps = []() consteval {
std::array<X86InstInfo, MAX_0F_38_TABLE_SIZE> Table{};
void InitializeH0F38Tables() {
#define OPD(prefix, opcode) (((prefix) << 8) | opcode)
constexpr uint16_t PF_38_NONE = 0;
constexpr uint16_t PF_38_66 = (1U << 0);
constexpr uint16_t PF_38_F2 = (1U << 1);
constexpr uint16_t PF_38_F3 = (1U << 2);
static constexpr U16U8InfoStruct H0F38Table[] = {
constexpr U16U8InfoStruct H0F38Table[] = {
{OPD(PF_38_NONE, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
{OPD(PF_38_66, 0x00), 1, X86InstInfo{"PSHUFB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
{OPD(PF_38_NONE, 0x01), 1, X86InstInfo{"PHADDW", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 0, nullptr}},
@@ -117,6 +118,8 @@ void InitializeH0F38Tables() {
};
#undef OPD
GenerateTable(&H0F38TableOps.at(0), H0F38Table, std::size(H0F38Table));
}
GenerateTable(&Table.at(0), H0F38Table, std::size(H0F38Table));
return Table;
}();
}
@@ -14,13 +14,13 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeH0F3ATables(Context::OperatingMode Mode) {
#define OPD(REX, prefix, opcode) ((REX << 9) | (prefix << 8) | opcode)
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = 1;
constexpr uint16_t PF_3A_NONE = 0;
constexpr uint16_t PF_3A_66 = 1;
static constexpr U16U8InfoStruct H0F3ATable[] = {
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> H0F3ATableOps = []() consteval {
std::array<X86InstInfo, MAX_0F_3A_TABLE_SIZE> Table{};
constexpr U16U8InfoStruct H0F3ATable[] = {
{OPD(0, PF_3A_NONE, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_SF_MMX, 1, nullptr}},
{OPD(0, PF_3A_66, 0x08), 1, X86InstInfo{"ROUNDPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(0, PF_3A_66, 0x09), 1, X86InstInfo{"ROUNDPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -54,6 +54,11 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
{OPD(0, PF_3A_66, 0xDF), 1, X86InstInfo{"AESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
};
GenerateTable(&Table.at(0), H0F3ATable, std::size(H0F3ATable));
return Table;
}();
void InitializeH0F3ATables(Context::OperatingMode Mode) {
static constexpr U16U8InfoStruct H0F3ATable_64[] = {
{OPD(1, PF_3A_66, 0x0F), 1, X86InstInfo{"PALIGNR", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(1, PF_3A_66, 0x16), 1, X86InstInfo{"PEXTRQ", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_DST_GPR | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -62,8 +67,6 @@ void InitializeH0F3ATables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(&H0F3ATableOps.at(0), H0F3ATable, std::size(H0F3ATable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(&H0F3ATableOps.at(0), H0F3ATable_64, std::size(H0F3ATable_64));
}
@@ -13,10 +13,10 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> PrimaryInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_GROUP_TABLE_SIZE> Table{};
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_1) << 6) | (prefix) << 3 | (Reg))
const U16U8InfoStruct PrimaryGroupOpTable[] = {
constexpr U16U8InfoStruct PrimaryGroupOpTable[] = {
// GROUP_1 | 0x80 | reg
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 0), 1, X86InstInfo{"ADD", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
{OPD(TYPE_GROUP_1, OpToIndex(0x80), 1), 1, X86InstInfo{"OR", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST, 1, nullptr}},
@@ -141,9 +141,13 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 0), 1, X86InstInfo{"MOV", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 1), 5, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_11, OpToIndex(0xC7), 7), 1, X86InstInfo{"XBEGIN", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT | FLAGS_SETS_RIP | FLAGS_DISPLACE_SIZE_DIV_2, 4, nullptr}},
};
GenerateTable(&Table.at(0), PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
return Table;
}();
void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
const U16U8InfoStruct PrimaryGroupOpTable_64[] = {
// Invalid in 64bit mode
{OPD(TYPE_GROUP_1, OpToIndex(0x82), 0), 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -163,7 +167,6 @@ void InitializePrimaryGroupTables(Context::OperatingMode Mode) {
#undef OPD
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable, std::size(PrimaryGroupOpTable));
if (Mode == Context::MODE_64BIT) {
GenerateTable(&PrimaryInstGroupOps.at(0), PrimaryGroupOpTable_64, std::size(PrimaryGroupOpTable_64));
}
@@ -12,15 +12,15 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeSecondaryGroupTables() {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = []() consteval {
std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> Table{};
#define OPD(group, prefix, Reg) (((group - FEXCore::X86Tables::TYPE_GROUP_6) << 5) | (prefix) << 3 | (Reg))
constexpr uint16_t PF_NONE = 0;
constexpr uint16_t PF_F3 = 1;
constexpr uint16_t PF_66 = 2;
constexpr uint16_t PF_F2 = 3;
static constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
constexpr U16U8InfoStruct SecondaryExtensionOpTable[] = {
// GROUP 1
// GROUP 2
// GROUP 3
@@ -487,7 +487,8 @@ void InitializeSecondaryGroupTables() {
};
#undef OPD
GenerateTable(&SecondInstGroupOps.at(0), SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
}
GenerateTable(&Table.at(0), SecondaryExtensionOpTable, std::size(SecondaryExtensionOpTable));
return Table;
}();
}
@@ -11,9 +11,9 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeSecondaryModRMTables() {
static constexpr U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps = []() consteval {
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
// REG /1
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
@@ -55,6 +55,8 @@ void InitializeSecondaryModRMTables() {
{((3 << 3) | 7), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(&SecondModRMTableOps.at(0), SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
}
GenerateTable(&Table.at(0), SecondaryModRMExtensionOpTable, std::size(SecondaryModRMExtensionOpTable));
return Table;
}();
}
@@ -13,9 +13,10 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
auto BaseOpsLambda = []() consteval {
std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> Table{};
void InitializeSecondaryTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct TwoByteOpTable[] = {
constexpr U8U8InfoStruct TwoByteOpTable[] = {
// Instructions
{0x00, 1, X86InstInfo{"", TYPE_GROUP_6, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
{0x01, 1, X86InstInfo{"", TYPE_GROUP_7, FLAGS_NO_OVERLAY, 0, nullptr}},
@@ -166,11 +167,13 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x9E, 1, X86InstInfo{"SETLE", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0, nullptr}},
{0x9F, 1, X86InstInfo{"SETNLE", TYPE_INST, GenFlagsSameSize(SIZE_8BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA0, 2, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0xA2, 1, X86InstInfo{"CPUID", TYPE_INST, FLAGS_SF_SRC_RAX | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA3, 1, X86InstInfo{"BT", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA4, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1, nullptr}},
{0xA5, 1, X86InstInfo{"SHLD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_SRC_RCX | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA6, 2, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA8, 2, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0xAA, 1, X86InstInfo{"RSM", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
{0xAB, 1, X86InstInfo{"BTS", TYPE_INST, FLAGS_DEBUG_MEM_ACCESS | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xAC, 1, X86InstInfo{"SHRD", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_NO_OVERLAY, 1, nullptr}},
@@ -265,23 +268,16 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0x3F, 1, X86InstInfo{"ALTINST", TYPE_INST, FLAGS_BLOCK_END | FLAGS_NO_OVERLAY | FLAGS_SETS_RIP, 0, nullptr}},
};
static constexpr U8U8InfoStruct TwoByteOpTable_32[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
GenerateTable(&Table.at(0), TwoByteOpTable, std::size(TwoByteOpTable));
{0xA8, 1, X86InstInfo{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
return Table;
};
static constexpr U8U8InfoStruct TwoByteOpTable_64[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
std::array<X86InstInfo, MAX_SECOND_TABLE_SIZE> SecondBaseOps = BaseOpsLambda();
std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> RepModOps = []() consteval {
std::array<X86InstInfo, MAX_REP_MOD_TABLE_SIZE> Table{};
{0xA8, 1, X86InstInfo{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
static constexpr U8U8InfoStruct RepModOpTable[] = {
constexpr U8U8InfoStruct RepModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -361,7 +357,15 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
static constexpr U8U8InfoStruct RepNEModOpTable[] = {
GenerateTableWithCopy(&Table.at(0), RepModOpTable, std::size(RepModOpTable), &BaseOpsLambda().at(0));
return Table;
}();
std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> RepNEModOps = []() consteval {
std::array<X86InstInfo, MAX_REPNE_MOD_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct RepNEModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -434,7 +438,15 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xF8, 8, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
};
static constexpr U8U8InfoStruct OpSizeModOpTable[] = {
GenerateTableWithCopy(&Table.at(0), RepNEModOpTable, std::size(RepNEModOpTable), &BaseOpsLambda().at(0));
return Table;
}();
std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> OpSizeModOps = []() consteval {
std::array<X86InstInfo, MAX_OPSIZE_MOD_TABLE_SIZE> Table{};
constexpr U8U8InfoStruct OpSizeModOpTable[] = {
{0x0, 16, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
{0x10, 1, X86InstInfo{"MOVUPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -581,19 +593,40 @@ void InitializeSecondaryTables(Context::OperatingMode Mode) {
{0xFF, 1, X86InstInfo{"", TYPE_COPY_OTHER, FLAGS_NONE, 0, nullptr}},
};
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable, std::size(TwoByteOpTable));
GenerateTableWithCopy(&Table.at(0), OpSizeModOpTable, std::size(OpSizeModOpTable), &BaseOpsLambda().at(0));
return Table;
}();
void InitializeSecondaryTables(Context::OperatingMode Mode) {
static constexpr U8U8InfoStruct TwoByteOpTable_32[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA8, 1, X86InstInfo{"PUSH GS", TYPE_INST, GenFlagsSrcSize(SIZE_16BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_DEF) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
static constexpr U8U8InfoStruct TwoByteOpTable_64[] = {
{0xA0, 1, X86InstInfo{"PUSH FS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA1, 1, X86InstInfo{"POP FS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA8, 1, X86InstInfo{"PUSH GS", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
{0xA9, 1, X86InstInfo{"POP GS", TYPE_INST, GenFlagsSizes(SIZE_16BIT, SIZE_64BIT) | FLAGS_DEBUG_MEM_ACCESS | FLAGS_NO_OVERLAY, 0, nullptr}},
};
if (Mode == Context::MODE_64BIT) {
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
LateInitCopyTable(&SecondBaseOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
LateInitCopyTable(&RepModOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
LateInitCopyTable(&RepNEModOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
LateInitCopyTable(&OpSizeModOps.at(0), TwoByteOpTable_64, std::size(TwoByteOpTable_64));
}
else {
GenerateTable(&SecondBaseOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
LateInitCopyTable(&SecondBaseOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
LateInitCopyTable(&RepModOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
LateInitCopyTable(&RepNEModOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
LateInitCopyTable(&OpSizeModOps.at(0), TwoByteOpTable_32, std::size(TwoByteOpTable_32));
}
GenerateTableWithCopy(&RepModOps.at(0), RepModOpTable, std::size(RepModOpTable), &SecondBaseOps.at(0));
GenerateTableWithCopy(&RepNEModOps.at(0), RepNEModOpTable, std::size(RepNEModOpTable), &SecondBaseOps.at(0));
GenerateTableWithCopy(&OpSizeModOps.at(0), OpSizeModOpTable, std::size(OpSizeModOpTable), &SecondBaseOps.at(0));
}
}
@@ -11,10 +11,10 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeVEXTables() {
std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> Table{};
#define OPD(map_select, pp, opcode) (((map_select - 1) << 10) | (pp << 8) | (opcode))
static constexpr U16U8InfoStruct VEXTable[] = {
constexpr U16U8InfoStruct VEXTable[] = {
// Map 0 (Reserved)
// VEX Map 1
{OPD(1, 0b00, 0x10), 1, X86InstInfo{"VMOVUPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
@@ -488,8 +488,15 @@ void InitializeVEXTables() {
};
#undef OPD
GenerateTable(&Table.at(0), VEXTable, std::size(VEXTable));
return Table;
}();
std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> VEXTableGroupOps = []() consteval {
std::array<X86InstInfo, MAX_VEX_GROUP_TABLE_SIZE> Table{};
#define OPD(group, pp, opcode) (((group - TYPE_VEX_GROUP_12) << 4) | (pp << 3) | (opcode))
static constexpr U8U8InfoStruct VEXGroupTable[] = {
constexpr U8U8InfoStruct VEXGroupTable[] = {
{OPD(TYPE_VEX_GROUP_12, 1, 0b010), 1, X86InstInfo{"VPSRLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b100), 1, X86InstInfo{"VPSRAW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
{OPD(TYPE_VEX_GROUP_12, 1, 0b110), 1, X86InstInfo{"VPSLLW", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
@@ -512,7 +519,8 @@ void InitializeVEXTables() {
};
#undef OPD
GenerateTable(&VEXTableOps.at(0), VEXTable, std::size(VEXTable));
GenerateTable(&VEXTableGroupOps.at(0), VEXGroupTable, std::size(VEXGroupTable));
}
GenerateTable(&Table.at(0), VEXGroupTable, std::size(VEXGroupTable));
return Table;
}();
}
@@ -507,26 +507,30 @@ using U8U8InfoStruct = X86TablesInfoStruct<uint8_t>;
using U16U8InfoStruct = X86TablesInfoStruct<uint16_t>;
template<typename OpcodeType>
static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
constexpr static inline void GenerateTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
ERROR_AND_DIE_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
}
FinalTable[OpNum + i] = Info;
}
}
};
template<typename OpcodeType>
static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, X86InstInfo *OtherLocal) {
constexpr static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize, X86InstInfo *OtherLocal) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
ERROR_AND_DIE_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
}
if (Info.Type == TYPE_COPY_OTHER) {
FinalTable[OpNum + i] = OtherLocal[OpNum + i];
}
@@ -538,13 +542,30 @@ static inline void GenerateTableWithCopy(X86InstInfo *FinalTable, X86TablesInfoS
};
template<typename OpcodeType>
static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
static inline void LateInitCopyTable(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *OtherLocal, size_t OtherTableSize) {
for (size_t j = 0; j < OtherTableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &OtherOp = OtherLocal[j];
auto OtherOpNum = OtherOp.first;
X86InstInfo const &OtherInfo = OtherOp.Info;
for (uint32_t i = 0; i < OtherOp.second; ++i) {
X86InstInfo &FinalOp = FinalTable[OtherOpNum + i];
if (FinalOp.Type == TYPE_COPY_OTHER) {
FinalOp = OtherInfo;
}
}
}
}
template<typename OpcodeType>
constexpr static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct<OpcodeType> const *LocalTable, size_t TableSize) {
for (size_t j = 0; j < TableSize; ++j) {
X86TablesInfoStruct<OpcodeType> const &Op = LocalTable[j];
auto OpNum = Op.first;
X86InstInfo const &Info = Op.Info;
for (uint32_t i = 0; i < Op.second; ++i) {
LOGMAN_THROW_AA_FMT(FinalTable[OpNum + i].Type == TYPE_UNKNOWN, "Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
if (FinalTable[OpNum + i].Type != TYPE_UNKNOWN) {
ERROR_AND_DIE_FMT("Duplicate Entry {}->{}", FinalTable[OpNum + i].Name, Info.Name);
}
if ((OpNum & 0b11'000'000) == 0b11'000'000) {
// If the mod field is 0b11 then it is a regular op
FinalTable[OpNum + i] = Info;
@@ -552,7 +573,9 @@ static inline void GenerateX87Table(X86InstInfo *FinalTable, X86TablesInfoStruct
else {
// If the mod field is !0b11 then this instruction is duplicated through the whole mod [0b00, 0b10] range
// and the modrm.rm space because that is used part of the instruction encoding
LOGMAN_THROW_AA_FMT((OpNum & 0b11'000'000) == 0, "Only support mod field of zero in this path");
if ((OpNum & 0b11'000'000) != 0) {
ERROR_AND_DIE_FMT("Only support mod field of zero in this path");
}
for (uint16_t mod = 0b00'000'000; mod < 0b11'000'000; mod += 0b01'000'000) {
for (uint16_t rm = 0b000; rm < 0b1'000; ++rm) {
FinalTable[(OpNum | mod | rm) + i] = Info;
@@ -11,11 +11,11 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeX87Tables() {
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> X87Ops = []() consteval {
std::array<X86InstInfo, MAX_X87_TABLE_SIZE> Table{};
#define OPD(op, modrmop) (((op - 0xD8) << 8) | modrmop)
#define OPDReg(op, reg) (((op - 0xD8) << 8) | (reg << 3))
static constexpr U16U8InfoStruct X87OpTable[] = {
constexpr U16U8InfoStruct X87OpTable[] = {
// 0xD8
{OPDReg(0xD8, 0), 1, X86InstInfo{"FADD", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
{OPDReg(0xD8, 1), 1, X86InstInfo{"FMUL", TYPE_X87, FLAGS_MODRM, 0, nullptr}},
@@ -263,6 +263,8 @@ void InitializeX87Tables() {
#undef OPD
#undef OPDReg
GenerateX87Table(&X87Ops.at(0), X87OpTable, std::size(X87OpTable));
}
GenerateX87Table(&Table.at(0), X87OpTable, std::size(X87OpTable));
return Table;
}();
}
@@ -12,14 +12,14 @@ $end_info$
namespace FEXCore::X86Tables {
using namespace InstFlags;
void InitializeXOPTables() {
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> XOPTableOps = []() consteval {
std::array<X86InstInfo, MAX_XOP_TABLE_SIZE> Table{};
#define OPD(group, pp, opcode) ( (group << 10) | (pp << 8) | (opcode))
constexpr uint16_t XOP_GROUP_8 = 0;
constexpr uint16_t XOP_GROUP_9 = 1;
constexpr uint16_t XOP_GROUP_A = 2;
static constexpr U16U8InfoStruct XOPTable[] = {
constexpr U16U8InfoStruct XOPTable[] = {
// Group 8
{OPD(XOP_GROUP_8, 0, 0x85), 1, X86InstInfo{"VPMAXSSWW", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(XOP_GROUP_8, 0, 0x86), 1, X86InstInfo{"VPMACSSWD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -104,8 +104,15 @@ void InitializeXOPTables() {
};
#undef OPD
GenerateTable(&Table.at(0), XOPTable, std::size(XOPTable));
return Table;
}();
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> XOPTableGroupOps = []() consteval {
std::array<X86InstInfo, MAX_XOP_GROUP_TABLE_SIZE> Table{};
#define OPD(subgroup, opcode) (((subgroup - 1) << 3) | (opcode))
static constexpr U8U8InfoStruct XOPGroupTable[] = {
constexpr U8U8InfoStruct XOPGroupTable[] = {
// Group 1
{OPD(1, 1), 1, X86InstInfo{"BLCFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
{OPD(1, 2), 1, X86InstInfo{"BLSFILL", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
@@ -129,7 +136,8 @@ void InitializeXOPTables() {
};
#undef OPD
GenerateTable(&XOPTableOps.at(0), XOPTable, std::size(XOPTable));
GenerateTable(&XOPTableGroupOps.at(0), XOPGroupTable, std::size(XOPGroupTable));
}
GenerateTable(&Table.at(0), XOPGroupTable, std::size(XOPGroupTable));
return Table;
}();
}
+26 -2
View File
@@ -6,12 +6,13 @@ tags: glue|thunks
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include <FEXCore/Config/Config.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/fextl/set.h>
#include <FEXCore/fextl/string.h>
@@ -180,6 +181,9 @@ namespace FEXCore {
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDI] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSI] = (uintptr_t)arg1;
} else {
if ((reinterpret_cast<uintptr_t>(arg1) >> 32) != 0) {
ERROR_AND_DIE_FMT("Tried to call guest function with arguments packed to a 64-bit address");
}
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RCX] = (uintptr_t)arg0;
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RDX] = (uintptr_t)arg1;
}
@@ -216,7 +220,7 @@ namespace FEXCore {
LogMan::Msg::DFmt("Thunks: Adding guest trampoline from address {:#x} to guest function {:#x}",
args->original_callee, args->target_addr);
auto Result = Thread->CTX->AddCustomIREntrypoint(
auto Result = CTX->AddCustomIREntrypoint(
args->original_callee,
[CTX, GuestThunkEntrypoint = args->target_addr](uintptr_t Entrypoint, FEXCore::IR::IREmitter *emit) {
auto IRHeader = emit->_IRHeader(emit->Invalid(), Entrypoint, 0, 0);
@@ -483,6 +487,26 @@ namespace FEXCore {
}
}
FEX_DEFAULT_VISIBILITY void* GetGuestStack() {
if (!Thread) {
ERROR_AND_DIE_FMT("Thunked library attempted to query guest stack pointer asynchronously");
}
return (void*)(uintptr_t)((Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSP]));
}
FEX_DEFAULT_VISIBILITY void MoveGuestStack(uintptr_t NewAddress) {
if (!Thread) {
ERROR_AND_DIE_FMT("Thunked library attempted to query guest stack pointer asynchronously");
}
if (NewAddress >> 32) {
ERROR_AND_DIE_FMT("Tried to set stack pointer for 32-bit guest to a 64-bit address");
}
Thread->CurrentFrame->State.gregs[FEXCore::X86State::REG_RSP] = NewAddress;
}
#else
fextl::unique_ptr<ThunkHandler> ThunkHandler::Create() {
ERROR_AND_DIE_FMT("Unsupported");
+19
View File
@@ -0,0 +1,19 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/ThreadPoolAllocator.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/fextl/memory.h>
#include <FEXCore/fextl/sstream.h>
namespace FEXCore::IR {
class RegisterAllocationData;
class IRListView;
class IREmitter;
void Dump(fextl::stringstream *out, IRListView const* IR, IR::RegisterAllocationData *RAData);
fextl::unique_ptr<IREmitter> Parse(FEXCore::Utils::IntrusivePooledAllocator &ThreadAllocator, fextl::stringstream &MapsStream);
}
+95 -11
View File
@@ -347,8 +347,7 @@
"Desc": ["Loads a value from the static-ra context with offset",
"Dest = Ctx[Offset]"
],
"DestSize": "Size",
"DynamicDispatch": true
"DestSize": "Size"
},
"StoreRegister SSA:$Value, i1:$IsPrewrite, u32:$Offset, RegisterClass:$Class, RegisterClass:$StaticClass, u8:#Size": {
@@ -359,7 +358,6 @@
"Truncates if value's type is too large"
],
"DestSize": "Size",
"DynamicDispatch": true,
"EmitValidation": [
"WalkFindRegClass($Value) == $Class"
]
@@ -953,14 +951,55 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AddNZCV OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Set NZCV for the sum of two GPRs"],
"HasSideEffects": true,
"GPR = Adc OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": [ "Integer Add with carry",
"Will truncate to 64 or 32bits"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = Sbb OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": [ "Integer Subtract with carry/borrow",
"Will truncate to 64 or 32bits"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AddShift OpSize:#Size, GPR:$Src1, GPR:$Src2, ShiftType:$Shift{ShiftType::LSL}, u8:$ShiftAmount{0}": {
"Desc": [ "Integer Add with shifted register",
"Will truncate to 64 or 32bits"
],
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit",
"_Shift != ShiftType::ROR"
]
},
"GPR = AddWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": [ "Integer add. Truncates and sets NZCV per AddNZCV"],
"DestSize": "Size",
"HasSideEffects": true,
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AddNZCV OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Set NZCV for the sum of two GPRs"],
"HasSideEffects": true,
"DestSize": "Size"
},
"SetSmallNZV OpSize:#Size, GPR:$Src": {
"Desc": ["Set NZV with a SETF instruction. Preserves CF."],
"HasSideEffects": true,
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i8Bit || Size == FEXCore::IR::OpSize::i16Bit"
]
},
"CarryInvert": {
"Desc": ["Invert carry flag in NZCV"],
"HasSideEffects": true
@@ -981,6 +1020,22 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AdcWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Adds and set NZCV for the sum of two GPRs and carry-in given as NZCV"],
"HasSideEffects": true,
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = SbbWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Subtracts and set NZCV for the difference of two GPRs and carry-in given as NZCV"],
"HasSideEffects": true,
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"AdcNZCV OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Set NZCV for the sum of two GPRs and carry-in given as NZCV"],
"HasSideEffects": true,
@@ -1016,16 +1071,21 @@
"_Shift != ShiftType::ROR"
]
},
"SubNZCV OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Set NZCV for the difference of two GPRs. ",
"Carry flag uses arm64 definition, inverted x86.",
""],
"GPR = SubWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": [ "Integer Sub. Truncates and sets NZCV per SubNZCV"],
"DestSize": "Size",
"HasSideEffects": true,
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"SubNZCV OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Set NZCV for the difference of two GPRs. ",
"Carry flag uses arm64 definition, inverted x86.",
""],
"DestSize": "Size",
"HasSideEffects": true
},
"GPR = Or OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer binary or"
],
@@ -1081,6 +1141,12 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = AndWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer binary and"
],
"DestSize": "Size",
"HasSideEffects": true
},
"GPR = Andn OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer binary AND NOT. Performs the equivalent of Src1 & ~Src2"],
"DestSize": "Size",
@@ -1141,6 +1207,18 @@
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
]
},
"GPR = UMull GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer unsigned multiplication long",
"Multiplies two 32-bit numbers, returning a 64-bit destination register."
],
"DestSize": "FEXCore::IR::OpSize::i64Bit"
},
"GPR = SMull GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer signed multiplication long",
"Multiplies two 32-bit numbers, returning a 64-bit destination register."
],
"DestSize": "FEXCore::IR::OpSize::i64Bit"
},
"GPR = Div OpSize:#Size, GPR:$Src1, GPR:$Src2": {
"Desc": ["Integer signed division"
],
@@ -1592,7 +1670,13 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VUMaxV u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"Desc": ["Does a horizontal vector unsigned maximum of elements across the source vector",
"Result is a zero extended scalar"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VFAbs u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
+2 -1
View File
@@ -6,8 +6,9 @@ tags: ir|emitter
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
+3 -3
View File
@@ -6,9 +6,9 @@ tags: ir|parser
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/StringUtils.h>
#include <FEXCore/fextl/sstream.h>
@@ -505,8 +505,8 @@ class IRParser: public FEXCore::IR::IREmitter {
}
if (Def.HasArgs) {
RemainingLine = FEXCore::StringUtils::Trim(RemainingLine.substr(CurrentPos));
CurrentPos = 0;
RemainingLine =
FEXCore::StringUtils::Trim(RemainingLine.substr(CurrentPos));
if (RemainingLine.empty()) {
// How did we get here?
Def.HasArgs = false;
+4 -4
View File
@@ -66,7 +66,7 @@ void PassManager::Finalize() {
}
}
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl *ctx, bool InlineConstants, bool StaticRegisterAllocation) {
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl *ctx, bool InlineConstants) {
FEX_CONFIG_OPT(DisablePasses, O0);
if (!DisablePasses()) {
@@ -82,7 +82,7 @@ void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl *ctx, bool Inli
InsertPass(CreatePassDeadCodeElimination());
InsertPass(CreateConstProp(InlineConstants, ctx->HostFeatures.SupportsTSOImm9));
////// InsertPass(CreateDeadFlagCalculationEliminination());
InsertPass(CreateDeadFlagCalculationEliminination());
InsertPass(CreateInlineCallOptimization(&ctx->CPUID));
InsertPass(CreatePassDeadCodeElimination());
@@ -101,8 +101,8 @@ void PassManager::AddDefaultValidationPasses() {
#endif
}
void PassManager::InsertRegisterAllocationPass(bool OptimizeSRA, bool SupportsAVX) {
InsertPass(IR::CreateRegisterAllocationPass(GetPass("Compaction"), OptimizeSRA, SupportsAVX), "RA");
void PassManager::InsertRegisterAllocationPass(bool SupportsAVX) {
InsertPass(IR::CreateRegisterAllocationPass(GetPass("Compaction"), SupportsAVX), "RA");
}
bool PassManager::Run(IREmitter *IREmit) {
+2 -9
View File
@@ -29,8 +29,6 @@ namespace FEXCore::IR {
class PassManager;
class IREmitter;
using ShouldExitHandler = std::function<void(void)>;
class Pass {
public:
virtual ~Pass() = default;
@@ -47,7 +45,7 @@ protected:
class PassManager final {
friend class InlineCallOptimization;
public:
void AddDefaultPasses(FEXCore::Context::ContextImpl *ctx, bool InlineConstants, bool StaticRegisterAllocation);
void AddDefaultPasses(FEXCore::Context::ContextImpl *ctx, bool InlineConstants);
void AddDefaultValidationPasses();
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
auto PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
@@ -58,14 +56,10 @@ public:
return PassPtr;
}
void InsertRegisterAllocationPass(bool OptimizeSRA, bool SupportsAVX);
void InsertRegisterAllocationPass(bool SupportsAVX);
bool Run(IREmitter *IREmit);
void RegisterExitHandler(ShouldExitHandler Handler) {
ExitHandler = std::move(Handler);
}
bool HasPass(fextl::string Name) const {
return NameToPassMaping.contains(Name);
}
@@ -86,7 +80,6 @@ public:
void Finalize();
protected:
ShouldExitHandler ExitHandler;
FEXCore::HLE::SyscallHandler *SyscallHandler;
private:
-1
View File
@@ -24,7 +24,6 @@ fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination(bool SupportsAVX
fextl::unique_ptr<FEXCore::IR::Pass> CreatePassDeadCodeElimination();
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction(FEXCore::Utils::IntrusivePooledAllocator &Allocator);
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass,
bool OptimizeSRA,
bool SupportsAVX);
fextl::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass();
@@ -13,10 +13,10 @@ $end_info$
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
@@ -833,12 +833,17 @@ bool ConstProp::ConstantPropagation(IREmitter *IREmit, const IRListView& Current
uint64_t Constant;
if (IREmit->IsValueConstant(Op->Src, &Constant)) {
// SBFE of a constant can be converted to a constant.
uint64_t SourceMask = Op->Width == 64 ? ~0ULL : ((1ULL << Op->Width) - 1);
uint64_t SourceMask =
Op->Width == 64 ? ~0ULL : ((1ULL << Op->Width) - 1);
uint64_t DestSizeInBits = IROp->Size * 8;
uint64_t DestMask =
DestSizeInBits == 64 ? ~0ULL : ((1ULL << DestSizeInBits) - 1);
SourceMask <<= Op->lsb;
int64_t NewConstant = (Constant & SourceMask) >> Op->lsb;
NewConstant <<= 64 - Op->Width;
NewConstant >>= 64 - Op->Width;
NewConstant &= DestMask;
IREmit->ReplaceWithConstant(CodeNode, NewConstant);
Changed = true;
@@ -960,20 +965,24 @@ bool ConstProp::ConstantInlining(IREmitter *IREmit, const IRListView& CurrentIR)
case OP_SUB:
case OP_ADDNZCV:
case OP_SUBNZCV:
case OP_ADDWITHFLAGS:
case OP_SUBWITHFLAGS:
{
auto Op = IROp->C<IR::IROp_Add>();
uint64_t Constant2{};
if (IREmit->IsValueConstant(Op->Header.Args[1], &Constant2)) {
if (IsImmAddSub(Constant2)) {
// We don't allow 8/16-bit operations to have constants, since no
// constant would be in bounds after the JIT's 24/16 shift.
if (IsImmAddSub(Constant2) && Op->Header.Size >= 4) {
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->Header.Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, CreateInlineConstant(IREmit, Constant2));
Changed = true;
}
} else if (IROp->Op == OP_SUBNZCV) {
// If the first source is zero, we can use a NEGS instruction.
} else if (IROp->Op == OP_SUBNZCV || IROp->Op == OP_SUBWITHFLAGS || IROp->Op == OP_SUB) {
// TODO: Generalize this
uint64_t Constant1{};
if (IREmit->IsValueConstant(Op->Header.Args[0], &Constant1)) {
if (Constant1 == 0) {
@@ -986,6 +995,22 @@ bool ConstProp::ConstantInlining(IREmitter *IREmit, const IRListView& CurrentIR)
break;
}
case OP_ADC:
case OP_ADCWITHFLAGS:
{
auto Op = IROp->C<IR::IROp_Adc>();
uint64_t Constant1{};
if (IREmit->IsValueConstant(Op->Header.Args[0], &Constant1)) {
if (Constant1 == 0) {
IREmit->SetWriteCursor(CurrentIR.GetNode(Op->Header.Args[0]));
IREmit->ReplaceNodeArgument(CodeNode, 0, CreateInlineConstant(IREmit, 0));
Changed = true;
}
}
break;
}
case OP_CONDADDNZCV:
{
auto Op = IROp->C<IR::IROp_CondAddNZCV>();
@@ -1134,6 +1159,7 @@ bool ConstProp::ConstantInlining(IREmitter *IREmit, const IRListView& CurrentIR)
case OP_OR:
case OP_XOR:
case OP_AND:
case OP_ANDWITHFLAGS:
case OP_ANDN:
{
auto Op = IROp->CW<IR::IROp_Or>();
@@ -5,10 +5,10 @@ tags: ir|opts
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/Profiler.h>
@@ -6,12 +6,12 @@ desc: Transforms ContextLoad/Store to temporaries, similar to mem2reg
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/EnumOperators.h>
#include <FEXCore/Utils/LogManager.h>
@@ -544,7 +544,8 @@ bool RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter *IREmit, ContextInfo *Loc
bool RCLSE::ClassifyContextStore(FEXCore::IR::IREmitter *IREmit, ContextInfo *LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset, uint8_t Size, FEXCore::IR::OrderedNode *CodeNode, FEXCore::IR::OrderedNode *ValueNode) {
auto Info = FindMemberInfo(LocalInfo, Offset, Size);
Info = RecordAccess(Info, Class, Offset, Size, LastAccessType::WRITE, ValueNode, CodeNode);
RecordAccess(Info, Class, Offset, Size, LastAccessType::WRITE, ValueNode,
CodeNode);
// TODO: Optimize redundant stores.
// ContextMemberInfo PreviousMemberInfoCopy = *Info;
return false;
@@ -6,11 +6,11 @@ desc: Cross block store-after-store elimination
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
@@ -6,11 +6,11 @@ desc: Sorts the ssa storage in memory, needed for RA and others
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/Core/OpcodeDispatcher.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
@@ -6,6 +6,8 @@ desc: Prints IR
$end_info$
*/
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "Interface/Core/OpcodeDispatcher.h"
@@ -6,12 +6,13 @@ desc: Sanity checking pass
$end_info$
*/
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/LogManager.h>
@@ -7,10 +7,10 @@ $end_info$
*/
#include "Interface/Core/CPUID.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/Utils/Profiler.h>
@@ -6,9 +6,9 @@ desc: Long divide elimination pass
$end_info$
*/
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/Profiler.h>
@@ -1,10 +1,12 @@
// SPDX-License-Identifier: MIT
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include "Interface/IR/Passes/IRValidation.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/Profiler.h>
@@ -6,8 +6,10 @@ desc: This is not used right now, possibly broken
$end_info$
*/
#include "FEXCore/Core/X86Enums.h"
#include "Interface/IR/IREmitter.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/Profiler.h>
@@ -16,54 +18,374 @@ $end_info$
#include <array>
#include <memory>
// Flag bit flags
#define FLAG_V (1U << 0)
#define FLAG_C (1U << 1)
#define FLAG_Z (1U << 2)
#define FLAG_N (1U << 3)
#define FLAG_A (1U << 4)
#define FLAG_P (1U << 5)
#define FLAG_ZCV (FLAG_Z | FLAG_C | FLAG_V)
#define FLAG_NZCV (FLAG_N | FLAG_ZCV)
#define FLAG_ALL (FLAG_NZCV | FLAG_A | FLAG_P)
namespace FEXCore::IR {
struct FlagInfo {
// If set, all following fields are zero, used for a quick exit.
bool Trivial;
// Set of flags read by the instruction.
unsigned Read;
// Set of flags written by the instruction. Happens AFTER the reads.
unsigned Write;
// If true, the instruction can be be eliminated if its flag writes can all be
// eliminated.
bool CanEliminate;
// If true, the opcode can be replaced with Replacement if its flag writes can
// all be eliminated.
bool CanReplace;
IROps Replacement;
};
class DeadFlagCalculationEliminination final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
private:
FlagInfo Classify(IROp_Header *Node);
unsigned FlagForOffset(unsigned Offset);
unsigned FlagsForCondClassType(CondClassType Cond);
};
unsigned
DeadFlagCalculationEliminination::FlagForOffset(unsigned Offset)
{
return Offset == offsetof(FEXCore::Core::CPUState, pf_raw) ? FLAG_P :
Offset == offsetof(FEXCore::Core::CPUState, af_raw) ? FLAG_A :
0;
};
unsigned
DeadFlagCalculationEliminination::FlagsForCondClassType(CondClassType Cond)
{
switch (Cond) {
case COND_AL:
return 0;
case COND_MI:
case COND_PL:
return FLAG_N;
case COND_EQ:
case COND_NEQ:
return FLAG_Z;
case COND_UGE:
case COND_ULT:
return FLAG_C;
case COND_VS:
case COND_VC:
case COND_FU:
case COND_FNU:
return FLAG_V;
case COND_UGT:
case COND_ULE:
return FLAG_Z | FLAG_C;
case COND_SGE:
case COND_SLT:
case COND_FLU:
case COND_FGE:
return FLAG_N | FLAG_V;
case COND_SGT:
case COND_SLE:
case COND_FLEU:
case COND_FGT:
return FLAG_N | FLAG_Z | FLAG_V;
default:
LOGMAN_THROW_AA_FMT(false, "unknown cond class type");
return FLAG_NZCV;
}
}
FlagInfo
DeadFlagCalculationEliminination::Classify(IROp_Header *IROp)
{
switch (IROp->Op) {
case OP_ANDWITHFLAGS:
return {
.Write = FLAG_NZCV,
.CanReplace = true,
.Replacement = OP_AND,
};
case OP_ADDWITHFLAGS:
return {
.Write = FLAG_NZCV,
.CanReplace = true,
.Replacement = OP_ADD,
};
case OP_SUBWITHFLAGS:
return {
.Write = FLAG_NZCV,
.CanReplace = true,
.Replacement = OP_SUB,
};
case OP_ADCWITHFLAGS:
return {
.Read = FLAG_C,
.Write = FLAG_NZCV,
.CanReplace = true,
.Replacement = OP_ADC,
};
case OP_SBBWITHFLAGS:
return {
.Read = FLAG_C,
.Write = FLAG_NZCV,
.CanReplace = true,
.Replacement = OP_SBB,
};
case OP_ADDNZCV:
case OP_SUBNZCV:
case OP_TESTNZ:
case OP_FCMP:
case OP_STORENZCV:
return {
.Write = FLAG_NZCV,
.CanEliminate = true,
};
case OP_AXFLAG:
// Per the Arm spec, axflag reads Z/V/C but not N. It writes all flags.
return {
.Read = FLAG_ZCV,
.Write = FLAG_NZCV,
.CanEliminate = true,
};
case OP_CARRYINVERT:
return {
.Read = FLAG_C,
.Write = FLAG_C,
.CanEliminate = true,
};
case OP_SETSMALLNZV:
return {
.Write = FLAG_N | FLAG_Z | FLAG_V,
.CanEliminate = true,
};
case OP_LOADNZCV:
return {.Read = FLAG_NZCV};
case OP_ADC:
case OP_SBB:
return {.Read = FLAG_C};
case OP_ADCNZCV:
case OP_SBBNZCV:
return {
.Read = FLAG_C,
.Write = FLAG_NZCV,
.CanEliminate = true,
};
case OP_NZCVSELECT: {
auto Op = IROp->CW<IR::IROp_NZCVSelect>();
return {.Read = FlagsForCondClassType(Op->Cond)};
}
case OP_NEG: {
auto Op = IROp->CW<IR::IROp_Neg>();
return {.Read = FlagsForCondClassType(Op->Cond)};
}
case OP_CONDJUMP: {
auto Op = IROp->CW<IR::IROp_CondJump>();
if (!Op->FromNZCV)
break;
return {.Read = FlagsForCondClassType(Op->Cond)};
}
case OP_CONDADDNZCV: {
auto Op = IROp->CW<IR::IROp_CondAddNZCV>();
return {
.Read = FlagsForCondClassType(Op->Cond),
.Write = FLAG_NZCV,
.CanEliminate = true,
};
}
case OP_RMIFNZCV: {
auto Op = IROp->CW<IR::IROp_RmifNZCV>();
static_assert(FLAG_N == (1 << 3), "rmif mask lines up with our bits");
static_assert(FLAG_Z == (1 << 2), "rmif mask lines up with our bits");
static_assert(FLAG_C == (1 << 1), "rmif mask lines up with our bits");
static_assert(FLAG_V == (1 << 0), "rmif mask lines up with our bits");
return {
.Write = Op->Mask,
.CanEliminate = true,
};
}
case OP_INVALIDATEFLAGS: {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
unsigned Flags = 0;
// TODO: Make this translation less silly
if (Op->Flags & (1u << X86State::RFLAG_SF_RAW_LOC))
Flags |= FLAG_N;
if (Op->Flags & (1u << X86State::RFLAG_ZF_RAW_LOC))
Flags |= FLAG_Z;
if (Op->Flags & (1u << X86State::RFLAG_CF_RAW_LOC))
Flags |= FLAG_C;
if (Op->Flags & (1u << X86State::RFLAG_OF_RAW_LOC))
Flags |= FLAG_V;
if (Op->Flags & (1u << X86State::RFLAG_PF_RAW_LOC))
Flags |= FLAG_P;
if (Op->Flags & (1u << X86State::RFLAG_AF_RAW_LOC))
Flags |= FLAG_A;
// The mental model of InvalidateFlags is writing undefined values to all
// of the selected flags, allowing the write-after-write optimizations to
// optimize invalidate-after-write for free.
return {
.Write = Flags,
.CanEliminate = true,
};
}
case OP_LOADREGISTER: {
auto Op = IROp->CW<IR::IROp_LoadRegister>();
if (Op->Class != GPRClass || Op->StaticClass != GPRFixedClass)
break;
return {.Read = FlagForOffset(Op->Offset)};
}
case OP_STOREREGISTER: {
auto Op = IROp->CW<IR::IROp_StoreRegister>();
if (Op->Class != GPRClass || Op->StaticClass != GPRFixedClass)
break;
LOGMAN_THROW_A_FMT(!Op->IsPrewrite, "PF/AF writes are fixed-form");
unsigned Flag = FlagForOffset(Op->Offset);
return {
.Write = Flag,
.CanEliminate = Flag != 0,
};
}
default:
break;
}
return {.Trivial = true};
}
/**
* @brief (UNSAFE) This pass removes flag calculations that will otherwise be unused INSIDE of that block
*
* Compilers don't really do any form of cross-block flag allocation like they do RA with GPRs.
* This ends up with them recalculating flags across blocks regardless of if it is actually possible to reuse the flags.
* This is an additional burden in x86 that most instructions change flags when called, so it is easier to recalculate anyway.
*
* This is unsafe since handwritten code can easily break this assumption.
* This may be more interesting with full function level recompilation since flags definitely won't be used across function boundaries.
*
* @brief This pass removes flag calculations that will otherwise be unused INSIDE of that block
*/
bool DeadFlagCalculationEliminination::Run(IREmitter *IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::DFE");
std::array<OrderedNode*, 32> LastValidFlagStores{};
bool Changed = false;
auto CurrentIR = IREmit->ViewIR();
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->CW<IR::IROp_StoreFlag>();
// Set this node as the last one valid for this flag
LastValidFlagStores[Op->Flag] = CodeNode;
}
else if (IROp->Op == OP_LOADFLAG) {
auto Op = IROp->CW<IR::IROp_LoadFlag>();
LastValidFlagStores[Op->Flag] = nullptr;
}
}
// We model all flags as read at the end of the block, since this pass is
// presently purely local. Optimizing this requires global anslysis.
uint32_t FlagsRead = FLAG_ALL;
// If any flags are stored but not loaded by the end of the block, then erase them
for (auto &Flag : LastValidFlagStores) {
if (Flag != nullptr) {
IREmit->Remove(Flag);
// Reverse iteration is not yet working with the iterators
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
// We grab these nodes this way so we can iterate easily
auto CodeBegin = CurrentIR.at(BlockIROp->Begin);
auto CodeLast = CurrentIR.at(BlockIROp->Last);
// Iterate the block in reverse
while (1) {
auto [CodeNode, IROp] = CodeLast();
// Optimizing flags can cause earlier flag reads to become dead but dead
// flag reads should not impede optimiation of earlier dead flag writes.
// We must DCE as we go to ensure we converge in a single iteration.
//
// TODO: This whole pass could be merged with DCE?
bool HasSideEffects = IR::HasSideEffects(IROp->Op);
if (!HasSideEffects && CodeNode->GetUses() == 0) {
Changed = true;
}
}
IREmit->Remove(CodeNode);
} else {
// Optimiation algorithm: For each flag written...
//
// If the flag has a later read (per FlagsRead), remove the flag from
// FlagsRead, since the reader is covered by this write.
//
// Else, there is no later read, so remove the flag write (if we can).
// This is the active part of the optimization.
//
// Then, add each flag read to FlagsRead.
//
// This order is important: instructions that read-modify-write flags
// (like adcs) first read flags, then write flags. Since we're iterating
// the block backwards, that means we handle the write first.
struct FlagInfo Info = Classify(IROp);
LastValidFlagStores.fill(nullptr);
if (!Info.Trivial) {
bool Eliminated = false;
if ((FlagsRead & Info.Write) == 0) {
if (Info.CanEliminate) {
IREmit->Remove(CodeNode);
Eliminated = true;
Changed = true;
} else if (Info.CanReplace) {
IROp->Op = Info.Replacement;
Changed = true;
}
} else {
FlagsRead &= ~Info.Write;
}
// If we eliminated the instruction, we eliminate its read too. This
// check is required to ensure the pass converges locally in a single
// iteration.
if (!Eliminated)
FlagsRead |= Info.Read;
}
}
// Iterate in reverse
if (CodeLast == CodeBegin) {
break;
}
--CodeLast;
}
}
return Changed;
@@ -5,16 +5,16 @@ tags: ir|opts
$end_info$
*/
#include "Utils/BucketList.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
#include "FEXCore/Core/X86Enums.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/Utils/BitUtils.h>
#include <FEXCore/Utils/BucketList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/Utils/Profiler.h>
@@ -24,6 +24,7 @@ $end_info$
#include <FEXCore/fextl/unordered_set.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <FEXHeaderUtils/TypeDefines.h>
#include <algorithm>
@@ -223,7 +224,7 @@ namespace {
class ConstrainedRAPass final : public RegisterAllocationPass {
public:
ConstrainedRAPass(FEXCore::IR::Pass* _CompactionPass, bool OptimizeSRA, bool SupportsAVX);
ConstrainedRAPass(FEXCore::IR::Pass* _CompactionPass, bool SupportsAVX);
~ConstrainedRAPass();
bool Run(IREmitter *IREmit) override;
@@ -248,7 +249,6 @@ namespace {
RegisterGraph *Graph;
FEXCore::IR::Pass* CompactionPass;
bool OptimizeSRA;
bool SupportsAVX;
fextl::vector<LiveRange> LiveRanges;
@@ -297,8 +297,8 @@ namespace {
bool RunAllocateVirtualRegisters(IREmitter *IREmit);
};
ConstrainedRAPass::ConstrainedRAPass(FEXCore::IR::Pass* _CompactionPass, bool _OptimizeSRA, bool _SupportsAVX)
: CompactionPass {_CompactionPass}, OptimizeSRA(_OptimizeSRA), SupportsAVX{_SupportsAVX} {
ConstrainedRAPass::ConstrainedRAPass(FEXCore::IR::Pass* _CompactionPass, bool _SupportsAVX)
: CompactionPass {_CompactionPass}, SupportsAVX{_SupportsAVX} {
}
ConstrainedRAPass::~ConstrainedRAPass() {
@@ -700,8 +700,10 @@ namespace {
// Marking here as written is overly agressive, but
// there might be write(s) later on the instruction stream
if ((*StaticMap)) {
SRA_DEBUG("Markng ssa{} as written because ssa{} re-loads sra{}, and we can't track possible future writes\n",
(*StaticMap) - &LiveRanges[0], Node, -1 /*vreg*/);
SRA_DEBUG(
"Marking ssa{} as written because ssa{} re-loads sra{}, "
"and we can't track possible future writes\n",
(*StaticMap) - &LiveRanges[0], Node, -1 /*vreg*/);
(*StaticMap)->Written = true;
}
@@ -1378,8 +1380,9 @@ namespace {
auto FilledInterference = IREmit->_FillRegister(InterferenceOrderedNode, SpillSlot, InterferenceRegClass);
FilledInterference.first->Header.Size = InterferenceIROp->Size;
FilledInterference.first->Header.ElementSize = InterferenceIROp->ElementSize;
IREmit->ReplaceUsesWithAfter(InterferenceOrderedNode, FilledInterference, FilledInterference);
Spilled = true;
IREmit->ReplaceUsesWithAfter(InterferenceOrderedNode,
FilledInterference,
FilledInterference);
}
}
}
@@ -1402,8 +1405,7 @@ namespace {
ResetRegisterGraph(Graph, SSACount);
FindNodeClasses(Graph, &IR);
CalculateLiveRange(&IR);
if (OptimizeSRA)
OptimizeStaticRegisters(&IR);
OptimizeStaticRegisters(&IR);
// Linear forward scan based interference calculation is faster for smaller blocks
// Smarter block based interference calculation is faster for larger blocks
@@ -1470,7 +1472,7 @@ namespace {
return Changed;
}
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass, bool OptimizeSRA, bool SupportsAVX) {
return fextl::make_unique<ConstrainedRAPass>(CompactionPass, OptimizeSRA, SupportsAVX);
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass, bool SupportsAVX) {
return fextl::make_unique<ConstrainedRAPass>(CompactionPass, SupportsAVX);
}
}
@@ -6,10 +6,11 @@ desc: Sanity Checking
$end_info$
*/
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
+6
View File
@@ -297,6 +297,12 @@ namespace FEXCore::Allocator {
if (c == ' ') {
STEAL_LOG("[%d] ParseEnd; RegionBegin: %016lX RegionEnd: %016lX\n", __LINE__, RegionBegin, RegionEnd);
if (RegionEnd > End) {
// Early return if we are completely beyond the allocation space.
close(MapsFD);
return Regions;
}
State = ScanEnd;
// If the previous map's ending and the region we just parsed overlap the stack then we need to save the stack mapping.
+77 -67
View File
@@ -1,4 +1,8 @@
// SPDX-License-Identifier: MIT
#include "Utils/SpinWaitLock.h"
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Telemetry.h>
@@ -522,9 +526,7 @@ static bool HandleAtomicVectorStore(uint32_t Instr, uintptr_t ProgramCounter) {
uint32_t *PC = (uint32_t*)ProgramCounter;
uint32_t Size = (Instr >> 30) & 1;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
uint32_t DataReg2 = (Instr >> 10) & 0x1F;
if(Size == 1) {
// 64-bit pair happens on paranoid vector stores
@@ -533,9 +535,9 @@ static bool HandleAtomicVectorStore(uint32_t Instr, uintptr_t ProgramCounter) {
// [2] cbnz(TMP3, &B); // < Overwritten with DMB
if (DataReg == 31) {
uint32_t NextInstr = PC[1];
AddrReg = (NextInstr >> 5) & 0x1F;
uint32_t AddrReg = (NextInstr >> 5) & 0x1F;
DataReg = NextInstr & 0x1F;
DataReg2 = (NextInstr >> 10) & 0x1F;
uint32_t DataReg2 = (NextInstr >> 10) & 0x1F;
uint32_t STP =
(0b10 << 30) |
(0b101001000000000 << 15) |
@@ -2025,7 +2027,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return NumInstructionsToSkip * 4;
}
[[nodiscard]] std::pair<bool, int32_t> HandleUnalignedAccess(bool ParanoidTSO, uintptr_t ProgramCounter, uint64_t *GPRs) {
[[nodiscard]] std::pair<bool, int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState *Thread, bool ParanoidTSO, uintptr_t ProgramCounter, uint64_t *GPRs) {
#ifdef _M_ARM_64
constexpr bool is_arm64 = true;
#else
@@ -2044,9 +2046,11 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
uint32_t Size = (Instr & 0xC000'0000) >> 30;
uint32_t AddrReg = (Instr >> 5) & 0x1F;
uint32_t DataReg = Instr & 0x1F;
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ParanoidTSO) {
// ParanoidTSO path doesn't modify any code.
if (ParanoidTSO) [[unlikely]] {
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, 0)) {
// Skip this instruction now
return std::make_pair(true, 4);
@@ -2056,20 +2060,7 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return NotHandled;
}
}
else {
uint32_t LDR = 0b0011'1000'0111'1111'0110'1000'0000'0000;
LDR |= Size << 30;
LDR |= AddrReg << 5;
LDR |= DataReg;
PC[0] = LDR;
PC[1] = DMB_LD; // Back-patch the half-barrier.
ClearICache(&PC[-1], 16);
// With the instruction modified, now execute again.
return std::make_pair(true, 0);
}
}
else if ( (Instr & LDAXR_MASK) == STLR_INST) { // STLR*
if (ParanoidTSO) {
else if ( (Instr & LDAXR_MASK) == STLR_INST) { // STLR*
if (ArchHelpers::Arm64::HandleAtomicStore(Instr, GPRs, 0)) {
// Skip this instruction now
return std::make_pair(true, 4);
@@ -2079,22 +2070,9 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return NotHandled;
}
}
else {
uint32_t STR = 0b0011'1000'0011'1111'0110'1000'0000'0000;
STR |= Size << 30;
STR |= AddrReg << 5;
STR |= DataReg;
PC[-1] = DMB; // Back-patch the half-barrier.
PC[0] = STR;
ClearICache(&PC[-1], 16);
// Back up one instruction and have another go
return std::make_pair(true, -4);
}
}
else if ((Instr & RCPC2_MASK) == LDAPUR_INST) { // LDAPUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
else if ((Instr & RCPC2_MASK) == LDAPUR_INST) { // LDAPUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ArchHelpers::Arm64::HandleAtomicLoad(Instr, GPRs, Offset)) {
// Skip this instruction now
return std::make_pair(true, 4);
@@ -2104,23 +2082,9 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return NotHandled;
}
}
else {
uint32_t LDUR = 0b0011'1000'0100'0000'0000'0000'0000'0000;
LDUR |= Size << 30;
LDUR |= AddrReg << 5;
LDUR |= DataReg;
LDUR |= Instr & (0b1'1111'1111 << 9);
PC[0] = LDUR;
PC[1] = DMB_LD; // Back-patch the half-barrier.
ClearICache(&PC[-1], 16);
// With the instruction modified, now execute again.
return std::make_pair(true, 0);
}
}
else if ((Instr & RCPC2_MASK) == STLUR_INST) { // STLUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ParanoidTSO) {
else if ((Instr & RCPC2_MASK) == STLUR_INST) { // STLUR*
// Extract the 9-bit offset from the instruction
int32_t Offset = static_cast<int32_t>(Instr) << 11 >> 23;
if (ArchHelpers::Arm64::HandleAtomicStore(Instr, GPRs, Offset)) {
// Skip this instruction now
return std::make_pair(true, 4);
@@ -2130,18 +2094,64 @@ static uint64_t HandleAtomicLoadstoreExclusive(uintptr_t ProgramCounter, uint64_
return NotHandled;
}
}
else {
uint32_t STUR = 0b0011'1000'0000'0000'0000'0000'0000'0000;
STUR |= Size << 30;
STUR |= AddrReg << 5;
STUR |= DataReg;
STUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB; // Back-patch the half-barrier.
PC[0] = STUR;
ClearICache(&PC[-1], 16);
// Back up one instruction and have another go
return std::make_pair(true, -4);
}
}
const auto Frame = Thread->CurrentFrame;
const uint64_t BlockBegin = Frame->State.InlineJITBlockHeader;
auto InlineHeader = reinterpret_cast<const CPU::CPUBackend::JITCodeHeader *>(BlockBegin);
auto InlineTail = reinterpret_cast<CPU::CPUBackend::JITCodeTail *>(Frame->State.InlineJITBlockHeader + InlineHeader->OffsetToBlockTail);
// Lock code mutex during any SIGBUS handling that potentially changes code.
// Need to be careful to not read any code part-way through modification.
FEXCore::Utils::SpinWaitLock::UniqueSpinMutex lk(&InlineTail->SpinLockFutex);
if ((Instr & LDAXR_MASK) == LDAR_INST || // LDAR*
(Instr & LDAXR_MASK) == LDAPR_INST) { // LDAPR*
uint32_t LDR = 0b0011'1000'0111'1111'0110'1000'0000'0000;
LDR |= Size << 30;
LDR |= AddrReg << 5;
LDR |= DataReg;
PC[0] = LDR;
PC[1] = DMB_LD; // Back-patch the half-barrier.
ClearICache(&PC[-1], 16);
// With the instruction modified, now execute again.
return std::make_pair(true, 0);
}
else if ( (Instr & LDAXR_MASK) == STLR_INST) { // STLR*
uint32_t STR = 0b0011'1000'0011'1111'0110'1000'0000'0000;
STR |= Size << 30;
STR |= AddrReg << 5;
STR |= DataReg;
PC[-1] = DMB; // Back-patch the half-barrier.
PC[0] = STR;
ClearICache(&PC[-1], 16);
// Back up one instruction and have another go
return std::make_pair(true, -4);
}
else if ((Instr & RCPC2_MASK) == LDAPUR_INST) { // LDAPUR*
// Extract the 9-bit offset from the instruction
uint32_t LDUR = 0b0011'1000'0100'0000'0000'0000'0000'0000;
LDUR |= Size << 30;
LDUR |= AddrReg << 5;
LDUR |= DataReg;
LDUR |= Instr & (0b1'1111'1111 << 9);
PC[0] = LDUR;
PC[1] = DMB_LD; // Back-patch the half-barrier.
ClearICache(&PC[-1], 16);
// With the instruction modified, now execute again.
return std::make_pair(true, 0);
}
else if ((Instr & RCPC2_MASK) == STLUR_INST) { // STLUR*
uint32_t STUR = 0b0011'1000'0000'0000'0000'0000'0000'0000;
STUR |= Size << 30;
STUR |= AddrReg << 5;
STUR |= DataReg;
STUR |= Instr & (0b1'1111'1111 << 9);
PC[-1] = DMB; // Back-patch the half-barrier.
PC[0] = STUR;
ClearICache(&PC[-1], 16);
// Back up one instruction and have another go
return std::make_pair(true, -4);
}
else if ((Instr & ArchHelpers::Arm64::LDAXP_MASK) == ArchHelpers::Arm64::LDAXP_INST) { // LDAXP
//Should be compare and swap pair only. LDAXP not used elsewhere
@@ -23,7 +23,7 @@ bool HandleAtomicMemOp(void *_ucontext, void *_info, uint32_t Instr) {
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
std::pair<bool, int32_t> HandleUnalignedAccess(bool ParanoidTSO, uintptr_t ProgramCounter, uint64_t *GPRs) {
std::pair<bool, int32_t> HandleUnalignedAccess(FEXCore::Core::InternalThreadState *Thread, bool ParanoidTSO, uintptr_t ProgramCounter, uint64_t *GPRs) {
ERROR_AND_DIE_FMT("HandleAtomicMemOp Not Implemented");
}
File renamed without changes.
+30 -2
View File
@@ -17,13 +17,15 @@ class MemberFunctionToPointerCast final {
public:
MemberFunctionToPointerCast(PointerToMemberType Function) {
memcpy(&PMF, &Function, sizeof(PMF));
}
uintptr_t GetConvertedPointer() const {
#ifdef _M_X86_64
// Itanium C++ ABI (https://itanium-cxx-abi.github.io/cxx-abi/abi.html#member-function-pointers)
// Low bit of ptr specifies if this Member function pointer is virtual or not
// Throw an assert if we were trying to cast a virtual member
LOGMAN_THROW_AA_FMT((PMF.ptr & 1) == 0, "C++ Pointer-To-Member representation didn't have low bit set to 0. Are you trying to cast a virtual member?");
#elif defined(_M_ARM_64 )
#elif defined(_M_ARM_64)
// C++ ABI for the Arm 64-bit Architecture (IHI 0059E)
// 4.2.1 Representation of pointer to member function
// Differs from Itanium specification
@@ -31,10 +33,36 @@ class MemberFunctionToPointerCast final {
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#endif
return PMF.ptr;
}
uintptr_t GetConvertedPointer() const {
// Gets the vtable entry position of a virtual member function.
size_t GetVTableOffset() const {
#ifdef _M_X86_64
// Itanium C++ ABI (https://itanium-cxx-abi.github.io/cxx-abi/abi.html#member-function-pointers)
// Low bit of ptr specifies if this Member function pointer is virtual or not
// Throw an assert if we are not loading a virtual member.
LOGMAN_THROW_AA_FMT((PMF.ptr & 1) == 1, "C++ Pointer-To-Member representation didn't have low bit set to 1. This cast only works for virtual members.");
return PMF.ptr & ~1ULL;
#elif defined(_M_ARM_64)
// C++ ABI for the Arm 64-bit Architecture (IHI 0059E)
// 4.2.1 Representation of pointer to member function
// Differs from Itanium specification
LOGMAN_THROW_AA_FMT((PMF.adj & 1) == 1, "C++ Pointer-To-Member representation didn't have adj == 1. This cast only works for virtual members.");
return PMF.ptr;
#else
#error Don't know how to cast Member to function here. Likely just Itanium
#endif
}
// Gets the pointer to the vtable entry for the object passed it.
template<typename Class>
uintptr_t GetVTableEntry(Class *VirtualClass) const {
// VTable is always stored at the beginning of a class object.
uintptr_t *VTable = *reinterpret_cast<uintptr_t**>(VirtualClass);
size_t Offset = GetVTableOffset() / sizeof(void*);
return VTable[Offset];
}
private:
+27
View File
@@ -0,0 +1,27 @@
#include "Utils/SpinWaitLock.h"
namespace FEXCore::Utils::SpinWaitLock {
#ifdef _M_ARM_64
constexpr uint64_t NanosecondsInSecond = 1'000'000'000ULL;
static uint32_t GetCycleCounterFrequency() {
uint64_t Result{};
__asm("mrs %[Res], CNTFRQ_EL0"
: [Res] "=r" (Result));
return Result;
}
static uint64_t CalculateCyclesPerNanosecond() {
// Snapdragon devices historically use a 19.2Mhz cycle counter frequency
// This means that the number of cycles per nanosecond ends up being 52.0833...
//
// ARMv8.6 and ARMv9.1 requires the cycle counter frequency to be 1Ghz.
// This means the number of cycles per nanosecond ends up being 1.
uint64_t CounterFrequency = GetCycleCounterFrequency();
return NanosecondsInSecond / CounterFrequency;
}
uint32_t CycleCounterFrequency = GetCycleCounterFrequency();
uint64_t CyclesPerNanosecond = CalculateCyclesPerNanosecond();
#endif
}
+300
View File
@@ -0,0 +1,300 @@
#include <atomic>
#include <chrono>
#include <mutex>
#include <type_traits>
namespace FEXCore::Utils::SpinWaitLock {
/**
* @brief This provides routines to implement implement an "efficient spin-loop" using ARM's WFE and exclusive monitor interfaces.
*
* Spin-loops on mobile devices with a battery can be a bad idea as they burn a bunch of power. This attempts to mitigate some of the impact
* by putting the CPU in to a lower-power state using WFE.
* On platforms tested, WFE will put the CPU in to a lower power state for upwards of 52ns per WFE. Which isn't a significant amount of time
* but should still have power savings. Ideally WFE would be able to keep the CPU in a lower power state for longer. This also has the added benefit
* that atomics aren't abusing the caches when spinning on a cacheline, which has knock-on powersaving benefits.
*
* FEAT_WFxT adds a new instruction with a timeout, but since the spurious wake-up is so aggressive it isn't worth using.
*
* It should be noted that this implementation has a few dozen cycles of start-up time. Which means the overhead for invoking this implementation is
* slightly higher than a true spin-loop. The hot loop body itself is only three instructions so it is quite efficient.
*
* On non-ARM platforms it is truly a spin-loop, which is okay for debugging only.
*/
#ifdef _M_ARM_64
#define LOADEXCLUSIVE(LoadExclusiveOp, RegSize) \
/* Prime the exclusive monitor with the passed in address. */ \
#LoadExclusiveOp " %" #RegSize "[Result], [%[Futex]];"
#define SPINLOOP_BODY(LoadAtomicOp, RegSize) \
/* WFE will wait for either the memory to change or spurious wake-up. */ \
"wfe;" \
/* Load with acquire to get the result of memory. */ \
#LoadAtomicOp " %" #RegSize "[Result], [%[Futex]]; "
#define SPINLOOP_WFE_LDX_8BIT LOADEXCLUSIVE(ldaxrb, w)
#define SPINLOOP_WFE_LDX_16BIT LOADEXCLUSIVE(ldaxrh, w)
#define SPINLOOP_WFE_LDX_32BIT LOADEXCLUSIVE(ldaxr, w)
#define SPINLOOP_WFE_LDX_64BIT LOADEXCLUSIVE(ldaxr, x)
#define SPINLOOP_8BIT SPINLOOP_BODY(ldarb, w)
#define SPINLOOP_16BIT SPINLOOP_BODY(ldarh, w)
#define SPINLOOP_32BIT SPINLOOP_BODY(ldar, w)
#define SPINLOOP_64BIT SPINLOOP_BODY(ldar, x)
extern uint32_t CycleCounterFrequency;
extern uint64_t CyclesPerNanosecond;
///< Get the raw cycle counter which is synchronizing.
/// `CNTVCTSS_EL0` also does the same thing, but requires the FEAT_ECV feature.
static inline uint64_t GetCycleCounter() {
uint64_t Result{};
__asm volatile(R"(
isb;
mrs %[Res], CNTVCT_EL0;
)"
: [Res] "=r" (Result));
return Result;
}
///< Converts nanoseconds to number of cycles.
/// If the cycle counter is 1Ghz then this is a direct 1:1 map.
static inline uint64_t ConvertNanosecondsToCycles(std::chrono::nanoseconds const &Nanoseconds) {
const auto NanosecondCount = Nanoseconds.count();
return NanosecondCount / CyclesPerNanosecond;
}
static inline uint8_t LoadExclusive(uint8_t *Futex) {
uint8_t Result{};
__asm volatile(SPINLOOP_WFE_LDX_8BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint16_t LoadExclusive(uint16_t *Futex) {
uint16_t Result{};
__asm volatile(SPINLOOP_WFE_LDX_16BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint32_t LoadExclusive(uint32_t *Futex) {
uint32_t Result{};
__asm volatile(SPINLOOP_WFE_LDX_32BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint64_t LoadExclusive(uint64_t *Futex) {
uint64_t Result{};
__asm volatile(SPINLOOP_WFE_LDX_64BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint8_t WFELoadAtomic(uint8_t *Futex) {
uint8_t Result{};
__asm volatile(SPINLOOP_8BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint16_t WFELoadAtomic(uint16_t *Futex) {
uint16_t Result{};
__asm volatile(SPINLOOP_16BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint32_t WFELoadAtomic(uint32_t *Futex) {
uint32_t Result{};
__asm volatile(SPINLOOP_32BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
static inline uint64_t WFELoadAtomic(uint64_t *Futex) {
uint64_t Result{};
__asm volatile(SPINLOOP_64BIT
: [Result] "=r" (Result)
, [Futex] "+r" (Futex)
:: "memory");
return Result;
}
template<typename T, typename TT = T>
static inline void Wait(T *Futex, TT ExpectedValue) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
// Early exit if possible.
if (Result == ExpectedValue) return;
do {
Result = LoadExclusive(Futex);
if (Result == ExpectedValue) return;
Result = WFELoadAtomic(Futex);
} while (Result != ExpectedValue);
}
template
void Wait<uint8_t>(uint8_t*, uint8_t);
template
void Wait<uint16_t>(uint16_t*, uint16_t);
template
void Wait<uint32_t>(uint32_t*, uint32_t);
template
void Wait<uint64_t>(uint64_t*, uint64_t);
template<typename T, typename TT>
static inline bool Wait(T *Futex, TT ExpectedValue, std::chrono::nanoseconds const &Timeout) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
// Early exit if possible.
if (Result == ExpectedValue) return true;
const auto TimeoutCycles = ConvertNanosecondsToCycles(Timeout);
const auto Begin = GetCycleCounter();
do {
Result = LoadExclusive(Futex);
if (Result == ExpectedValue) return true;
Result = WFELoadAtomic(Futex);
const auto CurrentCycleCounter = GetCycleCounter();
if ((CurrentCycleCounter - Begin) >= TimeoutCycles) {
// Couldn't get value before timeout.
return false;
}
} while (Result != ExpectedValue);
// We got our result.
return true;
}
template
bool Wait<uint8_t>(uint8_t*, uint8_t, std::chrono::nanoseconds const &);
template
bool Wait<uint16_t>(uint16_t*, uint16_t, std::chrono::nanoseconds const &);
template
bool Wait<uint32_t>(uint32_t*, uint32_t, std::chrono::nanoseconds const &);
template
bool Wait<uint64_t>(uint64_t*, uint64_t, std::chrono::nanoseconds const &);
#else
template<typename T, typename TT>
static inline void Wait(T *Futex, TT ExpectedValue) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
// Early exit if possible.
if (Result == ExpectedValue) return;
do {
Result = AtomicFutex->load();
} while (Result != ExpectedValue);
}
template<typename T, typename TT>
static inline bool Wait(T *Futex, TT ExpectedValue, std::chrono::nanoseconds const &Timeout) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Result = AtomicFutex->load();
// Early exit if possible.
if (Result == ExpectedValue) return true;
const auto Begin = std::chrono::high_resolution_clock::now();
do {
Result = AtomicFutex->load();
const auto CurrentCycleCounter = std::chrono::high_resolution_clock::now();
if ((CurrentCycleCounter - Begin) >= Timeout) {
// Couldn't get value before timeout.
return false;
}
} while (Result != ExpectedValue);
// We got our result.
return true;
}
#endif
template<typename T>
static inline void lock(T *Futex) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Expected{};
T Desired {1};
// Try to CAS immediately.
if (AtomicFutex->compare_exchange_strong(Expected, Desired)) return;
do {
// Wait until the futex is unlocked.
Wait(Futex, 0);
Expected = 0;
} while (!AtomicFutex->compare_exchange_strong(Expected, Desired));
}
template<typename T>
static inline bool try_lock(T *Futex) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
T Expected{};
T Desired {1};
// Try to CAS immediately.
if (AtomicFutex->compare_exchange_strong(Expected, Desired)) return true;
return false;
}
template<typename T>
static inline void unlock(T *Futex) {
std::atomic<T> *AtomicFutex = reinterpret_cast<std::atomic<T>*>(Futex);
AtomicFutex->store(0);
}
#undef SPINLOOP_8BIT
#undef SPINLOOP_16BIT
#undef SPINLOOP_32BIT
#undef SPINLOOP_64BIT
template<typename T>
class UniqueSpinMutex final {
public:
UniqueSpinMutex(T *Futex)
: Futex {Futex} {
FEXCore::Utils::SpinWaitLock::lock(Futex);
}
~UniqueSpinMutex() {
FEXCore::Utils::SpinWaitLock::unlock(Futex);
}
private:
T *Futex;
};
}
+1 -1
View File
@@ -2,7 +2,7 @@
---
The FEXCore frontend's job is to translate an incoming x86-64 instruction stream in to a more easily digested version of x86.
This effectively expands x86-64 instruction encodings to be more easily ingested later on in the process.
This ends up being essential to allowing our IR translation step to be less strenious. It can decode a "common" expanded instruction format rather than various things that x86-supports.
This ends up being essential to allowing our IR translation step to be less strenuous. It can decode a "common" expanded instruction format rather than various things that x86-supports.
For a simple example, x86-64's primary op table has ALU ops that duplicate themselves at least six times with minor differences between each. The frontend is able to decode a large amount of these ops to the "same" op that the IR translation understands more readily.
This works for most instructions that follow a common decoding scheme, although there are instructions that don't follow the rules and must be handled explicitly elsewhere.
+1 -1
View File
@@ -32,7 +32,7 @@ BeginBlock
%360 i64 = LoadContext 0x8, 0x58
%392 i64 = LoadContext 0x8, 0x48
%424 i64 = LoadContext 0x8, 0x50
%456 i64 = Syscall%232, %264, %296, %328, %360, %392, %424
%456 i64 = Syscall %232, %264, %296, %328, %360, %392, %424
StoreContext 0x8, 0x8, %456
BeginBlock
EndBlock 0x1e
Loaded 100 of 396 files, more files were not shown because too many files have changed in this diff. Show more