Compare commits

...
548 Commits
Author SHA1 Message Date
Ryan Houdek aa0f2c3975 Docs: Update for release FEX-2406 2024-06-12 18:41:54 -07:00
Tony Wasserka 4dc8648d81 Merge pull request #3630 from neobrain/feature_libfwd_gl32
Library Forwarding: Add support for 32-bit OpenGL
2024-06-11 17:33:29 +02:00
Tony Wasserka 7a703e1176 Library Forwarding/GL: Enable stricter pointer parameter checks 2024-06-11 17:14:23 +02:00
Tony Wasserka 02df1a2924 Library Forwarding/GL: Avoid pointer array repacking for 64-bit guests 2024-06-11 17:14:23 +02:00
Tony Wasserka efac7efc97 Library Forwarding/GL: Remap _XDisplay returned by glXGetCurrentDisplay 2024-06-11 17:14:23 +02:00
Tony Wasserka d99b4a80c8 Library Forwarding/GL: Add glX support for 32-bit 2024-06-11 17:14:23 +02:00
Tony Wasserka 2a76744d30 Library Forwarding/GL: Add 32-bit support for most core GL APIs 2024-06-11 17:13:45 +02:00
Tony Wasserka 843b2d1969 Library Forwarding/GL: Assume void* always points to compatible data 2024-06-11 16:58:37 +02:00
Tony Wasserka b275c96889 Library Forwarding: Use the fixed-size guest type for passthrough parameters 2024-06-11 16:58:36 +02:00
Ryan Houdek 033b1ce449 Merge pull request #3688 from catfella/extend_bextr_tests
unittests/bextr: add SrcSize tests
2024-06-09 23:18:57 -07:00
Mikhail Nitenko 99a43283be unittests/bextr: add SrcSize tests
dougallj mentioned that adding these tests might expose
a bug in bextr. Since bextr implementation was changed
apparently it now works correctly, that's good.
2024-06-10 05:45:12 +00:00
Mai 55bfd6394b Merge pull request #3640 from Sonicadvance1/cleanup_execve_envp
LinuxSyscalls: Cleanup envp copying in execve
2024-06-07 21:40:13 -04:00
Ryan Houdek 14bfe6016e Merge pull request #3684 from alyssarosenzweig/constprop/cleanup
Constprop: clean up
2024-06-04 13:00:32 -07:00
Alyssa Rosenzweig 0d4ad70875 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig a8bf3859ea ConstProp: rm pointless constant folding
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig aa7dcffcea ConstProp: drop const pool heuristic
slightly worse for compile time, slightly better output, honestly I'll take the
win because this is easier to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig be1a5cea8e ConstProp: drop addressgen const pool stuff
I don't get the point, it should be handled by a combination of existing
passes/techniques just fine. no instcountci changes.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 402ea84aa0 RedundantFlagCalculationElimination: cleanup DCE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 19a7b06b91 ConstProp: swallow up LongDivideElimination
as usual.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 96bd643e5b ConstProp: always inline constants
x86/interpreter leftover, I think.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 6b9293979c ConstProp: swallow up InlineCallOptimization
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 7d5cee4384 InlineCallOptimization: rm x86 leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig c0bab70161 Merge pull request #3682 from alyssarosenzweig/ir/ref
Find-and-replace OrderedNode* with Ref
2024-06-04 10:09:36 -04:00
Alyssa Rosenzweig 32f5a28433 IR: use Ref instead of OrderedNode
find-and-replace across the tree, excluding IR.h itself.

also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig ce30179ed1 IR: add Ref typedef
To put new IR lipstick on the old IR pig.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig a515b707f3 Merge pull request #3679 from Sonicadvance1/memory_model_emulation_programmer_documentation
FEXCore/docs: Adds programmer documentation about memory model emulation
2024-06-03 09:24:37 -04:00
Ryan Houdek 9ab0fa01bd Merge pull request #3681 from alyssarosenzweig/opt/shld
optimize shld
2024-06-01 11:53:42 -07:00
Alyssa Rosenzweig c3bffa2929 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 14:44:24 -04:00
Alyssa Rosenzweig 951fee361f OpcodeDispatcher: optimize shld
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 14:44:24 -04:00
Ryan Houdek 8c4860b9a7 Merge pull request #3680 from alyssarosenzweig/opt/sib 2024-06-01 11:17:54 -07:00
Ryan Houdek ee221e6a8c Syscalls: Removes unnecessary lambda that was only called once.
Local refactor.
2024-06-01 11:08:34 -07:00
Alyssa Rosenzweig f5625093bb InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:50 -04:00
Alyssa Rosenzweig abfd974d70 OpcodeDispatcher: select hardware addressing modes
Now that we have a framework to do this in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:50 -04:00
Alyssa Rosenzweig 97966930e9 OpcodeDispatcher/x87f64: fuse addr calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig a52a2e3ae4 OpcodeDispatcher/x87: fuse addr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig c49b30f105 OpcodeDispatcher/Vector: fuse addr calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig b0b4ad2083 OpcodeDispatcher: fuse xlat address
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig ee4bee4fef OpcodeDispatcher: fuse BT address
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig c3a0f5a2f6 OpcodeDispatcher: fuse sgdt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig 0413a6bf68 OpcodeDispatcher: improve bmi2 shift
allow upper garbage, use simpler clean.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig 7bd036d1ae OpcodeDispatcher: refactor address modes
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:32 -04:00
Alyssa Rosenzweig 112c49a348 ConstProp: fix inlining shifted imm to mem instructions
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 80878ae611 ConstProp: rework mem immediate inlining
deduplicate all the things.

functional change:
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 85a69be5b6 ConstProp: drop address fusion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:38:03 -04:00
Ryan Houdek 8dbfd1635a FEXCore/docs: Adds programmer documentation about memory model emulation
I keep needing to look these up to remember the limitations. Add a doc
file so I can more easily point to the information.
2024-05-31 10:36:48 -07:00
Alyssa Rosenzweig 8b5ca303e3 JIT: add asserts for invalid TSO load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 12:12:36 -04:00
Ryan Houdek f90d2aeb6d Merge pull request #3678 from marysaka/fix/test-harness-runner-ags-segfault
Fix segfault when starting TestHarnessRunner with missing arguments
2024-05-31 08:06:54 -07:00
Mary Guillemard cd4b520f72 Fix segfault when starting TestHarnessRunner with missing arguments
Signed-off-by: Mary Guillemard <mary@mary.zone>
2024-05-31 16:54:31 +02:00
Ryan Houdek 20d5a26a72 Merge pull request #3674 from alyssarosenzweig/opt/logical-flags
Optimize logical flags
2024-05-30 12:11:21 -07:00
Alyssa Rosenzweig 9346116485 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-30 14:42:29 -04:00
Alyssa Rosenzweig bb8336fcad OpcodeDispatcher: optimize logical flags
fuse the PF write in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-30 14:42:22 -04:00
Ryan Houdek ee96d60983 Merge pull request #3673 from alyssarosenzweig/ra/tied
Track tied sources in the IR
2024-05-30 10:55:15 -07:00
Alyssa Rosenzweig 6052b335dc Merge pull request #3666 from Sonicadvance1/fix_initial_darwinia
FileManagement: Fix fstatat/statx with self and NOFOLLOW
2024-05-29 23:11:24 -04:00
Ryan Houdek ab0a6bbe9f Merge pull request #3669 from Sonicadvance1/fix_addshift_operation
ConstProp fixes for Darwinia
2024-05-29 19:43:13 -07:00
Ryan Houdek 9dd6d8ed94 Merge pull request #3639 from Sonicadvance1/cleanupFD
FEXLoader: Cleanup FD extraction from environment variables
2024-05-29 19:18:59 -07:00
Ryan Houdek 3b5d0e3e27 FEXLoader: Cleanup FD extraction from environment variables
In preparation for seccomp execve inheritance where we need to extract
another FD from a different environment variable.

- Small function to extract the FD and also unset the environment
  variable in the same place.
   - Keeping the fetch and unset together instead of spreading to
     another location in the source.
- Extract the FD upfront instead of passing the string_view around,
  since we are unsetting the environment variable at the same place.

Future seccomp inheritance will get the FD just after the FEXFD
   - `int FEXSeccompFD {GetFEXFDFromEnv("FEX_SECCOMPFD")};`
2024-05-29 18:47:28 -07:00
Ryan Houdek 37e13cf073 FileManagement: Fix fstatat with self and NOFOLLOW
When asked to not follow the symlink, FEX needs to return data about the
symlink itself rather than following to the target executable. In that
case we need to return symlink information otherwise games that sanity
check can break.

This is what happened with Darwinia in #3662.

We return the FEXInterpreter symlink information in this case since it
doesn't return any information that is relevent to leaking emulator
state. Once the application asks to follow through to the symlink target
is when we will replace.

Also adds a unit test to ensure we don't break it.
2024-05-29 18:41:24 -07:00
Alyssa Rosenzweig 9b1b9c26cc Merge pull request #3664 from Sonicadvance1/change_timestamp
FEXLogging: Changes representation of timestamp
2024-05-29 16:44:42 -04:00
Ryan Houdek f7f3024b92 unittests/ASM: Adds SIB transpose scale register test
With a bit of pointer math it will choose the incorrect address if the
base and offset registers were transposed.
2024-05-29 11:41:20 -07:00
Alyssa Rosenzweig 11ec71a4ce InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 665491adf8 OpcodeDispatcher: drop weird !flagm special case
now that bfi is coalesced, this is a win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 55391ccbc0 RegisterAllocationPass: try to coalesce tied sources
we'll do better in the future but this is already a win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 7790d7a0b7 IR: track tied sources
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Ryan Houdek f3d8c2cbac unittests/ASM: Adds unittest for bug encountered in Darwinia 2024-05-29 07:29:49 -07:00
Ryan Houdek 1226069b4c InstCountCI: Update for fixes
Only prefetch hit currently since ConstProp is limited to optimizing the
ADD IROp atm.
2024-05-29 04:42:40 -07:00
Ryan Houdek 80687c8d2d ConstProp: Limits which addressing modes can be used for vector loadstores
This was causing us to generate invalid code in Darwinia, resulting in a
crash. With assertions enabled this would be picked up in the emitter.

Only implement AddShift optimizations for now because I don't want to do
the remaining optimizations in a bug fix PR.

Fixes Darwinia.
2024-05-29 04:42:11 -07:00
Ryan Houdek 920fe60492 ConstProp: Fix bug with transposed elements from AddShift op
Accidentally we were swapping which sources were the base and which was
the one getting shifted. This wasn't super common so it usually didn't
matter.

Fixes one crash in Darwinia.
2024-05-29 04:32:51 -07:00
Ryan Houdek 8c6ce2cb3b Passes/ConstProp: Have MemExtendedAddressing return a struct rather than a tuple
Makes it less confusing about which variable is the base versus the
offset.

NFC
2024-05-29 04:32:14 -07:00
Ryan Houdek 61f30d004c IR: Document AddShift behaviour
Just to clarify that Src2 is the shifted operation.
2024-05-29 04:29:09 -07:00
Ryan Houdek 95919a1ddf InstcountCI: Add addressing limit tests for base + offset<<shift
These need to be tested.
2024-05-29 04:28:24 -07:00
Ryan Houdek 35ec54f920 Merge pull request #3667 from alyssarosenzweig/opt/pcmp
Optimize PCMPESTRI flags a bit
2024-05-28 22:06:37 -07:00
Alyssa Rosenzweig 32e8a56093 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:32:14 -04:00
Alyssa Rosenzweig 136f1d0a0b OpcodeDispatcher: drop pcmpestri zext
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:32:14 -04:00
Alyssa Rosenzweig 0c042d1e85 VectorFallbacks: optimize PCMP*STRI flags
Return an NZCV.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:19:43 -04:00
Alyssa Rosenzweig ad13442be4 Merge pull request #3665 from Sonicadvance1/sse42_instcountci
InstCountCI: Adds SSE4.2 operations
2024-05-28 08:57:45 -04:00
Ryan Houdek d6b9252760 InstCountCI: Adds SSE4.2 operations
Doesn't handle all 127 combinations of the control immediate for all
four instructions. Although supplies the instruction control instruction
that A Hat in Time abuses heavily.

The SSE2 implementation of the function in vcruntime140 is likely faster
than our currently implementation but we should be able to get something
comparable. Not bad considering this is a required extension and this is
the first game we found that abuses the instruction heavily.
2024-05-28 00:55:32 -07:00
Ryan Houdek 22222ebaf5 FEXLogging: Changes representation of timestamp
This was a bit confusing to read and I had always expected to change
this at some point.

Previous:
```
[INFO][1579518391560577][1601857.1601857] clone: Unsupported flags w/o CLONE_THREAD (Shared Resources), 4100
```

Now:
```
[INFO][1590468.992593376][1629501.1629501] clone: Unsupported flags w/o CLONE_THREAD (Shared Resources), 4100
```
2024-05-27 23:36:58 -07:00
Alyssa Rosenzweig 734258e23b Merge pull request #3661 from Sonicadvance1/remove_warnings2
Removes warnings
2024-05-25 11:52:37 -04:00
Ryan Houdek 74916b3757 RAPass: Remove warnings 2024-05-24 18:41:30 -07:00
Ryan Houdek c5359264a3 VixlUtils: Remove warnings 2024-05-24 18:41:19 -07:00
Ryan Houdek 9d0ff7929e Merge pull request #3660 from alyssarosenzweig/opt/smash-less
Delete a big chunk of IR/Passes/*
2024-05-24 16:23:48 -07:00
Alyssa Rosenzweig d3eed27d17 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:45:32 -04:00
Alyssa Rosenzweig bc1669b163 DeadStoreElimination: eliminate map
use a vec. block indices will be dense in the new IR. This is memory intensive
but seems faster in practice.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 83e417b2c6 DeadStoreElimination: combine GPR/FPR handling
slight speed up per profile.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cb00d9171f IR: merge general DCE with flag DCE
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.

Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cf77f2ae5d RedundantFlagCalculationElimination: fix convergence issue
If both the destination and the flags are dead for an AddWithFlags, we need to
eliminate it in one pass. If we only replace without elimiating, we would need a
second DCE pass to eliminate. We want DCE to finish in one pass, so fix this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 273d086a7b ConstProp: merge const pooling passes
walk the IR less.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 94d9cf54bc ConstProp: don't push/pop cursor
pointless

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3089e0e6de ConstProp: merge masking opts with const folding
Single pass over the IR now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3c088fb414 ConstProp: remove masking elimination opts
This has been deadcode since 2020. Drop it so we can focus on what *does* work
and what does matter.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 676c9a9be6 IR: do not return progress from passes
Generally, there are three reasons to track progress:

* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.

None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Ryan Houdek 314fea36b4 Merge pull request #3658 from Sonicadvance1/enable_afp 2024-05-24 12:23:29 -07:00
Ryan Houdek 3b7d30d26a Merge pull request #3637 from alyssarosenzweig/ra/mr 2024-05-24 12:14:17 -07:00
Alyssa Rosenzweig a8d32b9a2f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:26:02 -04:00
Alyssa Rosenzweig 24cb02f4ff FEXCore: remove IRCompaction
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 725d0e187a RegisterAllocationPass: rewrite RA
I recommend viewing the new source file as the diff is quite messy.

---

The old RA commits every "how not to write an RA" sin in the book.

Chaitin spill-one loop? Check.

Potential spilling caused by alignment issues since there's no live range
splitting? Check.

Panic spilling? Check.

Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.

...

It is wholly unsuitable to any application, and specifically unsuitable for FEX.

---

The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.

Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.

The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.

The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.

That's it :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig d8603cb9bd RAValidation: weaken fill validation
Fails with the new RA when copies are inserted, since RAValidation loses
visibility. Nontrivial to fix, but this particular assert hopefully isn't buying
us a ton. Weaken it for now, we can revisit later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 04c2cb5feb IR: add GPR copy/swap instructions
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.

Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Ryan Houdek 32f2decb24 Merge pull request #3656 from alyssarosenzweig/opt/asr-masking
Optimize asr
2024-05-23 21:35:48 -07:00
Ryan Houdek 6954ebe3a0 Merge pull request #3651 from Sonicadvance1/change_default_tso
Config: Change default TSO options
2024-05-23 21:35:33 -07:00
Ryan Houdek 2e40da3d6b HostFeatures: Enable AFP and RPRES
This has been investigated. Theoretically should work.
2024-05-23 13:26:05 -07:00
Alyssa Rosenzweig 559772bb03 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig 2bbcf72e27 OpcodeDispatcher: optimize asr masking
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig aa3a92aa60 OpcodeDispatcher: merge asr impls
so we can DRY the next patch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Mai 5497240a25 Merge pull request #3655 from alyssarosenzweig/opt/wacky-imul
Optimize large sign-extended constants
2024-05-22 23:27:09 -04:00
Alyssa Rosenzweig 79c609d0f5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:53:19 -04:00
Alyssa Rosenzweig b33788e765 Arm64Emitter: handle another class of constants
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:50:20 -04:00
Alyssa Rosenzweig 8340012466 InstCountCI: add wacky imul
found in bytemark, exposes an interesting constant case

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:50:19 -04:00
Ryan Houdek 0adcc779cf Merge pull request #3654 from alyssarosenzweig/opt/movsx
Optimize sign-extension
2024-05-22 19:29:51 -07:00
Alyssa Rosenzweig 9d0718fbc4 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Alyssa Rosenzweig 53567a6526 OpcodeDispatcher: optimize movsxd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Alyssa Rosenzweig 90fb5f038b OpcodeDispatcher: optimize movsx
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Mai 063b1eb936 Merge pull request #3652 from alyssarosenzweig/instcountci/more-bytemark
InstructionCountCI: add bytemark hot block
2024-05-22 14:20:13 -04:00
Alyssa Rosenzweig 9f243c8f7b InstructionCountCI: add bytemark hot block
this was basically practice for using perf with FEX, but a few things do stand
out in the assembly as suboptimal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 14:09:50 -04:00
Ryan Houdek 2dc600f283 Config: Change default TSO options
After two months of testing I finally have enough confidence that these
default options getting changed is safe enough that most games won't
notice the difference. But the performance differences can be wild.

Might want to come back and reenable this on platforms that support
hardware TSO and LRCPC3 (for the vector feature), but we can care once
those platforms come up to speed.
2024-05-21 18:08:58 -07:00
Ryan Houdek a01402d502 Merge pull request #3650 from alyssarosenzweig/unittests/xess
unittests: add XeSS test
2024-05-21 17:15:08 -07:00
Ryan Houdek c90036aeea Merge pull request #3649 from alyssarosenzweig/ra/validate-less-hard
Simplify/fix our validation passes
2024-05-21 17:12:49 -07:00
Alyssa Rosenzweig dfb751eea0 unittests: add XeSS test
Useful smoke test for the asymptoptic behaviour of our constant pooling code.
Not useful for correctness testing, so skip in CI.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:41:47 -04:00
Alyssa Rosenzweig 6e0f5eccb3 RAValidation: fix spillregister validation
It doesn't write to its node. Fixes spurious

  %7: Arg[0] expects reg0 to contain %4, but it actually contains %16

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 579fb42458 RAValidation: allow filling the same slot multiple times
This can be a reasonable thing to do!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig f4b487352c RAValidation: defeature control flow analysis
now that we've eliminated cross block liveness, we can do our validation locally
too for a massive simplification.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 4448f84f29 IRValidation: merge in ValueDominanceValidation
All we actually need to validate is that each source has been previously defined
within the block. That checks everything we care about now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 9e1e602e09 BitSet: fix memset/memclear logic
Missing a factored of 4, causing a buffer overflow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:32:54 -04:00
Ryan Houdek ca70e387ec Merge pull request #3648 from alyssarosenzweig/ra/pair-extract
Slightly improve pair coalescing + memcpy fix from RA branch
2024-05-21 16:19:35 -07:00
Ryan Houdek 9a483107e3 Merge pull request #3647 from alyssarosenzweig/ir/pass-simpler
ConstProp, RCLSE: simplifications
2024-05-21 16:11:36 -07:00
Ryan Houdek 3bac767866 Merge pull request #3645 from neobrain/refactor_aotir
AOTIR: Refactor interfaces to clarify ownership flow
2024-05-21 16:03:06 -07:00
Ryan Houdek 7b4e48480b Merge pull request #3646 from alyssarosenzweig/opt/minor-disp
OpcodeDispatcher: eliminate some Bfe's
2024-05-21 15:57:52 -07:00
Alyssa Rosenzweig 101bba4808 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig a31c3c1c15 OpcodeDispatcher: use ExtractPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 1a18e392f8 OpcodeDispatcher: add ExtractPair helper
terser and will aid coalescing, as well as eventual transition to multidest
extracts which is what we'll actually want.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 4c7595c68a JIT: allow coalescing ExtractElementPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 1a467f0ebd IR: reduce memcpy worstcase reg pressure
This avoids a bunch of sharp edges for RA at a small cost when obscure
segment registers are used.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:54:06 -04:00
Alyssa Rosenzweig 06e7360f4c RCLSE: only run once, do not DCE
From a theoretical perspective, we should not need to run RCLSE more than once.
If there are convergence issues with the current implementation, they should be
fixed instead of bandaged around. Fortunately, this has no instcountci changes.

Brings RCLSE cost down from like 12% to 5%.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig ec3b72e17e ConstProp: drop select folding
no instcountci changes.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig 259e1b75a4 ConstProp: rm printfs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig f9642cba7a ConstProp: rm pointless opcode casts
Just use the headers directly

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig c01c415030 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:40:59 -04:00
Alyssa Rosenzweig 50e56358c3 OpcodeDispatcher: eliminate Bfe's with cmpxchg
ConstProp was catching these but they're pointless.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:39:00 -04:00
Alyssa Rosenzweig 465dbc260f OpcodeDispatcher: eliminate Bfe's with lea
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:39:00 -04:00
Tony Wasserka 9f291f3adb AOTIR: Remove obsolete fields from serialized data 2024-05-21 17:54:28 +02:00
Tony Wasserka f2acc3da4c Core: Clean up ownership management for IRListView and RegisterAllocationData
IRListView is now purely a view type. Instead, ownership is managed on-demand
by a separate interface (IRStorageBase). Materialization of IRListViews to
owning types is moved to this interface as well.

This also avoids unneeded copies of the data.
2024-05-21 17:51:17 +02:00
Tony Wasserka f8f165d96d AOTIR: Remove redundant variable declarations
This code was already using structured bindings anyway and just reassigned
the values to different variables.
2024-05-21 17:38:41 +02:00
Tony Wasserka baef95992c AOTIR: Drop unneeded local variable 2024-05-21 17:38:41 +02:00
Tony Wasserka 97e18c8469 AOTIR: Drop effectively unused parameter from PreGenerateIRFetch 2024-05-21 17:38:41 +02:00
Tony Wasserka 6e04f7368b AOT: Use std::optional to replace a validity boolean in PreGenerateIRFetch 2024-05-21 17:38:41 +02:00
Tony Wasserka 55a835ebb8 AOTIR: Clarify serialization code
The comments weren't too helpful. Using the struct types directly conveys the
same information more clearly.
2024-05-21 17:38:41 +02:00
Alyssa Rosenzweig 85776c2537 Merge pull request #3643 from alyssarosenzweig/opt/shift-garbage
Allow garbage on more shifts
2024-05-21 11:05:47 -04:00
Ryan Houdek e3ec25d9db Merge pull request #3638 from alyssarosenzweig/jit/dedupe-vec
JIT/VectorOps: deduplicate common implementations
2024-05-20 14:44:52 -07:00
Alyssa Rosenzweig ebfcc1e835 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Alyssa Rosenzweig 769b2c2a46 OpcodeDispatcher: allow garbage on shift dests
doesn't matter for left shifts (we mask off the garbage), or 32-bit shifts, or
shifts where we explicitly sbfe after.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Alyssa Rosenzweig 3b2100307e OpcodeDispatcher: allow garbage on more shifts
we're masking anyway

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Ryan Houdek 28cc179214 LinuxSyscalls: Cleanup envp copying in execve
In preparation for seccomp execve inheritance.

We are going to need to add a new environment variable earlier in the
execve sequence to handle inheritance in the case of binfmt_misc.

No functional change in regards to envp handling.

Minor change around execveat with FD without binfmt_misc. In the case
that execveat returned an error and we did a `dup` of the FD then we
would have an FD leak. Make sure to close the duplicated FD in that
instance.
2024-05-20 07:27:37 -07:00
Alyssa Rosenzweig 3c0f243a2d JIT: dedupe MapCC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:56 -04:00
Alyssa Rosenzweig bbf1563f80 JIT/ALUOps: extract DEF_BINOP_WITH_CONSTANT
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:56 -04:00
Alyssa Rosenzweig ed6b1011f9 JIT/VectorOps: deduplicate common implementations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:55 -04:00
Ryan Houdek e3e7f0279c Merge pull request #3644 from alyssarosenzweig/clang-format/left
clang-format: left-align escaped newlines
2024-05-20 07:12:50 -07:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Ryan Houdek 663f3d8b5a Merge pull request #3641 from Sonicadvance1/instcountci_flake
InstCountCI: Hardcode the offset to load tests into
2024-05-20 06:45:47 -07:00
Ryan Houdek b1f7be2f6c InstCountCI: Update 2024-05-18 18:08:38 -07:00
Ryan Houdek b83cbcb33c ConstProp: Bandage fix for instcountci
Fixed offset x86 code doesn't quite solve the issue, so adjust this
heuristic just to get instcounci to stop flaking.

This code is going to heavily change soon anyway so +50 doesn't change
much.
2024-05-18 18:06:25 -07:00
Ryan Houdek ac1a096bae InstCountCI: Hardcode the offset to load tests into
Depending on where the assembly was getting loaded in to memory it was
causing slight code generation differences.

Map the entire file to the same fixed offset as our ASM tests to ensure
consistency and removing flakes in CI.
2024-05-18 17:00:28 -07:00
Ryan Houdek 048c8ded88 Merge pull request #3622 from Sonicadvance1/move_emitter 2024-05-17 10:41:51 -07:00
Alyssa Rosenzweig 948938bf4b Merge pull request #3636 from alyssarosenzweig/jit/factor-vec
JIT: factor out sub reg size conversion
2024-05-17 09:40:32 -04:00
Alyssa Rosenzweig bb064c7334 JIT/AtomicOps: factor out elementsize
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:28:58 -04:00
Alyssa Rosenzweig 2d3d49b900 JIT/ConversionOps: use ConvertSubRegSize*
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:18:59 -04:00
Alyssa Rosenzweig ea7096ed5b JIT/MemoryOps: use ConvertSubRegSize8
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:16:12 -04:00
Alyssa Rosenzweig 7a0f6c0a80 JIT: factor ConvertSize helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:13:00 -04:00
Alyssa Rosenzweig e4ee35a925 JIT: factor out sub reg size conversion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 15:56:05 -04:00
Ryan Houdek d3ab9bdef6 Remove Float16
We aren't using it. We won't be using it. We need unit tests in our
lives if we want this.
2024-05-16 12:06:54 -07:00
Ryan Houdek 926eefc86c Merge pull request #3635 from alyssarosenzweig/opt/flag-store
OpcodeDispatcher: reorder some moves
2024-05-16 10:58:40 -07:00
Ryan Houdek 3eb7a5b998 Merge pull request #3632 from pmatos/RemoveBlocks
Use erase-remove idiom to remove element
2024-05-16 10:50:50 -07:00
Ryan Houdek 58614ff131 Merge pull request #3634 from alyssarosenzweig/constprop/leftover
ConstProp: remove x86 jit leftover
2024-05-16 10:48:31 -07:00
Alyssa Rosenzweig bf3a09e5e3 ConstProp: remove x86 jit leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 09:21:41 -04:00
Alyssa Rosenzweig 83c536c47f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 08:32:11 -04:00
Alyssa Rosenzweig 7b39e57e72 OpcodeDispatcher: defer overwritten store
this can save moves, as it's a bit easier to reason about the live ranges.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 08:31:58 -04:00
Alyssa Rosenzweig 5bedf32666 Merge pull request #3633 from pmatos/UndefShift
Fix left shift undefined behaviour
2024-05-16 07:45:38 -04:00
Paulo Matos 1eb7be9870 Check BitOffset instead of zeroext 2024-05-16 10:50:40 +02:00
Paulo Matos 93e4288c57 Fix left shift undefined behaviour
Here BitOffset can have values higher than 32.
2024-05-16 10:22:03 +02:00
Paulo Matos 9ca4868833 Use erase-remove idiom to remove element
Fixes #3631
2024-05-16 10:16:01 +02:00
Alyssa Rosenzweig c5f8ea58e9 Merge pull request #3629 from pmatos/Unsup-typo
NFC: Fix typo
2024-05-15 10:31:38 -04:00
Paulo Matos 5bee17bee1 NFC: Fix typo 2024-05-15 15:10:00 +02:00
Ryan Houdek efe7c54374 Merge pull request #3625 from Sonicadvance1/restricted_inst
FEXCore: Fixes the difference between CPL-0 and undefined instructions
2024-05-14 07:11:19 -07:00
Ryan Houdek 9e1840e974 FEXCore: Moves CodeEmitter to FHU
Now that the vixl dependency is gone, this gets moved to FHU since the
frontend is going to need it for a microjit.
2024-05-13 12:48:10 -07:00
Ryan Houdek 1f40590f9a Emitter: Inline IsImmLogical from vixl
The only core vixl usage we use in the emitter. Is a complete pain to
reimplement so keep it around.
2024-05-13 12:48:10 -07:00
Ryan Houdek 64a3bc235d unittests/Emitter: Ensures coverage of imm float encodings
To ensure we round everything correctly for the new float16 class
2024-05-13 12:48:10 -07:00
Ryan Houdek 7d9af246ea CodeEmitter: Removes vixl Float16 usage
Creating local Float16 helper which handles our needs
2024-05-13 12:48:09 -07:00
Ryan Houdek 6d3471bcaa Merge pull request #3627 from Sonicadvance1/timeout_merge_base
Github: Support a timeout on checkout
2024-05-13 11:53:11 -07:00
Ryan Houdek a8714dbd49 Github: Support a timeout on checkout
Sometimes github or the CI runner times out trying to checkout the
source and stays timing out forever.

Give it a three minute timeout otherwise the CI runner will stall
forever.
2024-05-13 11:26:54 -07:00
Ryan Houdek 512312fa06 FEXLinuxTests: Implements a test for the new instructions 2024-05-13 11:12:26 -07:00
Ryan Houdek 010028e381 FEXCore: Fixes the difference between CPL-0 and undefined instructions
undefined instructions are expected to return SIGILL, while implemented
instructions that aren't available in CPL-3 are expected to SIGSEGV.

Noticed this while testing out CPU-Z, it installs a kernel module and
does a bunch of `RDMSR` and `OUTS` instructions. Decided to walk through
the rest of the instructions in the `System Instruction Reference`
section.

Turns out there's a bunch of oddities in there that we don't support.
First step is to go through all the explicitl SIGILL and SIGSEGV and
implement a test for them.

Next step will be implementing the remaining operations that are
considered "System" operations but are still available in CPL-3.
This list includes:
- lar
- lgdt
- lsl
- sidt
- sldt
- stac
- clac
- verr
- verw
2024-05-13 11:12:26 -07:00
Ryan Houdek f27f1871e4 Merge pull request #3624 from Sonicadvance1/faulty_mc_fault_face
FEXCore: Get rid of DeferredSignalFaultAddress and use the InterruptFaultPage
2024-05-13 03:22:43 -07:00
Ryan Houdek 3a7aa83ab1 Merge pull request #3626 from pmatos/TestClangIgnore
Fix exec path where file needs to be ignored
2024-05-13 02:55:33 -07:00
Paulo Matos bcc136c3b9 Fix exec path where file needs to be ignored
Ignored files were not being checked. Both clang-format.py wrapper
and code-format-helper where not aligned.
2024-05-13 11:41:44 +02:00
Ryan Houdek 3da31830d1 InstcountCI: Update 2024-05-10 15:34:13 -07:00
Ryan Houdek d19b57a52e FEXCore: Get rid of DeferredSignalFaultAddress and use the InterruptFaultPage
Arm64ec introduced the InterruptFaultPage which is lower overhead since
instead of ldr+str it just turns in to a single str. We were already
allocating the space, FEXCore and the frontend signal delegator just
needed to be updated to understand the new location.

We can additionally use this in the future if we want to make deferred
async signals INSIDE the JIT only cost a single str as well.
2024-05-10 15:31:28 -07:00
Ryan Houdek ef6d640a8c Merge pull request #3612 from Sonicadvance1/threadmanager_move
FEXLoader: Changes frontend thread management to wrap FEXCore thread objects
2024-05-09 09:27:11 -07:00
Ryan Houdek 2cae2f2462 Merge pull request #3617 from bylaws/arm64ec-dispatcher
FEXCore: ARM64EC x64 entry/exit support
2024-05-08 12:25:26 -07:00
Ryan Houdek 1fde5d7fca Merge pull request #3621 from alyssarosenzweig/ra/drop-avx
RegisterAllocationPass: drop AVX flag
2024-05-08 11:42:53 -07:00
Ryan Houdek 10de2f83ac Merge pull request #3620 from alyssarosenzweig/ir/burn-parser
IR: drop IRParser
2024-05-08 11:29:37 -07:00
Alyssa Rosenzweig 9d86e11a47 RegisterAllocationPass: drop AVX flag
RA should not depend on whether we support AVX, that's a huge layering
violation! and fortunately, it does not.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:26:31 -04:00
Alyssa Rosenzweig 4d503d3155 RegisterAllocationPass: drop prewritable check
always true.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:24:41 -04:00
Alyssa Rosenzweig 7e663b91df IR: drop IRParser
Aside from its own self-test, the parser is unused and should remain that way,
since it's a maintenance burden with no real benefit. Burn it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:16:54 -04:00
Alyssa Rosenzweig 47242dc190 Merge pull request #3616 from alyssarosenzweig/sra/simplify-1
SRA controlled burn
2024-05-08 14:10:48 -04:00
Alyssa Rosenzweig e13c8e3295 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3c3ba62c10 MemoryOps: optimize 32-bit SRA case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 34fe56dfb2 DeadStoreElimination: CSE block info
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig ecf8cde5e0 DeadStoreElimination: group common logic
slightly less obnoxious copypaste.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 55284aad7e DeadStoreElimination: don't handle partial stores
SRA replaces the whole contents of the destination.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3afc35f7b4 DeadStoreElimination: simplify
use registers internally, not synthesized offsets

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1058428a51 IR: document invariant on SRA
This lets us simplify a lot!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig a2fc51fc7b IR: specify registers, not offsets for SRA
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1848629ba5 RegisterAllocationPass: drop aliasable check
always true with the new ir invariants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3399577330 JIT: clean up fpr sra
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 18bfc8afd0 JIT: clean up gpr sra handling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 76b023ed3e JIT: drop unaligned and partial SRA handling
This is all dead, assert as much so it stays that way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig b91b0e9d65 IR: infer SRA static class
no need to stick it in the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 74489a4177 IR: remove dead SRA flags
I don't know what these were meant for, and I don't care (-:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Ryan Houdek 55d1d6bcd4 Merge pull request #3615 from bylaws/wow64-fix
Fix WOW64 frontend with recent wine versions
2024-05-07 22:35:32 -07:00
Ryan Houdek cd249e2c3a Merge pull request #3614 from Sonicadvance1/remove_temporary_allocation
FEXServer: Removes temporary variable allocation
2024-05-07 22:35:23 -07:00
Billy Laws 61cd835754 Update InstCountCI 2024-05-06 17:37:43 +00:00
Billy Laws bd24364c1b FEXCore: Switch stacks before exiting the JIT on ARM64EC
This removes the need for the frontend to have any knowledge of FEX's
SRA layout.
2024-05-06 15:41:34 +00:00
Billy Laws ab516d7b79 Dispatcher: Implement ARM64EC SRA setup entrypoints
While the ARM64EC ABI mostly matches FEX's SRA, the stack still needs to
be switched to the emulator stack and target RIP stored into the FEX
context before jumping to the dispatcher loop.
2024-05-06 15:41:34 +00:00
Billy Laws d25ed4b0bf Dispatcher: Block system call callbacks when compiling code
These callbacks are used for code invalidation and setting the right
emulated CPU features, neither of which are necessary for syscalls made
from within FEX. Avoid calling them to prevent deadlocks caused by
nested locks during compilation.
2024-05-06 15:41:28 +00:00
Billy Laws c521d2b48d WOW64: Support unwinding past FEX from within syscall handlers
This is required by recent wine changes to use longjmp for user
callbacks. Switch to saving the context at every simulate call and
setting the unwind SP/PC to that context with a small SEH trampoline
for the syscall handler.
2024-05-06 15:26:36 +00:00
Billy Laws 9ed8165405 WOW64: Dynamically allocate unixcall/syscall entrypoints
Removes the requirement that FEX needs to be loaded as part of the lower
32-bit address space.
2024-05-06 14:55:59 +00:00
Ryan Houdek 5099b2b5dc FEXServer: Removes temporary variable allocation
Was causing unnecessary memory allocation churn when a FEXInterpreter
was asking for the rootfs folder path.
2024-05-05 14:11:26 -07:00
Ryan Houdek d372552593 FEXLoader: Changes frontend thread management to wrap FEXCore thread objects
A bit of refactoring necessary before we can move the remaining Linux
specific code to the frontend.

Most of this taken from #3535 but attempting to be NFC as much as
possible.
2024-05-05 07:43:09 -07:00
Ryan Houdek 729e32ccc2 Linux: Move ThreadManager to its own header 2024-05-05 06:32:59 -07:00
Mai 170204d6f1 Merge pull request #3609 from neobrain/fix_catch2_setting
CMake: Remove obsolete Catch2 setting
2024-05-04 00:04:52 -04:00
Mai f7bfecd3f1 Merge pull request #3610 from Sonicadvance1/support_oryon_named
CPUID: Adds Qualcomm Oryon product name
2024-05-04 00:04:29 -04:00
Ryan Houdek 5f0427c253 CPUID: Adds Qualcomm Oryon product name
From https://github.com/llvm/llvm-project/pull/91022

Easy enough
2024-05-03 20:16:46 -07:00
Tony Wasserka 472860a840 CMake: Remove obsolete Catch2 setting 2024-05-03 15:25:42 +02:00
Mai eddb7d12cc Merge pull request #3608 from Sonicadvance1/readme_remove_x86
Readme: Remove misleading text about x86 hosts being supported
2024-05-03 00:04:53 -04:00
Ryan Houdek 789a9f19c0 Readme: Remove misleading text about x86 hosts being supported
This is no longer the case as x86-64 hosts is purely a development
vehicle and it is not expected for users to try this.
2024-05-02 20:38:13 -07:00
Ryan Houdek f70aafb211 Merge pull request #3607 from teohhanhui/fix/open-mode
Pass compulsory `mode` argument to `open` when `O_CREAT` is used
2024-05-02 13:08:47 -07:00
Teoh Han Hui 7519af2819 Pass compulsory mode argument to open when O_CREAT is used
From `man 2 open`:

> The mode argument must be supplied if O_CREAT or O_TMPFILE is
> specified in flags; if it is not supplied, some arbitrary bytes
> from the stack will be applied as the file mode.
2024-05-03 03:16:29 +08:00
Ryan Houdek 7c79e5dea1 Docs: Update for release FEX-2405 2024-05-02 11:35:14 -07:00
Ryan Houdek faa494c288 Merge pull request #3605 from Sonicadvance1/move_fex_versionstring_cpuid
CPUID: Removes FEX version string from CPU model name
2024-05-02 11:20:49 -07:00
Ryan Houdek b33e0e3839 Merge pull request #3584 from neobrain/feature_libfwd_guestx11
Library Forwarding: Support libGL/libvulkan without forwarding libX11
2024-05-02 11:20:34 -07:00
Tony Wasserka a15ed4c9da Library Forwarding: Drop support for libX11
The implementation of this has been brittle and is architecturally
incompatible with 32-bit guests. It's unlikely this could be fixed with
incremental improvements.

Since libGL and libvulkan can be forwarded independently of libX11 now,
these libX11 bits can be dropped without negative impact on compatibility.
2024-05-02 20:02:02 +02:00
Tony Wasserka 4b76eb0b3f Library Forwarding/vulkan: Reload pointers for extension functions more aggressively
Some applications create multiple Vulkan instances with different sets of
extensions. We might hence miss some of these pointers during the initial
function pointer query.
2024-05-02 20:02:02 +02:00
Tony Wasserka 86315027c3 Library Forwarding: Support Vulkan forwarding with guest-libX11 2024-05-02 18:06:54 +02:00
Tony Wasserka a114850c2a Library Forwarding: Support GL forwarding with guest-libX11 2024-05-02 17:59:28 +02:00
Tony Wasserka 997d1bf04b Library Forwarding: Add helper class for interception X11 communication
X11 displays and xcb connections managed by the guest libX11 can't be used by
the host, but we can create intermediary objects using the host libX11. This
allows to connect guest-managed objects to the host window system integration
APIs in OpenGL/Vulkan.
2024-05-02 17:59:28 +02:00
Tony Wasserka c294782a60 Library Forwarding: Support annotating function return types 2024-05-02 17:59:28 +02:00
Mai 9781b957d0 Merge pull request #3606 from Sonicadvance1/fix_inverted_rdtscp
CPUID: Fix inverted RDTSCP check
2024-05-01 22:21:09 -04:00
Ryan Houdek 6228226c08 CPUID: Fix inverted RDTSCP check
This was inverted and always enabling the RDTSCP cpuid bit for wine.
Thus always disabling it elsewhere.
2024-05-01 18:31:41 -07:00
Ryan Houdek 31341bb7c2 CPUID: Removes FEX version string from CPU model name
Moves it to the hypervisor leafs.

Before:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
```

After:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
```

Now the FEX string is in the hypervisor functions as a leaf, so if some
utility wants the FEX version they can query that directly

Ex:
```bash
$ ./Bin/FEXInterpreter get_cpuid_fex
Maximum 4000_0001h sub-leaf: 2
We are running under FEX on host: 2
FEX version string is: 'FEX-2404-113-g820494d'
```
2024-05-01 16:27:13 -07:00
Ryan Houdek 3fda47e870 Merge pull request #3604 from alyssarosenzweig/clang-format-path
clang-format: allow overriding clang-format
2024-04-29 19:09:23 -07:00
Ryan Houdek 9c6f749ebe Merge pull request #3601 from alyssarosenzweig/jit/sbc-adc
Fix 8/16-bit ADC/SBC
2024-04-29 18:47:34 -07:00
Alyssa Rosenzweig 711ae84175 clang-format: allow overriding clang-format
it's in a weird path for me.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 21:43:49 -04:00
Alyssa Rosenzweig bfb06b2d55 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 20:00:36 -04:00
Ryan Houdek 640e911af4 Adds unit tests for ADC/SBB garbage upper data bug
Reproduces broken rendering in Final Fantasy 7 (SteamID 39140)
2024-04-29 20:00:36 -04:00
Alyssa Rosenzweig 76b5ca4bcc OpcodeDispatcher: optimize 8/16-bit adc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 20:00:34 -04:00
Alyssa Rosenzweig 28fa88ff39 OpcodeDispatcher: fix 8/16-bit adc/sbc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 20:00:34 -04:00
Ryan Houdek 1069cabad0 Merge pull request #3598 from Sonicadvance1/half_barrier_delete_hack
Arm64: Adds another TSO hack to disable half-barrier TSO
2024-04-26 18:24:49 -07:00
Ryan Houdek fe70ec7277 Merge pull request #3599 from alyssarosenzweig/jit/fix-faddv
JIT: fix neon vec4 faddv
2024-04-24 18:20:11 -07:00
Alyssa Rosenzweig 4a4fa64254 JIT: fix neon vec4 faddv
We were previously genrating nonsense code if the destination != source:

         faddp v2.4s, v4.4s, v4.4s
         faddp s2, v4.2s

The result of the first faddp is ignored, so the second merely calculates the
sum of the first 2 sources (not all 4 as needed).

The correct fix is to feed the first add into the second, regardless of the
final destination:

         faddp v2.4s, v4.4s, v4.4s
         faddp s2, v2.2s

Hit in an ASM test with new RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-24 21:13:02 -04:00
Ryan Houdek 6463054fa3 Arm64: Adds another TSO hack to disable half-barrier TSO
A feature of FEX's JIT is that when an unaligned atomic load/store
operation occurs, the instructions will be backpatched in to a barrier
plus a non-atomic memory instruction. This is the half-barrier technique
that still ensures correct visibility of loadstores in an unaligned
context.

The problem with this approach is that the dmb instructions are HEAVY,
because they effectively stop the world until all memory operations in
flight are visible. But it is a necessary evil since unaligned atomics
aren't a thing on ARM processors. FEAT_LSE only gives you unaligned
atomics inside of a 16-byte granularity, which doesn't match x86
behaviour of cacheline size (effectively always 64B).

This adds a new TSO option to disable the half-barrier on unaligned
atomic and instead only convert it to a regular loadstore instruction,
ommiting the half-barrier. This gives more insight in to how well a
CPU's LRCPC implementation is by not stalling on DMB instructions when
possible.

Originally implemented as a test to see if this makes Sonic Adventure 2
run full speed with TSO enabled (but all available TSO options disabled)
on NVIDIA Orin. Unfortunately this basically makes the code no longer
stall on dmb instructions and instead just showing how bad the LRCPC
implementation is, since the stalls show up on `ldapur` instructions
instead.

Tested Sonic Adventure 2 on X13s and it ran at 60FPS there without the
hack anyway.
2024-04-24 13:09:00 -07:00
Ryan Houdek 81a4206805 Merge pull request #3581 from neobrain/fix_libfwd_x11_libname
Library Forwarding: Fix issues with libGL's fake X11 dependency
2024-04-23 14:07:09 -07:00
Ryan Houdek a0bf6a4255 Merge pull request #3595 from alyssarosenzweig/ir/before
Factor out SetWriteCursorBefore
2024-04-23 13:34:05 -07:00
Mai 7b88b0f271 Merge pull request #3596 from Sonicadvance1/fix_fedora
Allocator: Fixes compiling on Fedora 40
2024-04-23 16:03:52 -04:00
Ryan Houdek 308488c419 Allocator: Fixes compiling on Fedora 40
This header was missing.

Either libstdc++14 or clang-18 changed includes and we were only getting
this indirectly before.
2024-04-23 12:12:33 -07:00
Ryan Houdek 81c219c212 Merge pull request #3589 from bylaws/wow-inval
WOW64 backend code invalidation fixes
2024-04-23 10:50:41 -07:00
Alyssa Rosenzweig 68e543c81d Merge pull request #3594 from Sonicadvance1/enhanced_repmovs
CPUID: Enable enhanced rep movs in more situations
2024-04-23 13:10:39 -04:00
Alyssa Rosenzweig 2372c9458b ConstProp: use SetWriteCursorBefore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-23 13:09:45 -04:00
Alyssa Rosenzweig 1a11343f34 IREmitter: add SetWriteCursorBefore helper
This is subtle, add an ergonomic helper for it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-23 13:09:45 -04:00
Tony Wasserka cf75cdd16d Library Forwarding: Fix issues with libGL's fake X11 dependency
The library's soname is changed to libX11.so.6 and the CMake target is
renamed to libPlaceholderX11. This fixes two issues:
* Steam and mangohud can't find libX11 during startup if the library doesn't
  include a version suffix.
* Calling the CMake target libX11 overrode the true host X11 library used by
  unrelated targets (such as FEXConfig), which could cause link errors
2024-04-23 16:22:58 +02:00
Ryan Houdek 84d5b3ee59 CPUID: Enable enhanced rep movs in more situations
Instead of only enabling enhanced rep movs if software TSO is disabled,
Enable it if software tso is disabled OR memcpysettso is disabled. This
is because now we hit the fast path when memcpysettso is disabled alone
but global TSO is disabled.

Retested Hades and performance was fine in this configuration.
2024-04-21 18:50:17 -07:00
Ryan Houdek 376936c808 Merge pull request #3591 from alyssarosenzweig/ra/fix
JIT: fix ShiftFlags shuffles
2024-04-19 13:50:45 -07:00
Ryan Houdek 20bd86473b Merge pull request #3588 from Sonicadvance1/implement_smsw
OpcodeDispatcher: Implement support for SMSW
2024-04-19 07:18:45 -07:00
Alyssa Rosenzweig 932b8f38f4 JIT: fix ShiftFlags shuffles
messed up my RA.

fixes ShiftPF.asm with jit_1 with a pathological register allocation

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-18 14:11:03 -04:00
Billy Laws a7f4e99278 InvalidationTracker: Always invalidate all regions of a section on unmap
Unmapping a section will unmap the whole size initially allocated,
irrespective of how their protections are changed afterwards. Make sure
to follow this logic for invalidation too.
2024-04-18 15:16:28 +00:00
Billy Laws 7391456e48 Windows: Don't redefine existing MinGW ntdll exports 2024-04-18 15:15:11 +00:00
Billy Laws a6d061b711 InvalidationTracker: Invalidate code across all threads
When thread management was moved to the frontend, invalidation moved
from being a global operation to per-thread but the WOW64 backend wasn't
updated to account for this. Now for any invalidation event loop over
all threads tracked by the frontend and invalidate the appropriate
range.
2024-04-18 15:14:31 +00:00
Billy Laws d92580bccf WOW64: Keep track of all created threads on the frontend
This is necessary so that code can be invalidated across all threads
rather than just the initiator on any event that triggers invalidation.
2024-04-18 15:00:24 +00:00
Ryan Houdek c8704a7f71 OpcodeDispatcher: Implement support for SMSW
Found out that Far Cry uses this instruction and it is viable to use in
CPL-3. This only returns constant data but its behaviour is a little
quirky.

This instruction has a weird behaviour that the 32-bit operation does an
insert in to the 64-bit destination, which might be an Intel versus AMD
behaviour. I don't have an Intel machine available to test if that
theory is true although. This assumption would match similar behaviour
where segment registers are inserted instead of zext.

Gets the game farther but then it crashes in a `___ascii_strnicmp`
function where the arguments end up being `___ascii_strnicmp(nullptr, "Color", 5);`.
2024-04-18 07:41:39 -07:00
Ryan Houdek 5cb11aed3d Merge pull request #3587 from alyssarosenzweig/revert/waw
RCLSE: disable store-after-store optimization
2024-04-17 20:20:02 -07:00
Alyssa Rosenzweig 625d8ac177 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-17 14:54:19 -04:00
Alyssa Rosenzweig 352dcdb478 RCLSE: disable store-after-store optimization
Functional revert of 92f31648b ("RCLSE: optimize out pointless stores"), which
reportedly regressed some titles due to RA doom. We'll revisit later, leaving in
the code for when RA is ready to light this up.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-17 14:53:11 -04:00
Tony Wasserka 02ebb6e320 Merge pull request #3580 from neobrain/fix_libfwd_steamwebhelper_x11
AppConfig: Disable libGL forwarding for steamwebhelper
2024-04-17 07:41:25 +00:00
Tony Wasserka 53bed6de6a Merge pull request #3572 from pmatos/ClangFormatUtils
Clang Format file and script
2024-04-16 12:21:14 +00:00
Paulo Matos 00b81c8cbc Clang Format file and reformat target
Adds paths to ignore to clang-format, .clang-format file.
Also a wrapper script to clang-format to read .clang-format-ignore.

To format the whole tree at the root of the repository run:
`find . -iname '*.h' -o -iname '*.cpp' -exec python3 Scripts/clang-format.py -i \{\} \;`

Add reformat target to reformat the whole tree
2024-04-16 13:48:44 +02:00
Mai c126b209f1 Merge pull request #3582 from neobrain/fix_ci_x86flags
CI: Drop use of obsolete ENABLE_X86_HOST_DEBUG setting
2024-04-15 22:42:03 -04:00
Tony Wasserka 97386a7260 AppConfig: Disable libGL forwarding for steamwebhelper
This app bypasses the glX functions exported by libGL and instead sends GLX
requests directly via xcb. FEX won't support forwarding this usage pattern in
the foreseeable future.
2024-04-15 18:33:14 +02:00
Tony Wasserka a054b998c5 CI: Drop use of obsolete ENABLE_X86_HOST_DEBUG setting 2024-04-15 18:31:50 +02:00
Ryan Houdek f9097c0a92 Merge pull request #3561 from pmatos/cformat-PRCI
CI workflow to check clang-format usage on pull requests
2024-04-15 06:32:42 -07:00
Ryan Houdek 55c054bd8e Merge pull request #3578 from pmatos/Reformat2Ignore
Add second reformat to git blame ignore file
2024-04-15 06:31:13 -07:00
Paulo Matos 8799a48a0f Add second reformat to git blame ignore file 2024-04-15 15:23:49 +02:00
Paulo Matos e1efde9605 CI workflow to check clang-format usage on pull requests
Adapted from LLVM version of pr-code-format.yml.
Copies a few scripts from LLVM to External/.
Runs self-hosted on X64.

Assumes clang-format 16.0.6 for formatting.
2024-04-15 14:05:48 +02:00
Ryan Houdek 07e58f8db7 Merge pull request #3577 from pmatos/Reformat2
Reformat until fixed-point
2024-04-15 00:47:11 -07:00
Paulo Matos 905aa935f5 Reformat until fixed-point
Followup to 2b4ec88dae.
Some files needed a couple of calls to clang-format 16.0.6 to
reach a fixed point.
2024-04-15 09:40:00 +02:00
Ryan Houdek 7e97db8d98 Merge pull request #3575 from pmatos/GitBlameIgnoreRevs
Create .git-blame-ignore-revs with whole-tree reformat sha
2024-04-12 23:51:51 -07:00
Paulo Matos 340630a937 Create .git-blame-ignore-revs with whole-tree reformat sha
To enable it:
`git config blame.ignoreRevsFile .git-blame-ignore-revs`
2024-04-13 08:44:48 +02:00
Ryan Houdek 7614ac9f14 Merge pull request #3573 from pmatos/RemoveClangTidy
Remove trace of clang-tidy experiment from CMakeLists.txt
2024-04-12 17:13:53 -07:00
Ryan Houdek 1d806ac218 Merge pull request #3571 from pmatos/ReformatWholesale
Whole-tree reformat
2024-04-12 17:11:59 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Tony Wasserka 028c220041 Merge pull request #3574 from pmatos/ConstPlacementLeft
Move const to the left in preparation for reformatting
2024-04-12 14:16:46 +00:00
Paulo Matos 6524716404 Move const to the left in preparation for reformatting
clang-format-16 had some issues with const placement, so we are manually changing these.
2024-04-12 16:06:57 +02:00
Paulo Matos 20559853ee Remove trace of clang-tidy experiment from CMakeLists.txt 2024-04-12 12:31:04 +02:00
Ryan Houdek a9b7ad841c Merge pull request #3570 from bylaws/ec_pt8
Enable jemalloc for ARM64EC
2024-04-11 12:57:36 -07:00
Ryan Houdek ae5e388e1a Merge pull request #3569 from bylaws/ec_pt7
FHU: Switch over win32 file operations to std::filesystem
2024-04-11 00:33:04 -07:00
Ryan Houdek 271700e9f6 Merge pull request #3568 from lioncash/const
X87: Simplify constant loading for FLD family
2024-04-11 00:32:26 -07:00
Ryan Houdek 1ba678f631 Merge pull request #3562 from Sonicadvance1/fix_rsp_store_tso
OpcodeDispatcher: Fixes disabling TSO access on RSP SIB stores
2024-04-11 00:32:14 -07:00
Billy Laws ef48700f3e FHU: Switch over win32 file operations to std::filesystem
These were broken to varying degrees across the board on win32,
just switch to their std::filesystem as windows doesn't have the
same allocator issues that require their reimplementation.
2024-04-10 22:28:12 +00:00
Ryan Houdek a0f2cae1cb Merge pull request #3567 from lioncash/veczero
IR: Remove VectorZero
2024-04-09 20:04:26 -07:00
Ryan Houdek 66cbb66732 Merge pull request #3566 from lioncash/address
OpcodeDispatcher: Add helper for making segment offset addresses
2024-04-09 17:31:43 -07:00
Billy Laws 7ff4cd9108 Update jemalloc submodule 2024-04-09 23:43:49 +00:00
Billy Laws 5ed593b59f CMake: Force enable jemalloc when targetting ARM64EC 2024-04-09 23:42:23 +00:00
Billy Laws f1f0c47f16 AllocatorHooks: Allow using jemalloc on win32 2024-04-09 23:42:23 +00:00
Lioncache 4cb2432b5c OpcodeDispatcher: Make use of new x87 constants
Now we can load these directly instead of needing to manually materialize them.
2024-04-09 10:17:15 -04:00
Lioncache 27ba66a181 IRDumper: Extend printer for NamedVectorConstant
Makes it aware of the x87 constants.
2024-04-09 10:13:35 -04:00
Lioncache 65b5281d7c IR: Add constants for FLD variants 2024-04-09 10:13:33 -04:00
Ryan Houdek 1a8b61b9fc Merge pull request #3560 from bylaws/ec-pt6
FEXCore: Support x64 -> arm64ec calls
2024-04-09 07:08:38 -07:00
Ryan Houdek f0dad86332 Merge pull request #3559 from bylaws/ec-pt5
LookupCache: Track ARM64EC page state in the code cache
2024-04-09 07:08:29 -07:00
Mai eedb120fd0 Merge pull request #3563 from Sonicadvance1/fill_spill_pairs
JIT: Adds support for spilling/Filling GPRPair
2024-04-08 23:05:11 -04:00
Lioncache 98841fe07a IR: Remove VectorZero
We have LoadNamedVectorConstant that now performs this behavior while
also being more flexible.
2024-04-08 22:42:58 -04:00
Lioncache b0aeb501f4 OpcodeDispatcher: Add helper for making segment offset addresses
There's quite a few places where the segment offset appending is open-coded
throughout the opcode dispatcher, but we can pull these out into a few
helpers to make the sites a little more compact and declarative.
2024-04-08 17:50:58 -04:00
Ryan Houdek 1616b4e77c Merge pull request #3564 from lioncash/header
DebugData: Remove header
2024-04-08 14:33:42 -07:00
Lioncache b26bf2eaf6 DebugData: Remove header
This isn't included or used anywhere, so it can be removed.
2024-04-08 16:05:40 -04:00
Ryan Houdek e91e1d5533 Merge pull request #3550 from alyssarosenzweig/ra/validate
Validate that we have no crossblock liveness
2024-04-08 12:07:23 -07:00
Ryan Houdek 574fdcef32 JIT: Adds support for spilling/Filling GPRPair
Tony noticed this last week. I encountered it this week.
Add support for spilling and filling GPR pairs.
2024-04-08 11:55:29 -07:00
Ryan Houdek 0e93fd0f3e OpcodeDispatcher: Fixes disabling TSO access on RSP SIB stores
GPR Direct/Indirect already had this and SIB version also already
supported on the load side. Fixes this missed behaviour.
2024-04-08 11:32:01 -07:00
Alyssa Rosenzweig 063954c9b3 ValueDominanceValidation: rm deadcode
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:49:35 -04:00
Alyssa Rosenzweig 7b3e031678 ValueDominanceValidation: forbid crossblock liveness
Now that we have successfully eliminated crossblock liveness from the IR we
generate, validate as much to ensure it doesn't come back. We will take
advantage of this new invariant in RA in the future.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:49:35 -04:00
Alyssa Rosenzweig a775e474d5 ValueDominanceValidation: do not validate inline constants
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:49:35 -04:00
Alyssa Rosenzweig 0e99019586 ValueDominanceValidation: actually validate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:49:35 -04:00
Alyssa Rosenzweig d0ff3e64d9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:49:31 -04:00
Alyssa Rosenzweig d9493e5d9b OpcodeDispatcher: fix xblock liveness in xsave/xrstr
didn't fix this hard enough before. caught by validation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:48:29 -04:00
Alyssa Rosenzweig eb83c9e7f2 Core: use safe CondJump for self-modifying code
this ensures we put the StoreNZCV in the right block, which will fix validation
fails later in the series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-08 13:48:29 -04:00
Ryan Houdek bb24e1419c Merge pull request #3558 from bylaws/ec-pt3
AllocatorHooks: Mark JIT code memory as EC code on ARM64EC
2024-04-08 09:32:51 -07:00
Ryan Houdek cbfa426b59 Merge pull request #3548 from alyssarosenzweig/ra/shifts-rework
Eliminate xblock liveness for shifts
2024-04-08 09:31:49 -07:00
Billy Laws 526e3e654f LookupCache: Track ARM64EC page state in the code cache
Rather than checking the actual EC bitmap in the dispatcher (~6 instrs), this
indirection through the code cache allows just 1 instr for the hot path
of calling repeated EC code/x64 code.
2024-04-08 16:08:17 +01:00
Billy Laws e03434ebde Update InstCountCI 2024-04-08 16:04:59 +01:00
Billy Laws 243bb45a68 FEXCore: Support x64 -> arm64ec calls
The frontend will provide the return logic via ExitFunctionEC, which
will be jumped to whenever there is an indirect branch/return to an addr
such that RtlIsEcCode(addr) returns true.
2024-04-06 13:20:48 +00:00
Billy Laws bd5b817c3a AllocatorHooks: Mark JIT code memory as EC code on ARM64EC
Executable mapped memory is treated as x86 code by default when
running under EC, VirtualAlloc2 needs to be used together with a
special flag to map JIT arm64 code.
2024-04-06 12:40:52 +00:00
Alyssa Rosenzweig b54d4931fb InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 95589f6172 OpcodeDispatcher: rm deferred variable shift flag calcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 098859caf7 OpcodeDispatcher: use _ShiftFlags for ASHR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig c632543451 OpcodeDispatcher: use _ShiftFlags for SHRD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 801cf72f95 OpcodeDispatcher: use _ShiftFlags for SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 650cd2c46e OpcodeDispatcher: use _ShiftFlags for SHR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 5b48ce2228 OpcodeDispatcher: use _ShiftFlags for SHL
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 2173c26fd8 OpcodeDispatcher: add HandleShift helper
all the variable shift impls need to do this dance, make it common.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 982391ba9d IR: add ShiftFlags op
Generates flags for a variable shift as a dedicated IR op. This lets us optimize
around it (without generating control flow, relying on deferred flag infra,
etc). And it neatly solves our RA problem for shifts.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig a99c48b7a3 RedundantFlagCalculationElimination: do not eliminate if there are uses
we'll hit this with _ShiftFlags.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 90d5bd3aec unittests: add test for shift PF bug
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 59c3f96a23 unittests: add test for deferred flags + shift with cl=0
this failed on an earlier version of the series that otherwise passed ci.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 3661da1bc6 OpcodeDispatcher: calculate deferred flags before RMW on NZCV
otherwise we might have the wrong input NZCV.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 20:34:05 -04:00
Alyssa Rosenzweig 859df5e0b2 OpcodeDispatcher: optimize shl flag
This is something the new shift flag code will do. Backporting the opt since
that's stalled and this reduces the diff.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-05 19:38:44 -04:00
Ryan Houdek e8127b92e8 Docs: Update for release FEX-2404 2024-04-05 15:20:24 -07:00
Ryan Houdek 7786c23405 Merge pull request #3556 from Sonicadvance1/move_app_config
FEXCore: Fixes priority of FEX_APP_CONFIG
2024-04-05 15:18:24 -07:00
Ryan Houdek 904646e93b FEXCore: Fixes priority of FEX_APP_CONFIG
This environment variable had an incorrect priority on the configuration
system. The expectation was higher priority than most other layers.

Now the only layer that has higher priority is the environment
variables.
2024-04-05 13:10:43 -07:00
Alyssa Rosenzweig c43af8e975 Merge pull request #3553 from alyssarosenzweig/ra/shifts-rework-easy
OpcodeDispatcher: clean up shifts
2024-04-05 11:36:33 -04:00
Alyssa Rosenzweig a787daae41 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:27 -04:00
Alyssa Rosenzweig a05cc06ab4 OpcodeDispatcher: unify imm/1-bit ASHR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 031e756a78 OpcodeDispatcher: unify imm/1-bit SHR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 2a9f1ce8cb OpcodeDispatcher: unify imm/1-bit SHL
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 8c53a9f051 OpcodeDispatcher: use LoadConstantShift for rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig cf26ec7898 OpcodeDispatcher: use LoadConstantShift for SHRD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 582c3dae6e OpcodeDispatcher: use LoadConstantShift for SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 2abac03ab0 OpcodeDispatcher: add LoadConstantShift helper
shows up a bunch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 8cc684fa12 OpcodeDispatcher: drop misinformed comment
tbnz only tests a single bit, not a mask.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig d92de1d947 OpcodeDispatcher: drop result masking for shifts
flag calcs are fine with upper garbage.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-04 07:42:15 -04:00
Alyssa Rosenzweig 202a60b77a Merge pull request #3549 from alyssarosenzweig/constprop/dce
ConstProp: drop dead code
2024-04-03 11:22:30 -04:00
Alyssa Rosenzweig aa8d04c341 Merge pull request #3551 from alyssarosenzweig/opt/negate-adds
Negate more to inline constants
2024-04-03 11:22:01 -04:00
Alyssa Rosenzweig d6425d05f3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:57:06 -04:00
Alyssa Rosenzweig e07c81a5e7 ConstProp: also negate sub -> add
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:56:59 -04:00
Alyssa Rosenzweig fa76961873 ConstProp: negate adds -> subs
the arm ops are equiv, even though the x86 isn't (due to inverted carry).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:53:21 -04:00
Alyssa Rosenzweig b92c206db9 ConstProp: rm your deadcode
not sure who this is supposed to be helping.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:21:48 -04:00
Alyssa Rosenzweig efff942724 ConstProp: drop my deadcode
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:21:48 -04:00
Alyssa Rosenzweig 8d32113521 ConstProp: rm relic of x86 jit
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-02 13:21:48 -04:00
Alyssa Rosenzweig 37f2b417e4 Merge pull request #3546 from alyssarosenzweig/flag/cleanup
Minor cleanups around flags
2024-04-02 11:29:19 -04:00
Alyssa Rosenzweig bd0b5eceb8 Merge pull request #3545 from alyssarosenzweig/opt/pf-scalar
Use scalar integer code to calculate PF
2024-04-02 11:28:53 -04:00
Alyssa Rosenzweig b632f7215c Merge pull request #3544 from alyssarosenzweig/ra/zero-multiple
OpcodeDispatcher: drop ZeroMultipleFlags
2024-04-02 11:27:45 -04:00
Ryan Houdek e8abc88702 Merge pull request #3542 from alyssarosenzweig/ra/rep
Eliminate xblock liveness with rep cmp/lod/scas
2024-04-02 04:24:24 -07:00
Ryan Houdek 29c6281e11 Merge pull request #3539 from alyssarosenzweig/ra/rol-ror2
rewrite ROL/ROR
2024-04-02 00:17:08 -07:00
Ryan Houdek 4214d9bda0 Merge pull request #3538 from pmatos/OffsetofOoB
Fix reference to out of bounds address in offsetof
2024-04-01 19:41:57 -07:00
Alyssa Rosenzweig 0a7b3efb41 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 16:42:50 -04:00
Alyssa Rosenzweig 067a5444dc OpcodeDispatcher: use HandleNZ00Write
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 16:42:38 -04:00
Alyssa Rosenzweig c7f159972d OpcodeDispatcher: rm pointless NZCV loads
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 16:42:38 -04:00
Alyssa Rosenzweig a70d0a5dd4 OpcodeDispatcher: rm unnecessary NZCV dirtying
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 16:42:38 -04:00
Ryan Houdek cd9ffd2045 Merge pull request #3536 from alyssarosenzweig/ra/rcl-rcr
OpcodeDispatcher: eliminate xblock liveness for rcl/rcr
2024-04-01 11:44:37 -07:00
Alyssa Rosenzweig 5c590b9a50 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 14:13:09 -04:00
Alyssa Rosenzweig eb4bb5875e OpcodeDispatcher: absorb invert into PF calculation
with xorn

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 14:12:33 -04:00
Alyssa Rosenzweig 3b052e826f OpcodeDispatcher: calculate PF with integer ops
based on clang's __builtin_parity

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 14:12:32 -04:00
Alyssa Rosenzweig 65ec191dc1 IR: add XornShift
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 14:12:32 -04:00
Alyssa Rosenzweig b1ddd8cd3b Merge pull request #3541 from alyssarosenzweig/opt/clc
optimize clc
2024-04-01 13:51:10 -04:00
Alyssa Rosenzweig f2d001e721 Merge pull request #3543 from alyssarosenzweig/ra/dead-code
RA: drop dead block interference code
2024-04-01 13:51:00 -04:00
Alyssa Rosenzweig 7852909cc4 OpcodeDispatcher: simplify IsNZCV
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 13:50:00 -04:00
Alyssa Rosenzweig fad243d3f6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 13:48:27 -04:00
Alyssa Rosenzweig f8b68d8b5a OpcodeDispatcher: drop ZeroMultipleFlags
lot of complexity for only a single interesting case. we can massively simplify.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-01 13:48:11 -04:00
Ryan Houdek e2a095372e Merge pull request #3534 from Sonicadvance1/move_ir_defines
FEXCore: Move nearly all IR definitions to internal
2024-04-01 10:00:20 -07:00
Ryan Houdek 5c29c9d464 Merge pull request #3527 from Sonicadvance1/move_type_defines
Moves FHU TypeDefines to FEXCore includes
2024-04-01 08:57:22 -07:00
Ryan Houdek 3bed305660 Merge pull request #3526 from Sonicadvance1/move_codeloader
FEXCore: Moves CodeLoader to frontend
2024-04-01 07:52:02 -07:00
Ryan Houdek f6639c3594 Merge pull request #3525 from Sonicadvance1/move_cpubackend
FEXCore: Moves CPUBackend definition internal
2024-04-01 06:47:34 -07:00
Paulo Matos 96087a69fa Fix reference to OoB address in offsetof and remove rflags printout
Adjust static array size to match new size.
Remove rflags from printing code and adjust offsets - fixes
printing off-by-one error.
2024-04-01 13:13:17 +02:00
Alyssa Rosenzweig ca1ec232c9 RA: drop dead block interference code
Unused, and new RA won't use it either. Torch it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:51:11 -04:00
Alyssa Rosenzweig ad0dd34412 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:31:38 -04:00
Alyssa Rosenzweig 7b1bb159fa OpcodeDispatcher: use ForeachDirection for scas
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:31:38 -04:00
Alyssa Rosenzweig 5c7f2934de OpcodeDispatcher: use ForeachDirection for lods
eliminates xblock live

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:29:29 -04:00
Alyssa Rosenzweig 5d79d4eb50 OpcodeDispatcher: use ForeachDirection for CMPS
eliminates xblock liveness

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:29:16 -04:00
Alyssa Rosenzweig 3f66173bc7 OpcodeDispatcher: add ForeachDirection helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:28:56 -04:00
Alyssa Rosenzweig b64a594b16 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:01:44 -04:00
Alyssa Rosenzweig 4452f0acba ConstProp: optimize rmif with 0 for clc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 20:01:44 -04:00
Alyssa Rosenzweig 784cdd7b6b InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 19:54:14 -04:00
Alyssa Rosenzweig 1a1545da0f OpcodeDispatcher: rework rep cmp
1. pull flag calculation out of the loop body for perf
2. fully rotate the inner loop to save an instruction per iteration
3. hoist the rcx=0 jump to avoid computing df when rcx=0

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 19:54:00 -04:00
Alyssa Rosenzweig a70ea30c02 IR: add CondSubNZCV (ccmp) instruction
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 17:50:57 -04:00
Alyssa Rosenzweig 2aa1fd7fa3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 14:47:19 -04:00
Alyssa Rosenzweig 15b86e4c5a OpcodeDispatcher: rewrite ROL/ROR
single unified implementation for ROL & ROR (instead of 4 cases). no more
deferred flags because it's easy to shoot ourselves in the foot with deferred
flags w.r.t the new RA design, and rotates are rare enough with very efficient
flag calculations such that the extra JIT overhead should be minimal to DCE the
resulting calculations later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-31 14:47:19 -04:00
Alyssa Rosenzweig 67baff8a57 Merge pull request #3537 from Sonicadvance1/remove_vla_ra
RA: Removes VLA usage
2024-03-31 14:44:37 -04:00
Ryan Houdek fedc24be1e RA: Removes VLA usage
Just like #3508, clang-18 complains about VLA usage.

This vector is relatively small, only around 18 elements but is
semi-dynamic depending on arch and if FEXCore is targeting Linux or
Win32.
2024-03-30 16:50:04 -07:00
Alyssa Rosenzweig 2a625a467b Merge pull request #3530 from alyssarosenzweig/opt/cmpxchg-flags2
Optimize cmpxchg with flagm
2024-03-30 15:22:40 -04:00
Alyssa Rosenzweig 6f5e4fd34b OpcodeDispatcher: add non-flag calc version of ShiftVariable
more correct for rcl, etc

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:30:36 -04:00
Alyssa Rosenzweig bdda99e44f ConstProp: constant fold Neg
will come up with rotate in the next patch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:15:11 -04:00
Alyssa Rosenzweig 9bca052146 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:12:59 -04:00
Alyssa Rosenzweig 706065b0e2 OpcodeDispatcher: accelerate cmpxchg with flagm
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:12:59 -04:00
Alyssa Rosenzweig 9fd32f07cb JIT: preserve nzcv for the slow atomic path
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:12:59 -04:00
Alyssa Rosenzweig deba6a1b76 JIT: add comment about unaligned backpatching
save future me some grief.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:12:59 -04:00
Alyssa Rosenzweig 6866c3d0ac InstCountCI: add dead cmpxchg case
not super optimizable but worth tracking.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 14:12:59 -04:00
Alyssa Rosenzweig d25ace43aa Merge pull request #3528 from alyssarosenzweig/ra/xsave-xrstor
Eliminate crossblock liveness in xsave/xrstor
2024-03-30 14:11:25 -04:00
Alyssa Rosenzweig d3b2ddf641 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:41 -04:00
Alyssa Rosenzweig 9010b3c117 OpcodeDispatcher: use neg trick for rcl smaller
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:41 -04:00
Alyssa Rosenzweig b0e001b660 OpcodeDispatcher: elim xblock live for smaller rcl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:41 -04:00
Alyssa Rosenzweig eadacbd67b OpcodeDispatcher: elim xblock live with smaller rcr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:03 -04:00
Alyssa Rosenzweig 7de29749be OpcodeDispatcher: eliminate xblock live for rcl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:03 -04:00
Alyssa Rosenzweig 1f3843ccad OpcodeDispatcher: eliminate xblock live for rcr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:03 -04:00
Alyssa Rosenzweig bc76df9901 OpcodeDispatcher: add non-flag calc version of ShiftVariable
more correct for rcl, etc

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:03 -04:00
Alyssa Rosenzweig 6e92cc454d ConstProp: constant fold Neg
will come up with rotate in the next patch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-30 13:42:03 -04:00
Ryan Houdek ed3af580c5 FEXCore: Move nearly all IR definitions to internal
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.

Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.

No functional change with this, just moving headers around.
2024-03-29 17:20:18 -07:00
Mai 2ad170b7d7 Merge pull request #3533 from Sonicadvance1/remove_debugstore
FEXCore: Remove DebugStore map
2024-03-29 20:19:47 -04:00
Ryan Houdek 8564290f76 FEXCore: Remove DebugStore map
This hasn't been used and is blocking refactoring more code.
2024-03-29 14:58:44 -07:00
Alyssa Rosenzweig 1d96631af7 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-29 09:58:03 -04:00
Alyssa Rosenzweig c513b9685d OpcodeDispatcher: eliminate crossblock liveness in xsave/xrstor
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-29 09:57:16 -04:00
Ryan Houdek d11a36eaea Moves FHU TypeDefines to FEXCore includes
FEXCore includes was including an FHU header which would result in
compilation failure for external projects trying to link to libFEXCore.

Moves it over to fix this, it was the only FHU usage in FEXCore/include
NFC
2024-03-29 02:54:54 -07:00
Ryan Houdek f46e88ebdb FEXCore: Moves CPUBackend definition internal
This is no longer necessary to be part of the public API. Moves the
header internally.

Needed to pass through `IsAddressInCodeBuffer` from CPUBackend through
the Context object, but otherwise no functional change.
2024-03-29 02:27:29 -07:00
Ryan Houdek 20eb338644 FEXCore: Moves CodeLoader to frontend
FEXCore no longer has a need for this since a bunch of related code was
already moved to the frontend. Move the CodeLoader now.
2024-03-29 02:24:53 -07:00
Ryan Houdek aa26b6288e Merge pull request #3522 from alyssarosenzweig/ra/cmpxchg8
OpcodeDispatcher: eliminate branch in cmpxchg pair
2024-03-27 21:56:19 -07:00
Mai 3d31291c3d Merge pull request #3510 from Sonicadvance1/fix_pthread_memleak
Linux/Threads: Fixes a stack memory leak for pthreads
2024-03-27 21:38:44 -04:00
Ryan Houdek 624bc3fce5 Merge pull request #3520 from Sonicadvance1/sleep_process
FEXLoader: Add a way to sleep a process on startup
2024-03-27 18:35:06 -07:00
Alyssa Rosenzweig d1722ab119 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-27 12:40:27 -04:00
Alyssa Rosenzweig 61758ea47d OpcodeDispatcher: eliminate branch in cmpxchg pair
In the old case:

* if we take the branch, 1 instruction
* if we don't take the branch, 3 instruction
* branch predictor fun
* 3 instructions of icache pressure

In the new case:

* unconditionally 2 instructions
* no branch predictor dependence
* 2 instructions of icache pressure

This should not be non-neglibly worse, and it simplifies things for RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-27 12:40:06 -04:00
Ryan Houdek 7b74ca1931 Merge pull request #3514 from alyssarosenzweig/opt/demon
rewrite Demon Addition Adjust (DAA) and other demonic opcodes
2024-03-26 23:24:00 -07:00
Ryan Houdek 24fd28ed9e Merge pull request #3511 from Sonicadvance1/more_tso_levers
FEXCore: Adds more TSO control levers
2024-03-26 23:23:41 -07:00
Ryan Houdek 970d5d5b13 Merge pull request #3509 from Sonicadvance1/allow_telemetry_redirect
Telemetry: Allow redirecting directory that data is written to
2024-03-26 23:23:05 -07:00
Ryan Houdek 79454ed8a6 Merge pull request #3507 from Sonicadvance1/fd_tracking_check
FEXLoader: Add some debug-only tracking for FEX owned FDs
2024-03-26 23:22:29 -07:00
Ryan Houdek 7f90ca53f7 Merge pull request #3505 from Sonicadvance1/telemetry_noncanonical
Telemetry: Adds tracker for non-canonical memory access crash
2024-03-26 23:21:32 -07:00
Mai 542f454630 Merge pull request #3517 from Sonicadvance1/remove_mman
FEXCore: Removes vestigial mman SMC checking
2024-03-26 23:28:05 -04:00
Ryan Houdek 4ea6305940 Merge pull request #3521 from neobrain/fix_libfwd_vkcreateinstance
Library Forwarding/vulkan: Fix query of vkCreateInstance function pointer
2024-03-26 08:59:23 -07:00
Tony Wasserka 4d8ffa2abb Library Forwarding/vulkan: Fix query of vkCreateInstance function pointer
The Vulkan specification states that querying "global commands" like
vkCreateInstance with a non-NULL instance is undefined behavior. Indeed, some
implementations will return null pointers in such cases.

Instead, we can drop the query from DoSetupWithInstance altogether, since
the library initializer will load the function pointer using dlsym instead.

Fixes #3519.
2024-03-26 16:25:51 +01:00
Ryan Houdek ade0c46845 FEXLoader: Add a way to sleep a process on startup
I find myself reimplementing this nearly monthly. Actually codify it so
I can stop reimplementing it.
2024-03-26 07:48:09 -07:00
Ryan Houdek 1450c92b60 Merge pull request #3518 from pmatos/20MFix
Put <20M in double quotes to avoid truncate error
2024-03-26 05:10:45 -07:00
Paulo Matos 53f02ee869 Put <20M in double quotes to avoid truncate error
Avoids bash assuming < is redirection.
See error in https://github.com/FEX-Emu/FEX/actions/runs/8434045431/job/23096421079#step:18:16
2024-03-26 11:47:57 +00:00
Ryan Houdek 6f29e75f67 FEXCore: Removes vestigial mman SMC checking
This wasn't actually wired up to anything ever since some refactoring
occured two years ago.
2024-03-26 02:56:26 -07:00
Ryan Houdek c1c797bcba FEXConfig: Add new TSO levers
Nice and convenient when testing applications.
2024-03-26 02:50:54 -07:00
Alyssa Rosenzweig bbc232741b InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:20 -04:00
Alyssa Rosenzweig dfe0bdd7f2 OpcodeDispatcher: rewrite DAS
exhaustively checked against the Intel pseudocode since this is tricky:

  def intel(AL, CF, AF):
      old_AL = AL
      old_CF = CF
      CF = False

      if (AL & 0x0F) > 9 or AF:
          Borrow = AL < 6
          AL = (AL - 6) & 0xff
          CF = old_CF or Borrow
          AF = True
      else:
          AF = False

      if (old_AL > 0x99) or old_CF:
          AL = (AL - 0x60) & 0xff
          CF = True

      return (AL & 0xff, CF, AF)

  def fex(AL, CF, AF):
      AF = AF | ((AL & 0xf) > 9)
      CF = CF | (AL > 0x99)
      NewCF = CF | (AF if (AL < 6) else CF)
      AL = (AL - 6) if AF else AL
      AL = (AL - 0x60) if CF else AL
      return (AL & 0xff, NewCF, AF)

  for AL in range(256):
      for CF in [False, True]:
          for AF in [False, True]:
              ref = intel(AL, CF, AF)
              test = fex(AL, CF, AF)
              print(AL, "CF" if CF else "", "AF" if AF else "", ref, test)
              assert(ref == test)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Alyssa Rosenzweig e26481e3cc OpcodeDispatcher: simplify AAM
in the area.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Alyssa Rosenzweig 86b5a2f352 OpcodeDispatcher: simplify AAD
noticed in the area.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Alyssa Rosenzweig 2bf880c43a OpcodeDispatcher: rewrite AAS
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Alyssa Rosenzweig 583d4f8f94 OpcodeDispatcher: factor out CalculateAFForDecimal
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Alyssa Rosenzweig 3ca2c4377f OpcodeDispatcher: rewrite AAA
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 19:43:10 -04:00
Ryan Houdek 32ec4a3c81 Merge pull request #3508 from Sonicadvance1/stop_using_vla
ELFParser: Stop using a VLA
2024-03-25 12:18:37 -07:00
Ryan Houdek 76983476b9 Merge pull request #3504 from Sonicadvance1/fix_loop_a16
OpcodeDispatcher: Fixes 32-bit mode LOOP RCX register usage
2024-03-25 12:18:14 -07:00
Alyssa Rosenzweig 150af80f3f Merge pull request #3512 from Sonicadvance1/panic_spilling_block
InstcountCI: Adds a block that is causing panic spilling
2024-03-25 13:13:48 -04:00
Alyssa Rosenzweig a8b59c16d6 Merge pull request #3513 from Sonicadvance1/panic_spilling_rip
RA: Adds RIP when a block panic spills
2024-03-25 13:13:15 -04:00
Alyssa Rosenzweig 949717a95f OpcodeDispatcher: rewrite DAA implementation
Based on https://www.righto.com/2023/01/

New implementation is branchless, which is theoretically easier to RA. It's also
massively simpler which is good for a demon opcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 13:00:59 -04:00
Alyssa Rosenzweig 693d86dd67 OpcodeDispatcher: add SetAFAndFixup helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 12:59:19 -04:00
Alyssa Rosenzweig ea4fce7a43 InstcountCI: add flagm primary 32-bit
track the demon opcodes

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-25 12:49:00 -04:00
Ryan Houdek 3034edb0aa RA: Adds RIP when a block panic spills
I find myself adding this every time I find a game that panic spills.
Let's just print it out.
2024-03-24 17:11:29 -07:00
Ryan Houdek c025039651 InstcountCI: Adds a block that is causing panic spilling 2024-03-24 17:10:00 -07:00
Ryan Houdek 64f47d1ec2 FEXCore: Adds more TSO control levers
Lets use control vector loadstores and memcpy/memset TSO visibility.
This just gives us a bit more configuration rather than TSO off or on.
2024-03-24 16:34:18 -07:00
Ryan Houdek ea31363221 Linux/Threads: Fixes a stack memory leak for pthreads
Same situation as the last stack leak memory fix, this is fairly tricky
since it is dealing with stack pivoting. Fixes the memory leak around
pthread stack allocations, making memory usage lower for applications
that constantly spin-up and destroy threads (Like Steam).

We need to let glibc allocate a minimum sized stack (128KB and we can't
control it) to work around a race condition with DTV/TLS regions. This
means we need to do a stack pivot once the thread starts executing.

We also need to be careful because the `PThread` object is deleted
inside of the execution thread, which was resulting in a use-after-free
bug.

There are definitely some more memory leaks that I'm still fighting, and I have
noticed in my abusive thread creation program that we might want to
change some jemalloc options to more aggressively cut down on residency.
This is just one out of many.
2024-03-24 05:22:22 -07:00
Ryan Houdek 70befc216f Telemetry: Allow redirecting directory that data is written to
This will be necessary
2024-03-24 00:47:35 -07:00
Ryan Houdek 50f62663ac ELFParser: Stop using a VLA
Clang-18 complains about this, use a vector instead.
2024-03-22 22:51:57 -07:00
Ryan Houdek 60755acef0 FEXLoader: Add some debug-only tracking for FEX owned FDs
I remember seeing some application last year where they closed a FEX
owned FD but now I don't remember what it was. This can really mess us
up so add some debug tracking so we can try and find it again.

Might be something specifically around flatpack, appimage, or chrome's
sandbox. I have some ideas about how to work around these problems if
they crop up but need to find the problem applications again.
2024-03-22 22:49:26 -07:00
Mai 002ca360f8 Merge pull request #3506 from Sonicadvance1/telemetry_rename
Telemetry: Rename old file instead of copying
2024-03-23 00:24:02 -04:00
Ryan Houdek 4952b2e16c Telemetry: Rename old file instead of copying
Since we do an immediate overwrite of the file we are copying, we can
instead do a rename. Failure on rename is fine, will either mean the
telemetry file didn't exist initially, or some other permission error so
the telemetry will get lost regardless.
2024-03-21 22:51:20 -07:00
Ryan Houdek cccf263080 InstCountCI: Update for Telemetry offset changes 2024-03-21 21:10:03 -07:00
Ryan Houdek 5a35e119fe Telemetry: Adds tracker for non-canonical memory access crash
This may be useful for tracking TSO faulting when it manages to fetch
stale data. While most TSO crashes are due to nullptr dereferences, this
can still check for the corruption case.
2024-03-21 20:47:36 -07:00
Ryan Houdek 9ab930cb26 unittests/ASM: Adds tests for loop instruction address size overrides
32-bit test would fail if the 16-bit address size override wasn't
respected.
2024-03-21 20:18:43 -07:00
Ryan Houdek 824f122680 OpcodeDispatcher: Fixes 32-bit mode LOOP RCX register usage
In 64-bit mode, the LOOP instruction's RCX register usage is 64-bit or
32-bit.
In 32-bit mode, the LOOP instruction's RCX register usage is 32-bit or
16-bit.

FEX wasn't handling the 16-bit case at all which was causing the LOOP
instruction to effectively always operate at 32-bit size. Now this is
correctly supported, and it also stops treating the operation as 64-bit.
2024-03-21 20:13:15 -07:00
Ryan Houdek 8852d94416 Merge pull request #3503 from alyssarosenzweig/opt/loop
OpcodeDispatcher: optimize LOOP/N/E
2024-03-21 20:05:50 -07:00
Mai 0c24aea27e Merge pull request #3502 from Sonicadvance1/remove_termux
Removes false termux support
2024-03-21 12:45:01 -04:00
Alyssa Rosenzweig 82ba16c6ed OpcodeDispatcher: optimize LOOP/N/E
Don't clobber NZCV.

Before/after assembly from the Primary_E1 unit test:

< 4340: [INFO] cset w20, ne
< 4340: [INFO] mrs x21, nzcv
< 4340: [INFO] cmp x5, #0x0 (0)
< 4340: [INFO] cset x22, ne
< 4340: [INFO] and x20, x22, x20
< 4340: [INFO] msr nzcv, x21
< 4340: [INFO] cbnz x20, #+0x8 (addr 0xffff896f8084)
< 4340: [INFO] b #+0x1c (addr 0xffff896f809c)
< 4340: [INFO] ldr x0, pc+8 (addr 0xffff896f808c)
---
> 4340: [INFO] csel x20, x5, xzr, ne
> 4340: [INFO] cbnz x20, #+0x8 (addr 0xfffed7308070)
> 4340: [INFO] b #+0x1c (addr 0xfffed7308088)
> 4340: [INFO] ldr x0, pc+8 (addr 0xfffed7308078)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-21 12:08:40 -04:00
Ryan Houdek 45ea0cd782 Removes false termux support
This was a funny joke that this was here, but it is fundamentally
incompatible with what we're doing. All those users are running proot
anyway because of how broken running under termux directly is.

Just remove this from here.
2024-03-20 22:04:32 -07:00
Ryan Houdek 6ce366ef35 Merge pull request #3499 from Sonicadvance1/overlapping_memcpy_unittest
unittests/ASM: Adds a test for overlapping memcpy using rep movs
2024-03-20 19:52:19 -07:00
Ryan Houdek 862d63adf2 unittests/ASM: Adds a test for overlapping memcpy using rep movs
Caused by #3478

This was missed in the review that it could cause issues. bylaws already
has a fix incoming that will get this unit test working.
2024-03-20 18:44:36 -07:00
Ryan Houdek 167896dc9d Merge pull request #3501 from bylaws/memcpy
FEXCore: Fallback to the memcpy slow path for overlaps within 32 bytes
2024-03-20 18:42:56 -07:00
Billy Laws 12fb26f9c0 Update InstCountCI 2024-03-20 21:08:59 +00:00
Billy Laws d490cb1b79 FEXCore: Fallback to the memcpy slow path for overlaps within 32 bytes
Take e.g a forward rep movsb copy from addr 0 to 1, the expected
behaviour since this is a bytewise copy is:
before: aaabbbb...
after: aaaaaaa...
but by copying in 32-byte chunks we end up with:
after: aaaabbbb...
due to the self overwrites not occuring within a single 32 bit copy.
2024-03-20 20:54:19 +00:00
Billy Laws 94fecb9dad FEXCore: Remove needless alignment checks for the mem{cpy,set} fastpath 2024-03-20 20:54:09 +00:00
Ryan Houdek 7dcacfe990 Merge pull request #3478 from bylaws/memcpy
FEXCore: Add non-atomic Memcpy and Memset IR fast paths
2024-03-18 18:56:44 -07:00
Billy Laws 29b05f6b90 Update InstCountCI 2024-03-18 23:30:19 +00:00
Billy Laws 8d4d8fe3e5 FEXCore: Add non-atomic Memcpy and Memset IR fast paths
When TSO is disabled, vector LDP/STP can be used for a two
instruction 32 byte memory copy which is significantly faster than the
current byte-by-byte copy. Performing two such copies directly after
oneanother also marginally increases copy speed for all sizes >=64.
2024-03-18 23:28:50 +00:00
Ryan Houdek ab8ee64352 Merge pull request #3497 from Sonicadvance1/movmaskb_constant
JIT: Optimize pmovmaskb with a named vector constant
2024-03-18 16:08:40 -07:00
Alyssa Rosenzweig 2a9fcc6a66 Merge pull request #3492 from Sonicadvance1/implement_prefetch
OpcodeDispatcher: Implement support for the various prefetch instructions
2024-03-18 07:49:47 -04:00
Ryan Houdek 20da1e4244 InstcountCI: Update for pmovmaskb 2024-03-17 18:52:21 -07:00
Ryan Houdek fd391b1b18 JIT: Optimize pmovmaskb with a named vector constant
I was looking at some other JIT overheads and this cropped up as some
overhead. Instead of materializing a constant using mov+movk+movk+movk,
load it from the named vector constant array.

In a micro-benchmark this improved performance by 34%.
In bytemark this improved on subbench by 0.82%
2024-03-17 18:40:46 -07:00
Mai ba3029b1f6 Merge pull request #3495 from Sonicadvance1/implement_rdpid
OpcodeDispatcher: Implement rdpid
2024-03-15 17:29:12 -04:00
Ryan Houdek 6757a80365 InstcountCI: Update for prefetch changes 2024-03-15 13:20:28 -07:00
Ryan Houdek f79991a9d8 OpcodeDispatcher: Implement rdpid
Missed this instruction when implementing rdtscp. Returns the same ID
result in a register just like rdtscp, but without the cycle counter
results. Doesn't touch any flags just like rdtscp.
2024-03-14 20:07:58 -07:00
Ryan Houdek ca6b2e43e6 Merge pull request #3491 from alyssarosenzweig/rclse/waw
RCLSE: Optimize store-after-store
2024-03-14 03:23:05 -07:00
Ryan Houdek 8a3d08e1d8 Merge pull request #3483 from neobrain/refactor_stealmemoryregion
Allocator: Cleanup StealMemoryRegions implementation
2024-03-14 03:21:09 -07:00
Ryan Houdek cd2a6ce820 Merge pull request #3469 from alyssarosenzweig/opt/df
Optimize DF representation
2024-03-14 03:18:41 -07:00
Ryan Houdek 4e269d8b80 Merge pull request #3494 from neobrain/fix_libfwd_float_as_int
Library Forwarding: Don't map float/double to fixed-size integers
2024-03-14 03:11:11 -07:00
Tony Wasserka 552e76c001 Library Forwarding: Don't map float/double to fixed-size integers
Fixes #3455.
2024-03-14 10:14:57 +01:00
Mai caff3cb799 Merge pull request #3493 from Sonicadvance1/bug_for_3478
unittests/ASM: Implements a unit test for #3478
2024-03-13 22:57:39 -04:00
Ryan Houdek 0d33dacc37 unittests/ASM: Implements a unit test for #3478
This unit test recreates the error condition that #3478 causes.
With a string operation that is a backwards copy then the optimization
will read past the end of the page and result in a crash.

Seemingly only happens with backwards string operations, but test
forward and backward in this test.
2024-03-13 18:36:19 -07:00
Ryan Houdek ba7b69eea2 InstCountCI: Adds prefetch addressing limits 2024-03-12 21:38:28 -07:00
Ryan Houdek 8056bee82b OpcodeDispatcher: Implement support for the various prefetch instructions
x86 has a few prefetch instructions.
- prefetch - One of two classic 3DNow! instructions
   - Prefetch in to L1 data cache
- prefetchw - One of two classic 3DNow! instructions
   - Implies prefetch in to L1 data cache
   - Prefetch cacheline with intent to write and exclusive ownership

- prefetchnta
   - Prefetch non-temporal data in respect to /all/ cache levels
   - Assumes inclusive caches?
- prefetch{t0,t1,t2}
   - Prefetch data with respect to each cache level
   - T0 = L1 and higher
   - T1 = L2 and higher
   - T2 = L3 and higher

**Some silly duplicates**
- prefetchwt1
   - Duplicate of prefetchw but explicitly L1 data cache
- prefetch_exclusive
   - Duplicate of prefetch

God Of War 2018 uses prefetchw as a hint for exclusive ownership of the
cacheline in some very aggressive spin-loops. Let's implement the
operations to help it along.
2024-03-12 21:37:31 -07:00
Ryan Houdek cc635a54f8 IR: Implements support for prefetch operation 2024-03-12 21:19:50 -07:00
Ryan Houdek 217d9d8c50 ARMEmitter: Fixes prfm with negative or unaligned offsets 2024-03-12 21:18:23 -07:00
Tony Wasserka a047ac1699 Allocator: Test CollectMemoryGaps instead of StealMemoryRegions and restore the original interfaces 2024-03-12 10:49:31 +01:00
Tony Wasserka bb0b114fc8 Allocator: Miscellaneous cleanups 2024-03-12 10:49:30 +01:00
Tony Wasserka ccd6c15316 Allocator: Use std::from_chars instead of parsing digits manually 2024-03-12 10:49:30 +01:00
Tony Wasserka 0a1fe1c8c2 Allocator: Parse process mappings per-line instead of per-character 2024-03-12 10:49:30 +01:00
Tony Wasserka f43fe5fd63 Allocator: Stop parsing more eagerly
This is a soft-revert of eaf83aa. That change is no longer needed, since the
stack special case is handled externally now.
2024-03-12 10:49:30 +01:00
Tony Wasserka dce9f651fd Allocator: Split off memory gap collection to a separate function
This function can be unit-tested more easily, and the stack special is more
cleanly handled as a post-collection step.

There is a minor functional change: The stack special case didn't trigger
previously if the range end was within the stack mapping. This is now fixed.
2024-03-12 10:49:30 +01:00
Tony Wasserka 0d71f169d0 Allocator: Adopt a more testable interface for StealMemoryRegions 2024-03-12 10:49:30 +01:00
Tony Wasserka 430ac0f70a Allocator: Fix format strings 2024-03-12 10:49:30 +01:00
Alyssa Rosenzweig 063b81da1d InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig 7629007cfa OpcodeDispatcher: allow upper garbage on STOS
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig 03c6abdad4 OpcodeDispatcher: optimize DF add
fuse the shift the right way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig c99cbe6d0a JIT: switch DF representation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig e3ee65e491 OpcodeDispatcher: use transformed DF for memset/memcpy
Use the 1/-1 representation instead of 0/1. This will be better by the end of
the series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig aee00f524c OpcodeDispatcher: use DF retrieval helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig a76321c6c1 OpcodeDispatcher: add DF retrieval helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Alyssa Rosenzweig f7586f4459 CoreState: use x86 enums for readability
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:50:31 -04:00
Ryan Houdek e33a76a2ad Merge pull request #3489 from Sonicadvance1/linux_v6.8
Linux: Expose support for v6.8
2024-03-11 15:48:37 -07:00
Ryan Houdek c37a12e806 Merge pull request #3490 from Sonicadvance1/disable_assert
Disable assert in release
2024-03-11 15:48:18 -07:00
Alyssa Rosenzweig ed59f73a65 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:41:33 -04:00
Alyssa Rosenzweig 92f31648b9 RCLSE: optimize out pointless stores
can help a lot of x86 code because x86 is 2-address and a64 is 3-address, so x86
ends up with piles of movs that end up dead after translation

It's not a win across the board because our RA isn't aware of tied registers so
sometimes we regress moves. But it's a win on average, and the RA bits can be
improved with time.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:41:23 -04:00
Alyssa Rosenzweig 85f8ad3842 JIT: fix sha256msg1 encoding
botched move in the !tied reg case.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:41:23 -04:00
Alyssa Rosenzweig 64f71d87bb unittests: disable rdtsc test on sim
Patch written by Sonicadvance1. Unclear how this wasn't already broken, but we
need this to keep CI happy with the rest of this series.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-03-11 18:41:23 -04:00
Ryan Houdek ff0c7637c9 Merge pull request #3421 from pmatos/AddressingModes32
Improve 32bit ld/st addressing mode propagation
2024-03-11 15:20:20 -07:00
Ryan Houdek 54403e2146 Disable assert in release
Arguments and conditional doesn't get optimized out in release builds
for the inline function call versus the define.

Was showing up an annoying amount of time when testing.
2024-03-10 22:01:50 -07:00
Ryan Houdek d8202335e0 Linux: Expose support for v6.8
The new syscalls for futexes are the most interesting part
2024-03-10 15:48:55 -07:00
Ryan Houdek 9ec20c4bef Linux/Ioctls: Update ioctl emulation for v6.8
- v3d added an ioctl
- drm base added a new ioctl
- pvr and xe are new drivers in v6.8
2024-03-10 15:46:21 -07:00
Ryan Houdek ce7acd9b71 External/drm-headers: Update to v6.8 2024-03-10 15:45:34 -07:00
Ryan Houdek 8a607135fd Linux: Update syscalls for v6.8 2024-03-10 15:22:51 -07:00
Tony Wasserka 26a66790ab Merge pull request #3486 from neobrain/fix_libfwd_accidental_copy
Library Forwarding: Fix accidental data copying when converting from host to guest layout
2024-03-08 12:40:44 +01:00
Mai 6d94d79409 Merge pull request #3485 from Sonicadvance1/nouveau_ioctl
IoctlEmulation: Add missing nouveau ioctl
2024-03-07 07:57:47 -05:00
Tony Wasserka 2359a9899c Library Forwarding: Fix accidental data copying when converting from host to guest layout 2024-03-07 10:52:13 +01:00
Ryan Houdek aeb41e9ae2 IoctlEmulation: Add missing nouveau ioctl
The NVIF ioctl isn't publicly described in the nouveau headers and it is
required for anything to work with Nouveau.

Pass the ioctl command through without modification and hope that this
ioctl is architecture agnostic.
2024-03-05 16:05:13 -08:00
Tony Wasserka b892da72f3 Merge pull request #3484 from neobrain/feature_catch2_v3
Externals: Update Catch2 to v3.5.3
2024-03-05 19:30:30 +01:00
Paulo Matos a86f2d3e2c Improve 32bit constant usage in memory addressing
Folds reg+const memory address into addressing mode,
if the constant is within 16Kb.
Update instcountci files.
Add test 32Bit_ASM/FEX_bugs/SubAddrBug.asm
2024-03-05 14:01:32 +00:00
Tony Wasserka 6edba49784 Update Catch2 to v3.5.3 2024-03-05 12:15:29 +01:00
580 changed files with 87413 additions and 95844 deletions

No files matched your search

+109
View File
@@ -0,0 +1,109 @@
Language: Cpp
BasedOnStyle: WebKit
AccessModifierOffset: -2
AlignAfterOpenBracket: Align
AlignArrayOfStructures: None
AlignConsecutiveAssignments: None
AlignConsecutiveBitFields: Consecutive
AlignConsecutiveDeclarations: None
AlignConsecutiveMacros: None
AlignEscapedNewlines: Left
AlignOperands: Align
AlignTrailingComments: true
AllowAllParametersOfDeclarationOnNextLine: false
AllowShortCaseLabelsOnASingleLine: true
AllowShortEnumsOnASingleLine: true
AllowShortFunctionsOnASingleLine: Empty
AllowShortIfStatementsOnASingleLine: WithoutElse
AllowShortLambdasOnASingleLine: Inline
AlwaysBreakAfterDefinitionReturnType: None
AlwaysBreakAfterReturnType: None
AlwaysBreakBeforeMultilineStrings: false
AlwaysBreakTemplateDeclarations: true
AttributeMacros:
- JEMALLOC_NOTHROW
- FEX_ALIGNED
- FEX_ANNOTATE
- FEX_DEFAULT_VISIBILITY
- FEX_NAKED
- FEX_PACKED
- FEXCORE_PRESERVE_ALL_ATTR
- GLIBC_ALIAS_FUNCTION
BinPackArguments: true
BinPackParameters: true
BitFieldColonSpacing: Both
BreakAfterAttributes: Always # clang 16 required
BreakBeforeBraces: Attach
BreakBeforeBinaryOperators: None
BreakBeforeInlineASMColon: OnlyMultiline # clang 16 required
BreakBeforeTernaryOperators: false
BreakConstructorInitializers: BeforeComma
BreakInheritanceList: BeforeColon
ColumnLimit: 140
CompactNamespaces: false
ConstructorInitializerIndentWidth: 2
ContinuationIndentWidth: 2
Cpp11BracedListStyle: true
DerivePointerAlignment: false
EmptyLineAfterAccessModifier: Leave
EmptyLineBeforeAccessModifier: Leave
ExperimentalAutoDetectBinPacking: false
FixNamespaceComments: true
IncludeBlocks: Preserve
IndentAccessModifiers: false
IndentCaseBlocks: false
IndentCaseLabels: false
IndentExternBlock: AfterExternBlock
IndentGotoLabels: false
IndentPPDirectives: None
IndentRequires: false
IndentWidth: 2
InsertBraces: true
KeepEmptyLinesAtTheStartOfBlocks: true
LambdaBodyIndentation: OuterScope
LineEnding: LF # clang 16 required
MaxEmptyLinesToKeep: 2
NamespaceIndentation: Inner
QualifierAlignment: Left
PackConstructorInitializers: Never
PenaltyBreakAssignment: 2
PenaltyBreakBeforeFirstCallParameter: 2
PenaltyBreakOpenParenthesis: 2
PenaltyBreakString: 10
PenaltyBreakTemplateDeclaration: 8
PenaltyExcessCharacter: 2
PenaltyReturnTypeOnItsOwnLine: 16
PointerAlignment: Left
RemoveBracesLLVM: false
ReferenceAlignment: Left
ReflowComments: true
RequiresClausePosition: WithPreceding
SeparateDefinitionBlocks: Leave
SortIncludes: Never
SpaceAfterCStyleCast: false
SpaceAfterLogicalNot: false
SpaceAfterTemplateKeyword: false
SpaceAroundPointerQualifiers: Default
SpaceBeforeAssignmentOperators: true
SpaceBeforeCaseColon: false
SpaceBeforeCpp11BracedList: true
SpaceBeforeInheritanceColon: true
SpaceBeforeParens: Custom
SpaceBeforeParensOptions:
AfterControlStatements: true
AfterFunctionDeclarationName: false
AfterFunctionDefinitionName: false
AfterOverloadedOperator: false
AfterRequiresInClause: true
BeforeNonEmptyParentheses: false
SpaceBeforeRangeBasedForLoopColon: true
SpaceBeforeSquareBrackets: false
SpaceInEmptyBlock: false
SpaceInEmptyParentheses: false
SpacesBeforeTrailingComments: 1
SpacesInAngles: Leave
SpacesInCStyleCastParentheses: false
SpacesInConditionalStatement: false
SpacesInParentheses: false
Standard: c++20
UseTab: Never
+12
View File
@@ -0,0 +1,12 @@
# This file is used to ignore files and directories from clang-format
# Ignore all files in the External directory
External/*
# SoftFloat-3e code doesn't belong to us
FEXCore/Source/Common/SoftFloat-3e/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
FEXCore/Source/Interface/Core/X86Tables/*
+15
View File
@@ -0,0 +1,15 @@
# Since version 2.23 (released in August 2019), git-blame has a feature
# to ignore or bypass certain commits.
#
# This file contains a list of commits that are not likely what you
# are looking for in a blame, such as mass reformatting or renaming.
# You can set this file as a default ignore file for blame by running
# the following command.
#
# $ git config blame.ignoreRevsFile .git-blame-ignore-revs
# Whole tree reformat PR#3571
2b4ec88daebd35fefb5bf5c73d7fc2b4155771ed
# Second reformat to find fixed point PR#3577
905aa935f5ce344a48ef4d5edab3c31efa8d793e
+2 -2
View File
@@ -64,7 +64,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -237,7 +237,7 @@ jobs:
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
+2 -2
View File
@@ -71,7 +71,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -171,7 +171,7 @@ jobs:
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Remove old SHM regions
if: ${{ always() }}
+2 -2
View File
@@ -64,7 +64,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -90,7 +90,7 @@ jobs:
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
+2 -2
View File
@@ -74,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -121,7 +121,7 @@ jobs:
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
+1 -1
View File
@@ -74,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+76
View File
@@ -0,0 +1,76 @@
# Inspired by LLVM's pr-code-format.yml at
# https://github.com/llvm/llvm-project/blob/main/.github/workflows/pr-code-format.yml
name: "Check code formatting"
on:
pull_request:
branches:
- main
jobs:
code_formatter:
runs-on: [self-hosted, X64]
if: github.repository == 'FEX-Emu/FEX'
steps:
- name: Fetch FEX sources
uses: actions/checkout@v4
with:
ref: ${{ github.event.pull_request.head.sha }}
- name: Checkout through merge base
uses: rmacklin/fetch-through-merge-base@v0
timeout-minutes: 3
with:
base_ref: ${{ github.event.pull_request.base.ref }}
head_ref: ${{ github.event.pull_request.head.sha }}
deepen_length: 500
- name: Get changed files
id: changed-files
uses: tj-actions/changed-files@v39
with:
separator: ","
skip_initial_fetch: true
- name: "Listed files"
env:
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
run: |
echo "Formatting files:"
echo "$CHANGED_FILES"
- name: Check for correct clang-format version
run: clang-format --version | grep -qF '16.0.6'
- name: Check git-clang-format-16 exists
run: which git-clang-format-16
- name: Setup Python env
uses: actions/setup-python@v4
with:
python-version: '3.11'
cache: 'pip'
cache-dependency-path: './External/code-format-helper/requirements_formatting.txt'
- name: Install python dependencies
run: pip install -r ./External/code-format-helper/requirements_formatting.txt
- name: Run code formatter
env:
CLANG_FORMAT_PATH: 'git-clang-format-16'
GITHUB_PR_NUMBER: ${{ github.event.pull_request.number }}
START_REV: ${{ github.event.pull_request.base.sha }}
END_REV: ${{ github.event.pull_request.head.sha }}
CHANGED_FILES: ${{ steps.changed-files.outputs.all_changed_files }}
# TODO(pmatos): Once we adopt v18, we should be able
# to take advantage of the new --diff_from_common_commit option
# explicitly in code-format-helper.py and not have to diff starting at
# the merge base.
run: |
python ./External/code-format-helper/code-format-helper.py \
--repo "FEX-emu/FEX" \
--issue-number $GITHUB_PR_NUMBER \
--start-rev $(git merge-base $START_REV $END_REV) \
--end-rev $END_REV \
--changed-files "$CHANGED_FILES"
+2 -2
View File
@@ -65,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -106,7 +106,7 @@ jobs:
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
run: truncate --size="<20M" ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
+4 -67
View File
@@ -9,7 +9,6 @@ option(BUILD_FEX_LINUX_TESTS "Build FEXLinuxTests, requires x86 compiler" FALSE)
option(BUILD_THUNKS "Build thunks" FALSE)
option(BUILD_FEXCONFIG "Build FEXConfig, requires SDL2 and X11" TRUE)
option(ENABLE_CLANG_THUNKS "Build thunks with clang" FALSE)
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
@@ -26,7 +25,6 @@ option(ENABLE_OFFLINE_TELEMETRY "Enables FEX offline telemetry" TRUE)
option(ENABLE_COMPILE_TIME_TRACE "Enables time trace compile option" FALSE)
option(ENABLE_LIBCXX "Enables LLVM libc++" FALSE)
option(ENABLE_CCACHE "Enables ccache for compile caching" TRUE)
option(ENABLE_TERMUX_BUILD "Forces building for Termux on a non-Termux build machine" FALSE)
option(ENABLE_VIXL_SIMULATOR "Forces the FEX JIT to use the VIXL simulator" FALSE)
option(ENABLE_VIXL_DISASSEMBLER "Enables debug disassembler output with VIXL" FALSE)
option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FALSE)
@@ -125,6 +123,9 @@ endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^arm64ec")
set(_M_ARM_64EC 1)
add_definitions(-D_M_ARM_64EC=1)
# Required as FEX is not allowed to lock the CRT heap lock during compilation or callbacks
set(ENABLE_JEMALLOC TRUE)
endif()
if (ENABLE_CCACHE)
@@ -163,18 +164,6 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if(DEFINED ENV{TERMUX_VERSION} OR ENABLE_TERMUX_BUILD)
add_definitions(-DTERMUX_BUILD=1)
set(TERMUX_BUILD 1)
# Termux doesn't support Jemalloc due to bad interactions between emutls, jemalloc, and scudo
set(ENABLE_JEMALLOC FALSE)
# Termux builds can't rely on X11 packages
# SDL2 isn't even compiled with GL support so our GUIs wouldn't even work
set(BUILD_FEXCONFIG FALSE)
endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
@@ -246,8 +235,6 @@ add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
option(CATCH_BUILD_STATIC_LIBRARY "" ON)
set(CATCH_BUILD_STATIC_LIBRARY ON)
add_subdirectory(External/Catch2/)
# Pull in catch_discover_tests definition
@@ -357,57 +344,6 @@ if (ENABLE_IWYU)
endif()
endif()
if (ENABLE_CLANG_FORMAT)
find_program(CLANG_TIDY_EXE "clang-tidy")
if (NOT CLANG_TIDY_EXE)
message(FATAL_ERROR "Couldn't find clang-tidy")
endif()
set(CLANG_TIDY_FLAGS
"-checks=*"
"-fuchsia*"
"-bugprone-macro-parentheses"
"-clang-analyzer-core.*"
"-cppcoreguidelines-pro-type-*"
"-cppcoreguidelines-pro-bounds-array-to-pointer-decay"
"-cppcoreguidelines-pro-bounds-pointer-arithmetic"
"-cppcoreguidelines-avoid-c-arrays"
"-cppcoreguidelines-avoid-magic-numbers"
"-cppcoreguidelines-pro-bounds-constant-array-index"
"-cppcoreguidelines-no-malloc"
"-cppcoreguidelines-special-member-functions"
"-cppcoreguidelines-owning-memory"
"-cppcoreguidelines-macro-usage"
"-cppcoreguidelines-avoid-goto"
"-google-readability-function-size"
"-google-readability-namespace-comments"
"-google-readability-braces-around-statements"
"-google-build-using-namespace"
"-hicpp-*"
"-llvm-namespace-comment"
"-llvm-include-order" # Messes up with case sensitivity
"-llvmlibc-*"
"-misc-unused-parameters"
"-modernize-loop-convert"
"-modernize-use-auto"
"-modernize-avoid-c-arrays"
"-modernize-use-nodiscard"
"readability-*"
"-readability-function-size"
"-readability-implicit-bool-conversion"
"-readability-braces-around-statements"
"-readability-else-after-return"
"-readability-magic-numbers"
"-readability-named-parameter"
"-readability-uppercase-literal-suffix"
"-cert-err34-c"
"-cert-err58-cpp"
"-bugprone-exception-escape"
)
string(REPLACE ";" "," CLANG_TIDY_FLAGS "${CLANG_TIDY_FLAGS}")
set(CMAKE_CXX_CLANG_TIDY ${CLANG_TIDY_EXE} "${CLANG_TIDY_FLAGS}")
endif()
add_compile_options(-Wall)
configure_file(
@@ -421,6 +357,7 @@ if (BUILD_TESTS)
endif()
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
# Binfmt_misc files must be installed prior to Source/ installs
+2
View File
@@ -0,0 +1,2 @@
add_library(CodeEmitter INTERFACE)
target_include_directories(CodeEmitter INTERFACE .)
File diff suppressed because it is too large. Load diff
@@ -8,11 +8,11 @@ public:
public:
// Conditional branch immediate
///< Branch conditional
void b(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void b(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
}
void b(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
void b(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
@@ -20,13 +20,13 @@ public:
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
void b(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
void b(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
void b(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
}
@@ -36,11 +36,11 @@ public:
}
///< Branch consistent conditional
void bc(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void bc(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
void bc(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
@@ -49,13 +49,13 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bc(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
void bc(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
}
@@ -65,7 +65,7 @@ public:
}
// Unconditional branch register
void br(FEXCore::ARMEmitter::Register rn) {
void br(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'000 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -74,7 +74,7 @@ public:
UnconditionalBranch(Op, rn);
}
void blr(FEXCore::ARMEmitter::Register rn) {
void blr(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'001 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -83,7 +83,7 @@ public:
UnconditionalBranch(Op, rn);
}
void ret(FEXCore::ARMEmitter::Register rn = FEXCore::ARMEmitter::Reg::r30) {
void ret(ARMEmitter::Register rn = ARMEmitter::Reg::r30) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'010 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -156,13 +156,13 @@ public:
}
// Compare and branch
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -173,7 +173,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0100 << 24;
@@ -181,7 +181,7 @@ public:
CompareAndBranch(Op, s, rt, 0);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
}
@@ -190,13 +190,13 @@ public:
}
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -207,7 +207,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0101 << 24;
@@ -215,7 +215,7 @@ public:
CompareAndBranch(Op, s, rt, 0);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
}
@@ -225,12 +225,12 @@ public:
}
// Test and branch immediate
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -241,7 +241,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0110 << 24;
@@ -249,7 +249,7 @@ public:
TestAndBranch(Op, rt, Bit, 0);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
}
@@ -258,12 +258,12 @@ public:
}
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -274,14 +274,14 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
}
@@ -292,7 +292,7 @@ public:
private:
// Conditional branch immediate
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, ARMEmitter::Condition Cond, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Op1 << 24;
@@ -304,7 +304,7 @@ private:
}
// Unconditional branch register
void UnconditionalBranch(uint32_t Op, FEXCore::ARMEmitter::Register rn) {
void UnconditionalBranch(uint32_t Op, ARMEmitter::Register rn) {
uint32_t Instr = Op;
Instr |= Encode_rn(rn);
dc32(Instr);
@@ -318,8 +318,8 @@ private:
}
// Compare and branch
void CompareAndBranch(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
void CompareAndBranch(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -330,7 +330,7 @@ private:
}
// Test and branch - immediate
void TestAndBranch(uint32_t Op, FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void TestAndBranch(uint32_t Op, ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= (Bit >> 5) << 31;
+106
View File
@@ -0,0 +1,106 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
#include <cstdint>
#include <cstring>
namespace ARMEmitter {
class Buffer {
public:
Buffer() {
SetBuffer(nullptr, 0);
}
Buffer(uint8_t* Base, uint64_t BaseSize) {
SetBuffer(Base, BaseSize);
}
void SetBuffer(uint8_t* Base, uint64_t BaseSize) {
BufferBase = Base;
CurrentOffset = BufferBase;
Size = BaseSize;
}
void dc8(uint8_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc16(uint16_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc32(uint32_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc64(uint64_t Data) {
decltype(Data)* Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void EmitString(const char* String) {
const auto StringLength = strlen(String);
memcpy(CurrentOffset, String, StringLength);
CurrentOffset += StringLength;
}
void Align() {
// Align the buffer to instruction size
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & 0b11;
if (!CurrentAlignment) {
return;
}
CurrentOffset += 4 - CurrentAlignment;
}
template<typename T>
T GetCursorAddress() const {
return reinterpret_cast<T>(CurrentOffset);
}
static void ClearICache(void* Begin, std::size_t Length) {
__builtin___clear_cache(static_cast<char*>(Begin), static_cast<char*>(Begin) + Length);
}
size_t GetCursorOffset() const {
return static_cast<size_t>(CurrentOffset - BufferBase);
}
uint8_t* GetBufferBase() const {
return BufferBase;
}
void CursorIncrement(size_t Size) {
CurrentOffset += Size;
}
void SetCursorOffset(size_t Offset) {
CurrentOffset = BufferBase + Offset;
}
uint64_t GetBufferSize() const {
return Size;
}
template<typename T>
size_t GetCursorOffsetFromAddress(const T* Address) const {
return static_cast<size_t>(reinterpret_cast<const uint8_t*>(Address) - BufferBase);
}
protected:
void ResetBuffer() {
CurrentOffset = BufferBase;
}
uint8_t* BufferBase;
uint8_t* CurrentOffset;
uint64_t Size;
};
} // namespace ARMEmitter
+831
View File
@@ -0,0 +1,831 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <CodeEmitter/Buffer.h>
#include <CodeEmitter/Registers.h>
#include <array>
#include <cstdint>
#include <utility>
#include <type_traits>
/*
* Welcome to FEX-Emu's custom AArch64 emitter.
* This was written specifically to avoid the performance cost of the vixl emitter.
*
* There are some specific design constraints in this design to target a couple features:
* - High performance
* - Low CPU cache performance hit
* - Significantly reduced code footprint
* - Low number of branches
*
* These requirements are mostly achieved by removing a bunch of developer conveniences
* that vixl provides. The developer needs to take a lot of care to not shoot themselves in the foot.
*
* Misc design decisions:
* - Registers are encoded as basic uint32_t enums.
* - Converting between different registers is zero-cost.
* - Passing around as arguments are as cheap as registers
* - Contrast to vixl where every register requires living on the stack.
* - Registers can get encoded in to instructions with a simple `BFM` instruction.
*
* - Instructions are very simply emitted, allowing direct inlining most of the time.
* - These are simple enough that multiple back-to-back instructions get optimized to 128-bit load-store operations.
* - Contrast to vixl where pretty much no instruction emitter gets inlined.
*
* - Instruction emitters are /mostly/ unsized. Most instructions take a size argument first, which gets encoded
* directly in to the instruction.
* - Contrast to vixl where the register arguments are how the instructions determine operating size.
* - Size argument allows FEX to use `CSEL` to select a size at runtime, instead of branching.
* - Some instructions are explicitly sized based on register type. Read comments in the respective `inl` files to
* see why.
* Some scalar/vector operations are an example of this.
*
* - Almost zero helper functions.
* - Primary exception to this rule is load-store operations. These will use a helper to make
* it easier to select the correct load-store instruction. Mostly because these are a nightmare selecting
* the right instruction.
*/
namespace ARMEmitter {
/*
* This `Size` enum is used for most ALU operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class Size : uint32_t {
i32Bit = 0,
i64Bit,
};
// This allows us to get the `Size` enum in bits.
[[nodiscard]]
constexpr size_t RegSizeInBits(Size size) {
return size_t {32} << FEXCore::ToUnderlying(size);
}
/* This `SubRegSize` enum is used for most ASIMD operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class SubRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
i128Bit = 0b100,
};
// This allows us to get the `SubRegSize` in bits.
[[nodiscard]]
constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t {8} << FEXCore::ToUnderlying(size);
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
* This is specifically duplicated from `SubRegSize` to have strongly
* typed functions.
*
* `ScalarRegSize` specifically doesn't have `i128Bit` because scalar operations
* can't operate at 128-bit.
*/
enum class ScalarRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
};
// This allows us to get the `ScalarRegSize` in bits.
[[nodiscard]]
constexpr size_t ScalarRegSizeInBits(ScalarRegSize size) {
return size_t {8} << FEXCore::ToUnderlying(size);
}
/* This `VectorRegSizePair` union allows us to have an overlapping type
* to select a scalar operation or a vector depending on which operation
* we pass in.
* Useful in FEX's vector operations that behave as scalar or vector
* depending on various factors. But since the operation will have the sa,e
* element size, we want to choose the operation more easily
*/
union VectorRegSizePair {
ScalarRegSize Scalar;
SubRegSize Vector;
};
// This allows us to create a `VectorRegSizePair` union.
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(SubRegSize size) {
return VectorRegSizePair {.Vector = size};
}
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(ScalarRegSize size) {
return VectorRegSizePair {.Scalar = size};
}
// This `ShiftType` enum is used for ALU shift-register encoded instructions.
enum class ShiftType : uint32_t {
LSL = 0,
LSR,
ASR,
ROR,
};
// This `ExtendedType` enum is used for ALU extended-register encoded instructions.
enum class ExtendedType : uint32_t {
UXTB = 0b000,
UXTH = 0b001,
UXTW = 0b010,
UXTX = 0b011,
SXTB = 0b100,
SXTH = 0b101,
SXTW = 0b110,
SXTX = 0b111,
LSL_32 = UXTW,
LSL_64 = UXTX,
};
// This `Condition` enum is used for various conditional instructions.
enum class Condition : uint32_t {
// Meaning: Int - Float
CC_EQ = 0, // Equal - Equal
CC_NE, // Not Eq - Not Eq or unordered
CC_CS, // Carry set - Greater than, equal, or unordered
CC_CC, // Carry clear - Less than
CC_MI, // Minus/Negative - Less than
CC_PL, // Plus, positive or zero - GT, equal, or unordered
CC_VS, // Overflow - Unordered
CC_VC, // No Overflow - Ordered
CC_HI, // Unsigned higher - GT, or unordered
CC_LS, // Unsigned lower or same - LT or EQ
CC_GE, // Signed GT or EQ - GT or EQ
CC_LT, // Signed LT - LT or Unordered
CC_GT, // Signed GT - GT
CC_LE, // Signed LT or EQ - LT, EQ, or Unordered
CC_AL, // Always - Always
CC_NV, // Always - Always
// Aliases
CC_HS = CC_CS,
CC_LO = CC_CC,
};
/*
* This `StatusFlags` enum is used for conditional compare encoded instructions.
* These directly encode to the `nzcv` flags.
*/
enum class StatusFlags : uint32_t {
None = 0,
Flag_V = 0b0001,
Flag_C = 0b0010,
Flag_Z = 0b0100,
Flag_N = 0b1000,
Flag_NZCV = Flag_N | Flag_Z | Flag_C | Flag_V,
};
/*
* This `IndexType` enum is used for load-store instructions.
* Not all load-store instructions use this, so the user needs to be careful.
*/
enum class IndexType {
POST,
OFFSET,
PRE,
UNPRIVILEGED,
};
// Used with adr and scalar + vector load/store variants to denote
// a modifier operation.
enum class SVEModType : uint8_t {
MOD_UXTW,
MOD_SXTW,
MOD_LSL,
MOD_NONE,
};
/* This `SVEMemOperand` class is used for the helper SVE load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class SVEMemOperand final {
public:
enum class Type {
ScalarPlusScalar,
ScalarPlusImm,
ScalarPlusVector,
VectorPlusImm,
};
SVEMemOperand(XRegister rn, XRegister rm = XReg::zr)
: rn {rn}
, MemType {Type::ScalarPlusScalar}
, MetaType {.ScalarScalarType {
.rm = rm,
}} {}
SVEMemOperand(XRegister rn, int32_t imm = 0)
: rn {rn}
, MemType {Type::ScalarPlusImm}
, MetaType {.ScalarImmType {
.Imm = imm,
}} {}
SVEMemOperand(XRegister rn, ZRegister zm, SVEModType mod = SVEModType::MOD_NONE, uint8_t scale = 0)
: rn {rn}
, MemType {Type::ScalarPlusVector}
, MetaType {.ScalarVectorType {
.zm = zm,
.mod = mod,
.scale = scale,
}} {}
SVEMemOperand(ZRegister zn, uint32_t imm)
: rn {Register {zn.Idx()}}
, MemType {Type::VectorPlusImm}
, MetaType {.VectorImmType {
.Imm = imm,
}} {}
[[nodiscard]]
bool IsScalarPlusScalar() const {
return MemType == Type::ScalarPlusScalar;
}
[[nodiscard]]
bool IsScalarPlusImm() const {
return MemType == Type::ScalarPlusImm;
}
[[nodiscard]]
bool IsScalarPlusVector() const {
return MemType == Type::ScalarPlusVector;
}
[[nodiscard]]
bool IsVectorPlusImm() const {
return MemType == Type::VectorPlusImm;
}
union Data {
struct {
Register rm;
} ScalarScalarType;
struct {
int32_t Imm;
} ScalarImmType;
struct {
ZRegister zm;
SVEModType mod;
uint8_t scale;
} ScalarVectorType;
struct {
// rn will be a ZRegister
uint32_t Imm;
} VectorImmType;
};
Register rn;
Type MemType;
Data MetaType;
};
/* This `ExtendedMemOperand` class is used for the helper load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {.ExtendedType {
.Header = {.MemType = TYPE_EXTENDED},
.rm = rm,
.Option = Option,
.Shift = Shift,
}} {}
ExtendedMemOperand(XRegister rn, IndexType Index = IndexType::OFFSET, int32_t Imm = 0)
: rn {rn}
, MetaType {.ImmType {
.Header = {.MemType = TYPE_IMM},
.Index = Index,
.Imm = Imm,
}} {}
Register rn;
enum Type {
TYPE_EXTENDED,
TYPE_IMM,
};
struct HeaderStruct {
Type MemType;
};
union {
HeaderStruct Header;
struct {
HeaderStruct Header;
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
struct {
HeaderStruct Header;
IndexType Index;
int32_t Imm;
} ImmType;
} MetaType;
};
template<uint32_t op0, uint32_t op1, uint32_t CRn, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenSystemReg() {
return op0 << 19 | op1 << 16 | CRn << 12 | CRm << 8 | op2 << 5;
};
// This `SystemRegister` enum is used for the mrs/msr instructions.
enum class SystemRegister : uint32_t {
CTR_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b001>(),
DCZID_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b111>(),
TPIDR_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b010>(),
RNDR = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b000>(),
RNDRRS = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b001>(),
NZCV = GenSystemReg<0b11, 0b011, 0b0100, 0b0010, 0b000>(),
FPCR = GenSystemReg<0b11, 0b011, 0b0100, 0b0100, 0b000>(),
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>(),
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>(),
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenDCReg() {
return op1 << 16 | CRm << 8 | op2 << 5;
};
// This `DataCacheOperation` enum is used for the dc instruction.
enum class DataCacheOperation : uint32_t {
IVAC = GenDCReg<0b000, 0b0110, 0b001>(),
ISW = GenDCReg<0b000, 0b0110, 0b010>(),
CSW = GenDCReg<0b000, 0b1010, 0b010>(),
CISW = GenDCReg<0b000, 0b1110, 0b010>(),
ZVA = GenDCReg<0b011, 0b0100, 0b001>(),
CVAC = GenDCReg<0b011, 0b1010, 0b001>(),
CVAU = GenDCReg<0b011, 0b1011, 0b001>(),
CIVAC = GenDCReg<0b011, 0b1110, 0b001>(),
// MTE2
IGVAC = GenDCReg<0b000, 0b0110, 0b011>(),
IGSW = GenDCReg<0b000, 0b0110, 0b100>(),
IGDVAC = GenDCReg<0b000, 0b0110, 0b101>(),
IGDSW = GenDCReg<0b000, 0b0110, 0b110>(),
CGSW = GenDCReg<0b000, 0b1010, 0b100>(),
CGDSW = GenDCReg<0b000, 0b1010, 0b110>(),
CIGSW = GenDCReg<0b000, 0b1110, 0b100>(),
CIGDSW = GenDCReg<0b000, 0b1110, 0b110>(),
// MTE
GVA = GenDCReg<0b011, 0b0100, 0b011>(),
GZVA = GenDCReg<0b011, 0b0100, 0b100>(),
CGVAC = GenDCReg<0b011, 0b1010, 0b011>(),
CGDVAC = GenDCReg<0b011, 0b1010, 0b101>(),
CGVAP = GenDCReg<0b011, 0b1100, 0b011>(),
CGDVAP = GenDCReg<0b011, 0b1100, 0b101>(),
CGVADP = GenDCReg<0b011, 0b1101, 0b011>(),
CGDVADP = GenDCReg<0b011, 0b1101, 0b101>(),
CIGVAC = GenDCReg<0b011, 0b1110, 0b011>(),
CIGDVAC = GenDCReg<0b011, 0b1110, 0b101>(),
// DPB
CVAP = GenDCReg<0b011, 0b1100, 0b001>(),
// DPB2
CVADP = GenDCReg<0b011, 0b1101, 0b001>(),
};
template<uint32_t CRm, uint32_t op2>
constexpr uint32_t GenHintBarrierReg() {
return CRm << 8 | op2 << 5;
}
// This `HintRegister` enum is used for the hint instruction.
enum class HintRegister : uint32_t {
NOP = GenHintBarrierReg<0b0000, 0b000>(),
YIELD = GenHintBarrierReg<0b0000, 0b001>(),
WFE = GenHintBarrierReg<0b0000, 0b010>(),
WFI = GenHintBarrierReg<0b0000, 0b011>(),
SEV = GenHintBarrierReg<0b0000, 0b100>(),
SEVL = GenHintBarrierReg<0b0000, 0b101>(),
DGH = GenHintBarrierReg<0b0000, 0b110>(),
CSDB = GenHintBarrierReg<0b0010, 0b100>(),
};
// This `BarrierRegister` enum is used for the various barrier instructions.
enum class BarrierRegister : uint32_t {
CLREX = GenHintBarrierReg<0b0000, 0b010>(),
TCOMMIT = GenHintBarrierReg<0b0000, 0b011>(),
DSB = GenHintBarrierReg<0b0000, 0b100>(),
DMB = GenHintBarrierReg<0b0000, 0b101>(),
ISB = GenHintBarrierReg<0b0000, 0b110>(),
SB = GenHintBarrierReg<0b0000, 0b111>(),
};
// This `BarrierScope` enum is used for the dsb/dmb instructions.
enum class BarrierScope : uint32_t {
// Outer shareable
OSHLD = 0b0001,
OSHST = 0b0010,
OSH = 0b0011,
// Non shareable
NSHLD = 0b0101,
NSHST = 0b0110,
NSH = 0b0111,
// Inner shareable
ISHLD = 0b1001,
ISHST = 0b1010,
ISH = 0b1011,
// Full System visibility
LD = 0b1101,
ST = 0b1110,
SY = 0b1111,
};
// This `Prefetch` enum is used for prefetch instructions.
enum class Prefetch : uint32_t {
// Prefetch for load
PLDL1KEEP = 0b00000,
PLDL1STRM = 0b00001,
PLDL2KEEP = 0b00010,
PLDL2STRM = 0b00011,
PLDL3KEEP = 0b00100,
PLDL3STRM = 0b00101,
// Preload instructions
PLIL1KEEP = 0b01000,
PLIL1STRM = 0b01001,
PLIL2KEEP = 0b01010,
PLIL2STRM = 0b01011,
PLIL3KEEP = 0b01100,
PLIL3STRM = 0b01101,
// Preload for store
PSTL1KEEP = 0b10000,
PSTL1STRM = 0b10001,
PSTL2KEEP = 0b10010,
PSTL2STRM = 0b10011,
PSTL3KEEP = 0b10100,
PSTL3STRM = 0b10101,
};
// This `PredicatePattern` enun is used for some SVE instructions.
enum class PredicatePattern : uint32_t {
SVE_POW2 = 0b00000,
SVE_VL1 = 0b00001,
SVE_VL2 = 0b00010,
SVE_VL3 = 0b00011,
SVE_VL4 = 0b00100,
SVE_VL5 = 0b00101,
SVE_VL6 = 0b00110,
SVE_VL7 = 0b00111,
SVE_VL8 = 0b01000,
SVE_VL16 = 0b01001,
SVE_VL32 = 0b01010,
SVE_VL64 = 0b01011,
SVE_VL128 = 0b01100,
SVE_VL256 = 0b01101,
SVE_MUL4 = 0b11101,
SVE_MUL3 = 0b11110,
SVE_ALL = 0b11111,
};
// Used with SVE FP immediate arithmetic instructions
enum class SVEFAddSubImm : uint32_t {
_0_5,
_1_0,
};
enum class SVEFMulImm : uint32_t {
_0_5,
_2_0,
};
enum class SVEFMaxMinImm : uint32_t {
_0_0,
_1_0,
};
/* This `BackwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `below` an instruction that uses it.
* Which means that a branch would jump backwards.
*/
struct BackwardLabel {
uint8_t* Location {};
};
/* This `SingleUseForwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `above` an instruction that uses it.
* Which means that a branch would jump forwards.
*
* The `ForwardLabel` struct can be bound to multiple instructions, so it needs a vector for each bind instruction type.
*/
struct SingleUseForwardLabel {
enum class InstType {
UNKNOWN,
ADR,
ADRP,
B,
BC,
TEST_BRANCH,
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
uint8_t* Location {};
InstType Type = InstType::UNKNOWN;
};
struct ForwardLabel {
fextl::vector<SingleUseForwardLabel> Insts {};
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is in either direction of an instruction that uses it.
* Which means a branch could jump backwards or forwards depending on situation.
*/
struct BiDirectionalLabel {
BackwardLabel Backward;
ForwardLabel Forward;
};
static inline void AddLocationToLabel(SingleUseForwardLabel* Label, SingleUseForwardLabel&& Location) {
LOGMAN_THROW_A_FMT(Label->Type == SingleUseForwardLabel::InstType::UNKNOWN, "Trying to bind a SingleUseForwardLabel to multiple "
"locations. Use ForwardLabel instead.");
*Label = std::move(Location);
}
static inline void AddLocationToLabel(ForwardLabel* Label, SingleUseForwardLabel&& Location) {
Label->Insts.emplace_back(std::move(Location));
}
// Some FCMA ASIMD instructions support a rotation argument.
enum class Rotation : uint32_t {
ROTATE_0 = 0b00,
ROTATE_90 = 0b01,
ROTATE_180 = 0b10,
ROTATE_270 = 0b11,
};
// Concept for contraining some instructions to accept only an XRegister or WRegister.
// Particularly for operations that differ encodings depending on which one is used.
template<typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
// For example, a set of registers like:
//
// v1, v2, v3 and
// v31, v0, v1
//
// would both be considered sequential sequences, and some instructions in particular
// limit register lists to these kind of sequences.
//
template<typename T, typename... Args>
constexpr bool AreVectorsSequential(T first, const Args&... args) {
// Ensure we always have a pair of registers to compare against.
static_assert(sizeof...(args) >= 1, "Number of arguments must be greater than 1");
const auto fn = [](auto& lhs, const auto& rhs) {
const auto result = ((lhs.Idx() + 1) % 32) == rhs.Idx();
lhs = rhs;
return result;
};
return (fn(first, args) && ...);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
// Choices:
// - Size of ops passed as an argument rather than template to let the compiler use csel instead of branching.
// - Registers are unsized so they can be passed in a GPR and not need conversion operations
class Emitter : public ARMEmitter::Buffer {
public:
Emitter() = default;
Emitter(uint8_t* Base, uint64_t BaseSize)
: Buffer(Base, BaseSize) {}
// Bind a backward label to an address.
// Address that is bound is the current emitter location.
void Bind(BackwardLabel* Label) {
LOGMAN_THROW_AA_FMT(Label->Location == nullptr, "Trying to bind a label twice");
Label->Location = GetCursorAddress<uint8_t*>();
}
void Bind(const SingleUseForwardLabel* Label) {
uint8_t* CurrentAddress = GetCursorAddress<uint8_t*>();
// Patch up the instructions
switch (Label->Type) {
case SingleUseForwardLabel::InstType::ADR: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::ADRP: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::B: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= Offset;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::TEST_BRANCH: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::BC:
case SingleUseForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t* Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t* Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
} else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
} else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
} else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
break;
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
}
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
template<bool WarnAboutEmpty = false>
void Bind(ForwardLabel* Label) {
if constexpr (WarnAboutEmpty) {
LOGMAN_THROW_A_FMT(Label->Insts.empty() == false, "Binding forward label that didn't have any instructions using it");
}
for (auto& Inst : Label->Insts) {
Bind(&Inst);
}
}
// Bind a bidirectional location to a location.
// Binds both forwards and backwards depending on how the label was used.
void Bind(BiDirectionalLabel* Label) {
if (!Label->Backward.Location) {
Bind(&Label->Backward);
}
Bind<false>(&Label->Forward);
}
#include <CodeEmitter/VixlUtils.inl>
public:
// TODO: Implement SME when it matters.
#include <CodeEmitter/ALUOps.inl>
#include <CodeEmitter/BranchOps.inl>
#include <CodeEmitter/LoadstoreOps.inl>
#include <CodeEmitter/SystemOps.inl>
#include <CodeEmitter/ScalarOps.inl>
#include <CodeEmitter/ASIMDOps.inl>
#include <CodeEmitter/SVEOps.inl>
private:
template<typename T>
uint32_t Encode_ra(T Reg) const {
return Reg.Idx() << 10;
}
uint32_t Encode_ra(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rt2(T Reg) const {
return Reg.Idx() << 10;
}
template<>
uint32_t Encode_rt2(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rm(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rm(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rs(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rs(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rn(T Reg) const {
return Reg.Idx() << 5;
}
uint32_t Encode_rn(uint32_t Reg) const {
return Reg << 5;
}
template<typename T>
uint32_t Encode_rd(T Reg) const {
return Reg.Idx();
}
uint32_t Encode_rd(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_rt(T Reg) const {
return Reg.Idx();
}
template<>
uint32_t Encode_rt(Prefetch Reg) const {
return FEXCore::ToUnderlying(Reg);
}
uint32_t Encode_rt(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_pd(T Reg) const {
return FEXCore::ToUnderlying(Reg);
}
};
} // namespace ARMEmitter
File diff suppressed because it is too large. Load diff
@@ -1506,13 +1506,14 @@ public:
}
// SVE broadcast floating-point immediate (unpredicated)
void fdup(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_AA_FMT(size == FEXCore::ARMEmitter::SubRegSize::i16Bit ||
size == FEXCore::ARMEmitter::SubRegSize::i32Bit ||
size == FEXCore::ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_AA_FMT(size == ARMEmitter::SubRegSize::i16Bit ||
size == ARMEmitter::SubRegSize::i32Bit ||
size == ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
uint32_t Imm{};
if (size == SubRegSize::i16Bit) {
Imm = FP16ToImm8(vixl::Float16(Value));
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
} else if (size == SubRegSize::i32Bit) {
Imm = FP32ToImm8(Value);
} else if (size == SubRegSize::i64Bit) {
@@ -1521,7 +1522,7 @@ public:
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
}
void fmov(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
fdup(size, zd, Value);
}
@@ -3513,7 +3514,8 @@ private:
size == SubRegSize::i64Bit, "Unsupported fcpy/fmov size");
uint32_t imm{};
if (size == SubRegSize::i16Bit) {
imm = FP16ToImm8(vixl::Float16(value));
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
} else if (size == SubRegSize::i32Bit) {
imm = FP32ToImm8(value);
} else if (size == SubRegSize::i64Bit) {
@@ -3717,7 +3719,7 @@ private:
// SVE bitwise logical operations (predicated)
void SVEBitwiseLogicalPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zdn, ZRegister zm, ZRegister zd) {
LOGMAN_THROW_AA_FMT(size != FEXCore::ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
LOGMAN_THROW_A_FMT(zd == zdn, "zd needs to equal zdn");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
@@ -4743,7 +4745,7 @@ private:
dc32(Instr);
}
void SVEPermuteVector(uint32_t op0, FEXCore::ARMEmitter::ZRegister zd, FEXCore::ARMEmitter::ZRegister zm, uint32_t Imm) {
void SVEPermuteVector(uint32_t op0, ARMEmitter::ZRegister zd, ARMEmitter::ZRegister zm, uint32_t Imm) {
constexpr uint32_t Op = 0b0000'0101'0010'0000'000 << 13;
uint32_t Instr = Op;
@@ -5228,15 +5230,14 @@ private:
// Alias that returns the equivalently sized unsigned type for a floating-point type T.
template <typename T>
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, vixl::Float16>, uint16_t,
std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>>;
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>;
// Determines if a floating-point value is capable of being converted
// into an 8-bit immediate. See pseudocode definition of VFPExpandImm
// in ARM A-profile reference manual for a general overview of how this was derived.
template <typename T>
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
[[nodiscard, maybe_unused]] static bool IsValidFPValueForImm8(T value) {
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
@@ -5277,18 +5278,6 @@ private:
return true;
}
static uint32_t FP16ToImm8(vixl::Float16 value) {
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
"Value cannot be encoded into an 8-bit immediate");
const uint32_t bits = vixl::Float16ToRawbits(value);
const uint32_t sign = (bits & 0x8000) >> 8;
const uint32_t expb2 = (bits & 0x2000) >> 7;
const uint32_t b5_to_0 = (bits >> 6) & 0x3F;
return sign | expb2 | b5_to_0;
}
static uint32_t FP32ToImm8(float value) {
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
"Value ({}) cannot be encoded into an 8-bit immediate", value);
@@ -33,7 +33,7 @@ public:
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
}
void mov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn, uint32_t Index) {
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
dup(size, rd, rn, Index);
}
@@ -1052,21 +1052,21 @@ public:
}
// Floating-point immediate
void fmov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, float Value) {
void fmov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, float Value) {
uint32_t M = 0;
uint32_t S = 0;
uint32_t ptype;
uint32_t imm8;
uint32_t imm5 = 0b0'0000;
if (size == FEXCore::ARMEmitter::ScalarRegSize::i16Bit) {
ptype = 0b11;
imm8 = FP16ToImm8(vixl::Float16(Value));
if (size == ARMEmitter::ScalarRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
}
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i32Bit) {
else if (size == ARMEmitter::ScalarRegSize::i32Bit) {
ptype = 0b00;
imm8 = FP32ToImm8(Value);
}
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i64Bit) {
else if (size == ARMEmitter::ScalarRegSize::i64Bit) {
ptype = 0b01;
imm8 = FP64ToImm8(Value);
}
@@ -1077,7 +1077,7 @@ public:
FloatScalarImmediate(M, S, ptype, imm8, imm5, rd);
}
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, FEXCore::ARMEmitter::VRegister rd) {
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, ARMEmitter::VRegister rd) {
constexpr uint32_t Op = 0b0001'1110'0010'0000'0001'00 << 10;
uint32_t Instr = Op;
@@ -1286,7 +1286,7 @@ public:
private:
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn) {
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
Instr |= Q << 30;
@@ -11,7 +11,7 @@ public:
// TODO: AT
// TODO: CFP
// TODO: CPP
void dc(FEXCore::ARMEmitter::DataCacheOperation DCOp, FEXCore::ARMEmitter::Register rt) {
void dc(ARMEmitter::DataCacheOperation DCOp, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
}
@@ -48,67 +48,67 @@ public:
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
}
// System instructions with register argument
void wfet(FEXCore::ARMEmitter::Register rt) {
void wfet(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b000, rt);
}
void wfit(FEXCore::ARMEmitter::Register rt) {
void wfit(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b001, rt);
}
// Hints
void nop() {
Hint(FEXCore::ARMEmitter::HintRegister::NOP);
Hint(ARMEmitter::HintRegister::NOP);
}
void yield() {
Hint(FEXCore::ARMEmitter::HintRegister::YIELD);
Hint(ARMEmitter::HintRegister::YIELD);
}
void wfe() {
Hint(FEXCore::ARMEmitter::HintRegister::WFE);
Hint(ARMEmitter::HintRegister::WFE);
}
void wfi() {
Hint(FEXCore::ARMEmitter::HintRegister::WFI);
Hint(ARMEmitter::HintRegister::WFI);
}
void sev() {
Hint(FEXCore::ARMEmitter::HintRegister::SEV);
Hint(ARMEmitter::HintRegister::SEV);
}
void sevl() {
Hint(FEXCore::ARMEmitter::HintRegister::SEVL);
Hint(ARMEmitter::HintRegister::SEVL);
}
void dgh() {
Hint(FEXCore::ARMEmitter::HintRegister::DGH);
Hint(ARMEmitter::HintRegister::DGH);
}
void csdb() {
Hint(FEXCore::ARMEmitter::HintRegister::CSDB);
Hint(ARMEmitter::HintRegister::CSDB);
}
// Barriers
void clrex(uint32_t imm = 15) {
LOGMAN_THROW_AA_FMT(imm < 16, "Immediate out of range");
Barrier(FEXCore::ARMEmitter::BarrierRegister::CLREX, imm);
Barrier(ARMEmitter::BarrierRegister::CLREX, imm);
}
void dsb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
void dsb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
}
void dmb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
void dmb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
}
void isb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(FEXCore::ARMEmitter::BarrierScope::SY));
Barrier(ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(ARMEmitter::BarrierScope::SY));
}
void sb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::SB, 0);
Barrier(ARMEmitter::BarrierRegister::SB, 0);
}
void tcommit() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::TCOMMIT, 0);
Barrier(ARMEmitter::BarrierRegister::TCOMMIT, 0);
}
// System register move
void msr(FEXCore::ARMEmitter::SystemRegister reg, FEXCore::ARMEmitter::Register rt) {
void msr(ARMEmitter::SystemRegister reg, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
SystemRegisterMove(Op, rt, reg);
}
void mrs(FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::SystemRegister reg) {
void mrs(ARMEmitter::Register rd, ARMEmitter::SystemRegister reg) {
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
SystemRegisterMove(Op, rd, reg);
}
@@ -130,7 +130,7 @@ private:
}
// System instructions with register argument
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, FEXCore::ARMEmitter::Register rt) {
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, ARMEmitter::Register rt) {
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
Instr |= CRm << 8;
@@ -140,13 +140,13 @@ private:
}
// Hints
void Hint(FEXCore::ARMEmitter::HintRegister Reg) {
void Hint(ARMEmitter::HintRegister Reg) {
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Barriers
void Barrier(FEXCore::ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
void Barrier(ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
Instr |= CRm << 8;
Instr |= FEXCore::ToUnderlying(Reg);
@@ -154,7 +154,7 @@ private:
}
// System Instruction
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, FEXCore::ARMEmitter::Register rt) {
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, ARMEmitter::Register rt) {
uint32_t Instr = Op;
Instr |= L << 21;
@@ -165,7 +165,7 @@ private:
}
// System register move
void SystemRegisterMove(uint32_t Op, FEXCore::ARMEmitter::Register rt, FEXCore::ARMEmitter::SystemRegister reg) {
void SystemRegisterMove(uint32_t Op, ARMEmitter::Register rt, ARMEmitter::SystemRegister reg) {
uint32_t Instr = Op;
Instr |= FEXCore::ToUnderlying(reg);
+311
View File
@@ -0,0 +1,311 @@
// Collection of utilities from vixl.
// Following is the vixl license.
// Copyright 2015, VIXL authors
// All rights reserved.
//
// Redistribution and use in source and binary forms, with or without
// modification, are permitted provided that the following conditions are met:
//
// * Redistributions of source code must retain the above copyright notice,
// this list of conditions and the following disclaimer.
// * Redistributions in binary form must reproduce the above copyright notice,
// this list of conditions and the following disclaimer in the documentation
// and/or other materials provided with the distribution.
// * Neither the name of ARM Limited nor the names of its contributors may be
// used to endorse or promote products derived from this software without
// specific prior written permission.
//
// THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS CONTRIBUTORS "AS IS" AND
// ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
// WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
// DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
// FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
// DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
// SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
// CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
// OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
// OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
// Test if a given value can be encoded in the immediate field of a logical
// instruction.
// If it can be encoded, the function returns true, and values pointed to by n,
// imm_s and imm_r are updated with immediates encoded in the format required
// by the corresponding fields in the logical instruction.
// If it can not be encoded, the function returns false, and the values pointed
// to by n, imm_s and imm_r are undefined.
static bool IsImmLogical(uint64_t value,
unsigned width,
unsigned* n,
unsigned* imm_s,
unsigned* imm_r) {
[[maybe_unused]] constexpr auto kBRegSize = 8;
[[maybe_unused]] constexpr auto kHRegSize = 16;
[[maybe_unused]] constexpr auto kSRegSize = 32;
[[maybe_unused]] constexpr auto kDRegSize = 64;
constexpr auto kWRegSize = 32;
constexpr auto kXRegSize = 64;
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) ||
(width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
bool negate = false;
// Logical immediates are encoded using parameters n, imm_s and imm_r using
// the following table:
//
// N imms immr size S R
// 1 ssssss rrrrrr 64 UInt(ssssss) UInt(rrrrrr)
// 0 0sssss xrrrrr 32 UInt(sssss) UInt(rrrrr)
// 0 10ssss xxrrrr 16 UInt(ssss) UInt(rrrr)
// 0 110sss xxxrrr 8 UInt(sss) UInt(rrr)
// 0 1110ss xxxxrr 4 UInt(ss) UInt(rr)
// 0 11110s xxxxxr 2 UInt(s) UInt(r)
// (s bits must not be all set)
//
// A pattern is constructed of size bits, where the least significant S+1 bits
// are set. The pattern is rotated right by R, and repeated across a 32 or
// 64-bit value, depending on destination register width.
//
// Put another way: the basic format of a logical immediate is a single
// contiguous stretch of 1 bits, repeated across the whole word at intervals
// given by a power of 2. To identify them quickly, we first locate the
// lowest stretch of 1 bits, then the next 1 bit above that; that combination
// is different for every logical immediate, so it gives us all the
// information we need to identify the only logical immediate that our input
// could be, and then we simply check if that's the value we actually have.
//
// (The rotation parameter does give the possibility of the stretch of 1 bits
// going 'round the end' of the word. To deal with that, we observe that in
// any situation where that happens the bitwise NOT of the value is also a
// valid logical immediate. So we simply invert the input whenever its low bit
// is set, and then we know that the rotated case can't arise.)
if (value & 1) {
// If the low bit is 1, negate the value, and set a flag to remember that we
// did (so that we can adjust the return values appropriately).
negate = true;
value = ~value;
}
if (width <= kWRegSize) {
// To handle 8/16/32-bit logical immediates, the very easiest thing is to repeat
// the input value to fill a 64-bit word. The correct encoding of that as a
// logical immediate will also be the correct encoding of the value.
// Avoid making the assumption that the most-significant 56/48/32 bits are zero by
// shifting the value left and duplicating it.
for (unsigned bits = width; bits <= kWRegSize; bits *= 2) {
value <<= bits;
uint64_t mask = (UINT64_C(1) << bits) - 1;
value |= ((value >> bits) & mask);
}
}
// The basic analysis idea: imagine our input word looks like this.
//
// 0011111000111110001111100011111000111110001111100011111000111110
// c b a
// |<--d-->|
//
// We find the lowest set bit (as an actual power-of-2 value, not its index)
// and call it a. Then we add a to our original number, which wipes out the
// bottommost stretch of set bits and replaces it with a 1 carried into the
// next zero bit. Then we look for the new lowest set bit, which is in
// position b, and subtract it, so now our number is just like the original
// but with the lowest stretch of set bits completely gone. Now we find the
// lowest set bit again, which is position c in the diagram above. Then we'll
// measure the distance d between bit positions a and c (using CLZ), and that
// tells us that the only valid logical immediate that could possibly be equal
// to this number is the one in which a stretch of bits running from a to just
// below b is replicated every d bits.
uint64_t a = LowestSetBit(value);
uint64_t value_plus_a = value + a;
uint64_t b = LowestSetBit(value_plus_a);
uint64_t value_plus_a_minus_b = value_plus_a - b;
uint64_t c = LowestSetBit(value_plus_a_minus_b);
int d, clz_a, out_n;
uint64_t mask;
if (c != 0) {
// The general case, in which there is more than one stretch of set bits.
// Compute the repeat distance d, and set up a bitmask covering the basic
// unit of repetition (i.e. a word with the bottom d bits set). Also, in all
// of these cases the N bit of the output will be zero.
clz_a = CountLeadingZeros(a, kXRegSize);
int clz_c = CountLeadingZeros(c, kXRegSize);
d = clz_a - clz_c;
mask = ((UINT64_C(1) << d) - 1);
out_n = 0;
} else {
// Handle degenerate cases.
//
// If any of those 'find lowest set bit' operations didn't find a set bit at
// all, then the word will have been zero thereafter, so in particular the
// last lowest_set_bit operation will have returned zero. So we can test for
// all the special case conditions in one go by seeing if c is zero.
if (a == 0) {
// The input was zero (or all 1 bits, which will come to here too after we
// inverted it at the start of the function), for which we just return
// false.
return false;
} else {
// Otherwise, if c was zero but a was not, then there's just one stretch
// of set bits in our word, meaning that we have the trivial case of
// d == 64 and only one 'repetition'. Set up all the same variables as in
// the general case above, and set the N bit in the output.
clz_a = CountLeadingZeros(a, kXRegSize);
d = 64;
mask = ~UINT64_C(0);
out_n = 1;
}
}
// If the repeat period d is not a power of two, it can't be encoded.
if (!IsPowerOf2(d)) {
return false;
}
if (((b - a) & ~mask) != 0) {
// If the bit stretch (b - a) does not fit within the mask derived from the
// repeat period, then fail.
return false;
}
// The only possible option is b - a repeated every d bits. Now we're going to
// actually construct the valid logical immediate derived from that
// specification, and see if it equals our original input.
//
// To repeat a value every d bits, we multiply it by a number of the form
// (1 + 2^d + 2^(2d) + ...), i.e. 0x0001000100010001 or similar. These can
// be derived using a table lookup on CLZ(d).
static const uint64_t multipliers[] = {
0x0000000000000001UL,
0x0000000100000001UL,
0x0001000100010001UL,
0x0101010101010101UL,
0x1111111111111111UL,
0x5555555555555555UL,
};
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
uint64_t candidate = (b - a) * multiplier;
if (value != candidate) {
// The candidate pattern doesn't match our input value, so fail.
return false;
}
// We have a match! This is a valid logical immediate, so now we have to
// construct the bits and pieces of the instruction encoding that generates
// it.
// Count the set bits in our basic stretch. The special case of clz(0) == -1
// makes the answer come out right for stretches that reach the very top of
// the word (e.g. numbers like 0xffffc00000000000).
int clz_b = (b == 0) ? -1 : CountLeadingZeros(b, kXRegSize);
int s = clz_a - clz_b;
// Decide how many bits to rotate right by, to put the low bit of that basic
// stretch in position a.
int r;
if (negate) {
// If we inverted the input right at the start of this function, here's
// where we compensate: the number of set bits becomes the number of clear
// bits, and the rotation count is based on position b rather than position
// a (since b is the location of the 'lowest' 1 bit after inversion).
s = d - s;
r = (clz_b + 1) & (d - 1);
} else {
r = (clz_a + 1) & (d - 1);
}
// Now we're done, except for having to encode the S output in such a way that
// it gives both the number of set bits and the length of the repeated
// segment. The s field is encoded like this:
//
// imms size S
// ssssss 64 UInt(ssssss)
// 0sssss 32 UInt(sssss)
// 10ssss 16 UInt(ssss)
// 110sss 8 UInt(sss)
// 1110ss 4 UInt(ss)
// 11110s 2 UInt(s)
//
// So we 'or' (2 * -d) with our computed s to form imms.
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
*n = out_n;
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
*imm_r = r;
}
return true;
}
private:
template <typename V>
static inline bool IsPowerOf2(V value) {
return (value != 0) && ((value & (value - 1)) == 0);
}
// Some compilers dislike negating unsigned integers,
// so we provide an equivalent.
template <typename T>
static inline T UnsignedNegate(T value) {
static_assert(std::is_unsigned<T>::value);
return ~value + 1;
}
static inline uint64_t LowestSetBit(uint64_t value) {
return value & UnsignedNegate(value);
}
template <typename V>
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
#if COMPILER_HAS_BUILTIN_CLZ
if (width == 32) {
return (value == 0) ? 32 : __builtin_clz(static_cast<unsigned>(value));
} else if (width == 64) {
return (value == 0) ? 64 : __builtin_clzll(value);
}
#endif
return CountLeadingZerosFallBack(value, width);
}
static inline int CountLeadingZerosFallBack(uint64_t value, int width) {
LOGMAN_THROW_A_FMT(IsPowerOf2(width) && (width <= 64), "Invalid width");
if (value == 0) {
return width;
}
int count = 0;
value = value << (64 - width);
if ((value & UINT64_C(0xffffffff00000000)) == 0) {
count += 32;
value = value << 32;
}
if ((value & UINT64_C(0xffff000000000000)) == 0) {
count += 16;
value = value << 16;
}
if ((value & UINT64_C(0xff00000000000000)) == 0) {
count += 8;
value = value << 8;
}
if ((value & UINT64_C(0xf000000000000000)) == 0) {
count += 4;
value = value << 4;
}
if ((value & UINT64_C(0xc000000000000000)) == 0) {
count += 2;
value = value << 2;
}
if ((value & UINT64_C(0x8000000000000000)) == 0) {
count += 1;
}
count += (value == 0);
return count;
}
public:
+6
View File
@@ -0,0 +1,6 @@
{
"Comment": "Bypasses libGL's glX and instead sends GLX requests directly via xcb",
"ThunksDB": {
"GL": 0
}
}
-9
View File
@@ -2,9 +2,6 @@
"DB": {
"GL": {
"Library" : "libGL-guest.so",
"Depends": [
"X11"
],
"Overlay": [
"@PREFIX_LIB@/libGL.so",
"@PREFIX_LIB@/libGL.so.1",
@@ -33,16 +30,10 @@
},
"Vulkan": {
"Library": "libvulkan-guest.so",
"Depends": [
"xcb"
],
"Overlay": [
"@PREFIX_LIB@/libvulkan.so",
"@PREFIX_LIB@/libvulkan.so.1",
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
],
"Comment": [
"Vulkan library relies on xcb, otherwise it crashes with jemalloc"
]
},
"xcb": {
+1 -1
+404
View File
@@ -0,0 +1,404 @@
#!/usr/bin/env python3
#
# ====- code-format-helper, runs code formatters from the ci or in a hook --*- python -*--==#
#
# Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
# See https://llvm.org/LICENSE.txt for license information.
# SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
#
# ==--------------------------------------------------------------------------------------==#
import argparse
import os
import subprocess
import sys
from typing import List, Optional
"""
This script is run by GitHub actions to ensure that the code in PR's conform to
the coding style of LLVM. It can also be installed as a pre-commit git hook to
check the coding style before submitting it. The canonical source of this script
is in the LLVM source tree under llvm/utils/git.
For C/C++ code it uses clang-format and for Python code it uses darker (which
in turn invokes black).
You can learn more about the LLVM coding style on llvm.org:
https://llvm.org/docs/CodingStandards.html
You can install this script as a git hook by symlinking it to the .git/hooks
directory:
ln -s $(pwd)/llvm/utils/git/code-format-helper.py .git/hooks/pre-commit
You can control the exact path to clang-format or darker with the following
environment variables: $CLANG_FORMAT_PATH and $DARKER_FORMAT_PATH.
"""
class FormatArgs:
start_rev: str = None
end_rev: str = None
repo: str = None
changed_files: List[str] = []
token: str = None
verbose: bool = True
issue_number: int = 0
write_comment_to_file: str = None
def __init__(self, args: argparse.Namespace = None) -> None:
if not args is None:
self.start_rev = args.start_rev
self.end_rev = args.end_rev
self.repo = args.repo
self.token = args.token
self.changed_files = args.changed_files
self.issue_number = args.issue_number
self.write_comment_to_file = args.write_comment_to_file
class FormatHelper:
COMMENT_TAG = "<!--CODE FORMAT COMMENT: {fmt}-->"
name: str
friendly_name: str
comment: dict = None
@property
def comment_tag(self) -> str:
return self.COMMENT_TAG.replace("fmt", self.name)
@property
def instructions(self) -> str:
raise NotImplementedError()
def has_tool(self) -> bool:
raise NotImplementedError()
def format_run(self, changed_files: List[str], args: FormatArgs) -> Optional[str]:
raise NotImplementedError()
def pr_comment_text_for_diff(self, diff: str) -> str:
return f"""
:warning: {self.friendly_name}, {self.name} found issues in your code. :warning:
<details>
<summary>
You can test this locally with the following command:
</summary>
``````````bash
{self.instructions}
``````````
</details>
<details>
<summary>
View the diff from {self.name} here.
</summary>
``````````diff
{diff}
``````````
</details>
"""
# TODO: any type should be replaced with the correct github type, but it requires refactoring to
# not require the github module to be installed everywhere.
def find_comment(self, pr: any) -> any:
for comment in pr.as_issue().get_comments():
if self.comment_tag in comment.body:
return comment
return None
def update_pr(self, comment_text: str, args: FormatArgs, create_new: bool) -> None:
import github
from github import IssueComment, PullRequest
repo = github.Github(args.token).get_repo(args.repo)
pr = repo.get_issue(args.issue_number).as_pull_request()
comment_text = self.comment_tag + "\n\n" + comment_text
existing_comment = self.find_comment(pr)
if args.write_comment_to_file:
if create_new or existing_comment:
self.comment = {"body": comment_text}
if existing_comment:
self.comment["id"] = existing_comment.id
return
if existing_comment:
existing_comment.edit(comment_text)
elif create_new:
pr.as_issue().create_comment(comment_text)
def run(self, changed_files: List[str], args: FormatArgs) -> bool:
changed_files = [arg for arg in changed_files if "third-party" not in arg]
diff = self.format_run(changed_files, args)
should_update_gh = args.token is not None and args.repo is not None
if diff is None:
if should_update_gh:
comment_text = (
":white_check_mark: With the latest revision "
f"this PR passed the {self.friendly_name}."
)
self.update_pr(comment_text, args, create_new=False)
return True
elif len(diff) > 0:
if should_update_gh:
comment_text = self.pr_comment_text_for_diff(diff)
self.update_pr(comment_text, args, create_new=True)
else:
print(
f"Warning: {self.friendly_name}, {self.name} detected "
"some issues with your code formatting..."
)
return False
else:
# The formatter failed but didn't output a diff (e.g. some sort of
# infrastructure failure).
comment_text = (
f":warning: The {self.friendly_name} failed without printing "
"a diff. Check the logs for stderr output. :warning:"
)
self.update_pr(comment_text, args, create_new=False)
return False
class ClangFormatHelper(FormatHelper):
name = "clang-format"
friendly_name = "C/C++ code formatter"
@property
def cformat_wrapper_path(self) -> str:
relpath = "../../Scripts/clang-format.py"
curpath = os.path.dirname(os.path.abspath(__file__))
return os.path.abspath(os.path.normpath(os.path.join(curpath, relpath)))
@property
def instructions(self) -> str:
return " ".join(self.cf_cmd)
def should_include_extensionless_file(self, path: str) -> bool:
return path.startswith("libcxx/include")
def filter_changed_files(self, changed_files: List[str]) -> List[str]:
filtered_files = []
for path in changed_files:
_, ext = os.path.splitext(path)
if ext in (".cpp", ".c", ".h", ".hpp", ".hxx", ".cxx", ".inc", ".cppm"):
filtered_files.append(path)
elif ext == "" and self.should_include_extensionless_file(path):
filtered_files.append(path)
return filtered_files
@property
def clang_fmt_path(self) -> str:
if "CLANG_FORMAT_PATH" in os.environ:
return os.environ["CLANG_FORMAT_PATH"]
return "git-clang-format"
def has_tool(self) -> bool:
cmd = [self.clang_fmt_path, "-h"]
proc = None
try:
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
except:
return False
return proc.returncode == 0
def format_run(self, changed_files: List[str], args: FormatArgs) -> Optional[str]:
cpp_files = self.filter_changed_files(changed_files)
if not cpp_files:
return None
cf_cmd = [
self.clang_fmt_path,
f"--binary={self.cformat_wrapper_path}",
"--diff",
]
if args.start_rev and args.end_rev:
cf_cmd.append(args.start_rev)
cf_cmd.append(args.end_rev)
cf_cmd.append("--")
cf_cmd += cpp_files
if args.verbose:
print(f"Running: {' '.join(cf_cmd)}")
self.cf_cmd = cf_cmd
proc = subprocess.run(cf_cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
sys.stdout.write(proc.stderr.decode("utf-8"))
if proc.returncode != 0:
# formatting needed, or the command otherwise failed
if args.verbose:
print(f"error: {self.name} exited with code {proc.returncode}")
# Print the diff in the log so that it is viewable there
print(proc.stdout.decode("utf-8"))
return proc.stdout.decode("utf-8")
else:
return None
class DarkerFormatHelper(FormatHelper):
name = "darker"
friendly_name = "Python code formatter"
@property
def instructions(self) -> str:
return " ".join(self.darker_cmd)
def filter_changed_files(self, changed_files: List[str]) -> List[str]:
filtered_files = []
for path in changed_files:
name, ext = os.path.splitext(path)
if ext == ".py":
filtered_files.append(path)
return filtered_files
@property
def darker_fmt_path(self) -> str:
if "DARKER_FORMAT_PATH" in os.environ:
return os.environ["DARKER_FORMAT_PATH"]
return "darker"
def has_tool(self) -> bool:
cmd = [self.darker_fmt_path, "--version"]
proc = None
try:
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
except:
return False
return proc.returncode == 0
def format_run(self, changed_files: List[str], args: FormatArgs) -> Optional[str]:
py_files = self.filter_changed_files(changed_files)
if not py_files:
return None
darker_cmd = [
self.darker_fmt_path,
"--check",
"--diff",
]
if args.start_rev and args.end_rev:
darker_cmd += ["-r", f"{args.start_rev}...{args.end_rev}"]
darker_cmd += py_files
if args.verbose:
print(f"Running: {' '.join(darker_cmd)}")
self.darker_cmd = darker_cmd
proc = subprocess.run(
darker_cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE
)
if args.verbose:
sys.stdout.write(proc.stderr.decode("utf-8"))
if proc.returncode != 0:
# formatting needed, or the command otherwise failed
if args.verbose:
print(f"error: {self.name} exited with code {proc.returncode}")
# Print the diff in the log so that it is viewable there
print(proc.stdout.decode("utf-8"))
return proc.stdout.decode("utf-8")
else:
sys.stdout.write(proc.stdout.decode("utf-8"))
return None
ALL_FORMATTERS = (DarkerFormatHelper(), ClangFormatHelper())
def hook_main():
# fill out args
args = FormatArgs()
args.verbose = False
# find the changed files
cmd = ["git", "diff", "--cached", "--name-only", "--diff-filter=d"]
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
output = proc.stdout.decode("utf-8")
for line in output.splitlines():
args.changed_files.append(line)
failed_fmts = []
for fmt in ALL_FORMATTERS:
if fmt.has_tool():
if not fmt.run(args.changed_files, args):
failed_fmts.append(fmt.name)
if fmt.comment:
comments.append(fmt.comment)
else:
print(f"Couldn't find {fmt.name}, can't check " + fmt.friendly_name.lower())
if len(failed_fmts) > 0:
sys.exit(1)
sys.exit(0)
if __name__ == "__main__":
script_path = os.path.abspath(__file__)
if ".git/hooks" in script_path:
hook_main()
sys.exit(0)
parser = argparse.ArgumentParser()
parser.add_argument(
"--token", type=str, required=False, help="GitHub authentication token"
)
parser.add_argument(
"--repo",
type=str,
default=os.getenv("GITHUB_REPOSITORY", "llvm/llvm-project"),
help="The GitHub repository that we are working with in the form of <owner>/<repo> (e.g. llvm/llvm-project)",
)
parser.add_argument("--issue-number", type=int, required=True)
parser.add_argument(
"--start-rev",
type=str,
required=True,
help="Compute changes from this revision.",
)
parser.add_argument(
"--end-rev", type=str, required=True, help="Compute changes to this revision"
)
parser.add_argument(
"--changed-files",
type=str,
help="Comma separated list of files that has been changed",
)
parser.add_argument(
"--write-comment-to-file",
type=str,
help="Don't post comments on the PR, instead write the comments and metadata a file",
)
args = FormatArgs(parser.parse_args())
changed_files = []
if args.changed_files:
changed_files = args.changed_files.split(",")
failed_formatters = []
comments = []
for fmt in ALL_FORMATTERS:
if not fmt.run(changed_files, args):
failed_formatters.append(fmt.name)
if fmt.comment:
comments.append(fmt.comment)
if len(comments):
with open(args.write_comment_to_file, "w") as f:
import json
json.dump(comments, f)
if len(failed_formatters) > 0:
print(f"error: some formatters failed: {' '.join(failed_formatters)}")
sys.exit(1)
+52
View File
@@ -0,0 +1,52 @@
#
# This file is autogenerated by pip-compile with Python 3.11
# by the following command:
#
# pip-compile --output-file=llvm/utils/git/requirements_formatting.txt llvm/utils/git/requirements_formatting.txt.in
#
black==23.9.1
# via
# -r llvm/utils/git/requirements_formatting.txt.in
# darker
certifi==2023.7.22
# via requests
cffi==1.15.1
# via
# cryptography
# pynacl
charset-normalizer==3.2.0
# via requests
click==8.1.7
# via black
cryptography==41.0.3
# via pyjwt
darker==1.7.2
# via -r llvm/utils/git/requirements_formatting.txt.in
deprecated==1.2.14
# via pygithub
idna==3.4
# via requests
mypy-extensions==1.0.0
# via black
packaging==23.1
# via black
pathspec==0.11.2
# via black
platformdirs==3.10.0
# via black
pycparser==2.21
# via cffi
pygithub==1.59.1
# via -r llvm/utils/git/requirements_formatting.txt.in
pyjwt[crypto]==2.8.0
# via pygithub
pynacl==1.5.0
# via pygithub
requests==2.31.0
# via pygithub
toml==0.10.2
# via darker
urllib3==2.0.4
# via requests
wrapt==1.15.0
# via deprecated
-2
View File
@@ -13,8 +13,6 @@ if (CMAKE_SYSTEM_PROCESSOR MATCHES "^aarch64|^arm64|^armv8\.*")
set(_M_ARM_64 1)
endif()
option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
set(CMAKE_POSITION_INDEPENDENT_CODE ON)
cmake_policy(SET CMP0083 NEW) # Follow new PIE policy
include(CheckPIESupported)
+44 -12
View File
@@ -55,6 +55,7 @@ class OpDefinition:
DynamicDispatch: bool
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
Arguments: list
EmitValidation: list
Desc: list
@@ -77,6 +78,7 @@ class OpDefinition:
self.DynamicDispatch = False
self.JITDispatch = True
self.JITDispatchOverride = None
self.TiedSource = -1
self.Arguments = []
self.EmitValidation = []
self.Desc = []
@@ -248,6 +250,9 @@ def parse_ops(ops):
if "JITDispatchOverride" in op_val:
OpDef.JITDispatchOverride = op_val["JITDispatchOverride"]
if "TiedSource" in op_val:
OpDef.TiedSource = op_val["TiedSource"]
# Do some fixups of the data here
if len(OpDef.EmitValidation) != 0:
for i in range(len(OpDef.EmitValidation)):
@@ -372,13 +377,30 @@ def print_ir_sizes():
output_file.write("[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] std::string_view const& GetName(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetRAArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool ImplicitFlagClobber(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool GetHasDest(IROps Op);\n")
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);\n'
)
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
@@ -471,15 +493,25 @@ def print_ir_getraargs():
def print_ir_hassideeffects():
output_file.write("#ifdef IROP_HASSIDEEFFECTS_IMPL\n")
for array, prop in [("SideEffects", "HasSideEffects"),
("ImplicitFlagClobbers", "ImplicitFlagClobber")]:
output_file.write(f"constexpr std::array<uint8_t, OP_LAST + 1> {array} = {{\n")
for array, prop, T in [
("SideEffects", "HasSideEffects", "bool"),
("ImplicitFlagClobbers", "ImplicitFlagClobber", "bool"),
("TiedSources", "TiedSource", "int8_t"),
]:
output_file.write(
f"constexpr std::array<{'uint8_t' if T == 'bool' else T}, OP_LAST + 1> {array} = {{\n"
)
for op in IROps:
output_file.write("\t{},\n".format(("true" if getattr(op, prop) else "false")))
if T == "bool":
output_file.write(
"\t{},\n".format(("true" if getattr(op, prop) else "false"))
)
else:
output_file.write(f"\t{getattr(op, prop)},\n")
output_file.write("};\n\n")
output_file.write(f"bool {prop}(IROps Op) {{\n")
output_file.write(f"{T} {prop}(IROps Op) {{\n")
output_file.write(f" return {array}[Op];\n")
output_file.write("}\n")
+4 -7
View File
@@ -134,22 +134,16 @@ set (SRCS
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/DeadCodeElimination.cpp
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/InlineCallOptimization.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
@@ -194,12 +188,15 @@ endif()
# Some defines for the softfloat library
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils)
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils CodeEmitter)
if (NOT MINGW_BUILD)
list (APPEND LIBS dl)
else()
list (APPEND LIBS synchronization)
if (_M_ARM_64EC)
list (APPEND LIBS kernelbase)
endif()
endif()
if (ENABLE_JEMALLOC)
+11 -9
View File
@@ -18,14 +18,14 @@ struct BitSet final {
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
ElementType *Memory;
ElementType* Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
@@ -43,10 +43,13 @@ struct BitSet final {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, ToBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, ToBytes(Elements));
}
uint32_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / MinimumSize;
}
// This very explicitly doesn't let you take an address
@@ -62,11 +65,10 @@ struct BitSetView final {
constexpr static size_t MinimumSize = sizeof(ElementType);
constexpr static size_t MinimumSizeBits = sizeof(ElementType) * 8;
ElementType *Memory;
ElementType* Memory;
void GetView(BitSet<T> &Set, uint64_t ElementOffset) {
LOGMAN_THROW_AA_FMT((ElementOffset % MinimumSize) == 0,
"Bitset view offset needs to be aligned to size of backing element");
void GetView(BitSet<T>& Set, uint64_t ElementOffset) {
LOGMAN_THROW_AA_FMT((ElementOffset % MinimumSize) == 0, "Bitset view offset needs to be aligned to size of backing element");
Memory = &Set.Memory[ElementOffset / MinimumSizeBits];
}
+116 -107
View File
@@ -7,133 +7,142 @@
#include <unistd.h>
namespace FEXCore {
JITSymbols::JITSymbols() {
}
JITSymbols::JITSymbols() {}
JITSymbols::~JITSymbols() {
if (fd != -1) {
close(fd);
}
JITSymbols::~JITSymbols() {
if (fd != -1) {
close(fd);
}
}
void JITSymbols::InitFile() {
// We can't use FILE here since we must be robust against forking processes closing our FD from under us.
void JITSymbols::InitFile() {
// We can't use FILE here since we must be robust against forking processes closing our FD from under us.
#ifdef __ANDROID__
// Android simpleperf looks in /data/local/tmp instead of /tmp
const auto PerfMap = fextl::fmt::format("/data/local/tmp/perf-{}.map", getpid());
// Android simpleperf looks in /data/local/tmp instead of /tmp
const auto PerfMap = fextl::fmt::format("/data/local/tmp/perf-{}.map", getpid());
#else
const auto PerfMap = fextl::fmt::format("/tmp/perf-{}.map", getpid());
const auto PerfMap = fextl::fmt::format("/tmp/perf-{}.map", getpid());
#endif
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
fd = open(PerfMap.c_str(), O_CREAT | O_TRUNC | O_WRONLY | O_APPEND, 0644);
}
void JITSymbols::RegisterNamedRegion(const void* HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) {
return;
}
void JITSymbols::RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}\n", HostAddr, CodeSize, Name);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} {}\n", HostAddr, CodeSize, Name);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
void JITSymbols::RegisterJITSpace(const void* HostAddr, uint32_t CodeSize) {
if (fd == -1) {
return;
}
void JITSymbols::RegisterJITSpace(const void *HostAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} FEXJIT\n", HostAddr, CodeSize);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
}
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto Buffer = fextl::fmt::format("{} {:x} FEXJIT\n", HostAddr, CodeSize);
auto Result = write(fd, Buffer.c_str(), Buffer.size());
if (Result == -1 && errno == EBADF) {
fd = -1;
}
// Buffered JIT symbols.
void JITSymbols::Register(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) {
return;
}
// Buffered JIT symbols.
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, GuestAddr, CodeSize);
return;
}
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} JIT_0x{:x}_{}\n", HostAddr, CodeSize, GuestAddr, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, GuestAddr, CodeSize);
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::Register(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) {
return;
}
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult =
fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, CodeSize, Name, Offset);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::RegisterNamedRegion(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) {
return;
}
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}\n", HostAddr, CodeSize, Name);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
RegisterNamedRegion(Buffer, HostAddr, CodeSize, Name);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::WriteBuffer(Core::JITSymbolBuffer* Buffer, bool ForceWrite) {
auto Now = std::chrono::steady_clock::now();
if (!ForceWrite) {
if (((Buffer->LastWrite - Now) < Buffer->MAXIMUM_THRESHOLD) && Buffer->Offset < Buffer->NEEDS_WRITE_DISTANCE) {
// Still buffering, no need to write.
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}+0x{:x} ({})\n", HostAddr, CodeSize, Name, Offset, HostAddr);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
Register(Buffer, HostAddr, CodeSize, Name, Offset);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
Buffer->LastWrite = Now;
auto Result = write(fd, Buffer->Buffer, Buffer->Offset);
if (Result == -1 && errno == EBADF) {
fd = -1;
}
void JITSymbols::RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name) {
if (fd == -1) return;
// Calculate remaining sizes.
const auto RemainingSize = Buffer->BUFFER_SIZE - Buffer->Offset;
const auto CurrentBufferOffset = &Buffer->Buffer[Buffer->Offset];
// Linux perf format is very straightforward
// `<HostPtr> <Size> <Name>\n`
const auto FMTResult = fmt::format_to_n(CurrentBufferOffset, RemainingSize, "{} {:x} {}\n", HostAddr, CodeSize, Name);
if (FMTResult.out >= &Buffer->Buffer[Buffer->BUFFER_SIZE]) {
// Couldn't fit, need to force a write.
WriteBuffer(Buffer, true);
// Rerun
RegisterNamedRegion(Buffer, HostAddr, CodeSize, Name);
return;
}
Buffer->Offset += FMTResult.size;
WriteBuffer(Buffer);
}
void JITSymbols::WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite) {
auto Now = std::chrono::steady_clock::now();
if (!ForceWrite) {
if (((Buffer->LastWrite - Now) < Buffer->MAXIMUM_THRESHOLD) &&
Buffer->Offset < Buffer->NEEDS_WRITE_DISTANCE) {
// Still buffering, no need to write.
return;
}
}
Buffer->LastWrite = Now;
auto Result = write(fd, Buffer->Buffer, Buffer->Offset);
if (Result == -1 && errno == EBADF) {
fd = -1;
}
Buffer->Offset = 0;
}
Buffer->Offset = 0;
}
} // namespace FEXCore
+8 -8
View File
@@ -17,20 +17,20 @@ public:
~JITSymbols();
void InitFile();
void RegisterNamedRegion(const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void *HostAddr, uint32_t CodeSize);
void RegisterNamedRegion(const void* HostAddr, uint32_t CodeSize, std::string_view Name);
void RegisterJITSpace(const void* HostAddr, uint32_t CodeSize);
// Allocate JIT buffer.
static fextl::unique_ptr<Core::JITSymbolBuffer> AllocateBuffer() {
return fextl::make_unique<Core::JITSymbolBuffer>();
}
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(Core::JITSymbolBuffer *Buffer, const void *HostAddr, uint32_t CodeSize, std::string_view Name);
void Register(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint64_t GuestAddr, uint32_t CodeSize);
void Register(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint32_t CodeSize, std::string_view Name, uintptr_t Offset);
void RegisterNamedRegion(Core::JITSymbolBuffer* Buffer, const void* HostAddr, uint32_t CodeSize, std::string_view Name);
private:
int fd{-1};
void WriteBuffer(Core::JITSymbolBuffer *Buffer, bool ForceWrite = false);
int fd {-1};
void WriteBuffer(Core::JITSymbolBuffer* Buffer, bool ForceWrite = false);
};
}
} // namespace FEXCore
+92 -122
View File
@@ -45,13 +45,13 @@ struct FEX_PACKED X80SoftFloat {
uint16_t Exponent : 15;
uint16_t Sign : 1;
X80SoftFloat() { memset(this, 0, sizeof(*this)); }
X80SoftFloat() {
memset(this, 0, sizeof(*this));
}
X80SoftFloat(uint16_t _Sign, uint16_t _Exponent, uint64_t _Significand)
: Significand {_Significand}
, Exponent {_Exponent}
, Sign {_Sign}
{
}
, Sign {_Sign} {}
fextl::string str() const {
fextl::ostringstream string;
@@ -63,21 +63,19 @@ struct FEX_PACKED X80SoftFloat {
}
// Ops
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FADD(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FADD(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
faddp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -85,21 +83,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSUB(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSUB(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fsubp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -107,21 +103,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FMUL(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FMUL(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fmulp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -129,21 +123,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FDIV(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FDIV(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
fdivp;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -151,11 +143,10 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
@@ -163,10 +154,9 @@ struct FEX_PACKED X80SoftFloat {
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -174,11 +164,10 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FREM1(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM1(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
@@ -186,10 +175,9 @@ struct FEX_PACKED X80SoftFloat {
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -197,30 +185,27 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FRNDINT(X80SoftFloat const &lhs, uint_fast8_t RoundMode) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_SIG(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FXTRACT_SIG(const X80SoftFloat& lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fxtract;
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st", "st(1)");
return Result;
#else
@@ -231,20 +216,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FXTRACT_EXP(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FXTRACT_EXP(const X80SoftFloat& lhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fxtract;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st", "st(1)");
return Result;
#else
@@ -253,19 +237,17 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static void FCMP(X80SoftFloat const &lhs, X80SoftFloat const &rhs, bool *eq, bool *lt, bool *nan) {
FEXCORE_PRESERVE_ALL_ATTR static void FCMP(const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(lhs, rhs);
*lt = extF80_lt(lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSCALE(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st1
fldt %[lhs]; # st0
@@ -273,10 +255,9 @@ struct FEX_PACKED X80SoftFloat {
fstpt %[result];
ffreep %%st(0);
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -289,20 +270,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat F2XM1(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat F2XM1(const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
f2xm1; # st0 = 2^st(0) - 1
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st");
return Result;
#else
@@ -313,22 +293,20 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FYL2X(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FYL2X(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[rhs]; # st(1)
fldt %[lhs]; # st(0)
fyl2x; # st(1) * log2l(st(0))
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -339,22 +317,20 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FATAN(X80SoftFloat const &lhs, X80SoftFloat const &rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FATAN(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs];
fldt %[rhs];
fpatan;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
, [rhs] "m" (rhs)
: "st", "st(1)");
: [result] "=m"(Result)
: [lhs] "m"(lhs), [rhs] "m"(rhs)
: "st", "st(1)");
return Result;
#else
@@ -365,21 +341,20 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FTAN(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FTAN(const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fptan;
ffreep %%st(0);
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st");
return Result;
#else
@@ -389,20 +364,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSIN(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSIN(const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fsin;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st");
return Result;
#else
@@ -412,20 +386,19 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FCOS(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FCOS(const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fcos;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st");
return Result;
#else
@@ -435,19 +408,18 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat FSQRT(X80SoftFloat const &lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSQRT(const X80SoftFloat& lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm (R"(
asm(R"(
fninit;
fldt %[lhs]; # st0
fsqrt;
fstpt %[result];
)"
: [result] "=m" (Result)
: [lhs] "m" (lhs)
: "st");
: [result] "=m"(Result)
: [lhs] "m"(lhs)
: "st");
return Result;
#else
@@ -471,7 +443,7 @@ struct FEX_PACKED X80SoftFloat {
const float128_t Result = extF80_to_f128(*this);
return FEXCore::BitCast<BIGFLOAT>(Result);
#else
BIGFLOAT result{};
BIGFLOAT result {};
memcpy(&result, this, sizeof(result));
return result;
#endif
@@ -570,19 +542,17 @@ struct FEX_PACKED X80SoftFloat {
}
operator extFloat80_t() const {
extFloat80_t Result{};
extFloat80_t Result {};
Result.signif = Significand;
Result.signExp = Exponent | (Sign << 15);
return Result;
}
static bool IsNan(X80SoftFloat const &lhs) {
return (lhs.Exponent == 0x7FFF) &&
(lhs.Significand & IntegerBit) &&
(lhs.Significand & Bottom62Significand);
static bool IsNan(const X80SoftFloat& lhs) {
return (lhs.Exponent == 0x7FFF) && (lhs.Significand & IntegerBit) && (lhs.Significand & Bottom62Significand);
}
static bool SignBit(X80SoftFloat const &lhs) {
static bool SignBit(const X80SoftFloat& lhs) {
return lhs.Sign;
}
+47 -40
View File
@@ -7,44 +7,51 @@
#include <optional>
namespace FEXCore::StrConv {
[[maybe_unused]] static bool Conv(std::string_view Value, bool *Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint8_t *Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint16_t *Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint32_t *Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, int32_t *Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, uint64_t *Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
template <typename T,
typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]] static bool Conv(std::string_view Value, T *Result) {
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]] static bool Conv(std::string_view Value, fextl::string *Result) {
*Result = Value;
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, bool* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint8_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint16_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint32_t* Result) {
*Result = std::strtoul(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, int32_t* Result) {
*Result = std::strtol(Value.data(), nullptr, 0);
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, uint64_t* Result) {
*Result = std::strtoull(Value.data(), nullptr, 0);
return true;
}
template<typename T, typename = std::enable_if<std::is_enum<T>::value, T>>
[[maybe_unused]]
static bool Conv(std::string_view Value, T* Result) {
*Result = static_cast<T>(std::stoull(Value.data(), nullptr, 0));
return true;
}
[[maybe_unused]]
static bool Conv(std::string_view Value, fextl::string* Result) {
*Result = Value;
return true;
}
} // namespace FEXCore::StrConv
+410 -418
View File
@@ -29,7 +29,7 @@
#include <utility>
namespace FEXCore::Context {
class Context;
class Context;
}
namespace FEXCore::Config {
@@ -40,472 +40,464 @@ namespace DefaultValues {
#define OPT_STRARRAY(group, enum, json, default) OPT_STR(group, enum, json, default)
#define OPT_STRENUM(group, enum, json, default) const uint64_t P(enum) = FEXCore::ToUnderlying(P(default));
#include <FEXCore/Config/ConfigValues.inl>
} // namespace DefaultValues
enum Paths {
PATH_DATA_DIR = 0,
PATH_CONFIG_DIR_LOCAL,
PATH_CONFIG_DIR_GLOBAL,
PATH_CONFIG_FILE_LOCAL,
PATH_CONFIG_FILE_GLOBAL,
PATH_CONFIG_TELEMETRY_FOLDER,
PATH_LAST,
};
static std::array<fextl::string, Paths::PATH_LAST> Paths;
void SetDataDirectory(const std::string_view Path) {
Paths[PATH_DATA_DIR] = Path;
}
enum Paths {
PATH_DATA_DIR = 0,
PATH_CONFIG_DIR_LOCAL,
PATH_CONFIG_DIR_GLOBAL,
PATH_CONFIG_FILE_LOCAL,
PATH_CONFIG_FILE_GLOBAL,
PATH_LAST,
};
static std::array<fextl::string, Paths::PATH_LAST> Paths;
void SetConfigDirectory(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_DIR_LOCAL + Global] = Path;
}
void SetDataDirectory(const std::string_view Path) {
Paths[PATH_DATA_DIR] = Path;
}
void SetConfigFileLocation(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_FILE_LOCAL + Global] = Path;
}
void SetConfigDirectory(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_DIR_LOCAL + Global] = Path;
}
void SetConfigFileLocation(const std::string_view Path, bool Global) {
Paths[PATH_CONFIG_FILE_LOCAL + Global] = Path;
}
fextl::string const& GetDataDirectory() {
return Paths[PATH_DATA_DIR];
}
fextl::string const& GetConfigDirectory(bool Global) {
return Paths[PATH_CONFIG_DIR_LOCAL + Global];
}
fextl::string const& GetConfigFileLocation(bool Global) {
return Paths[PATH_CONFIG_FILE_LOCAL + Global];
}
fextl::string GetApplicationConfig(const std::string_view Program, bool Global) {
fextl::string ConfigFile = GetConfigDirectory(Global);
if (!Global &&
!FHU::Filesystem::Exists(ConfigFile) &&
!FHU::Filesystem::CreateDirectories(ConfigFile)) {
LogMan::Msg::DFmt("Couldn't create config directory: '{}'", ConfigFile);
// Let's go local in this case
return fextl::fmt::format("./{}.json", Program);
const fextl::string& GetTelemetryDirectory() {
auto& Path = Paths[PATH_CONFIG_TELEMETRY_FOLDER];
if (Path.empty()) {
FEX_CONFIG_OPT(TelemetryDirectory, TELEMETRYDIRECTORY);
if (!TelemetryDirectory().empty()) {
Path = TelemetryDirectory;
Path += "/";
} else {
Path = Config::GetDataDirectory() + "Telemetry/";
}
}
ConfigFile += "AppConfig/";
return Path;
}
// Attempt to create the local folder if it doesn't exist
if (!Global &&
!FHU::Filesystem::Exists(ConfigFile) &&
!FHU::Filesystem::CreateDirectories(ConfigFile)) {
// Let's go local in this case
return fextl::fmt::format("./{}.json", Program);
const fextl::string& GetDataDirectory() {
return Paths[PATH_DATA_DIR];
}
const fextl::string& GetConfigDirectory(bool Global) {
return Paths[PATH_CONFIG_DIR_LOCAL + Global];
}
const fextl::string& GetConfigFileLocation(bool Global) {
return Paths[PATH_CONFIG_FILE_LOCAL + Global];
}
fextl::string GetApplicationConfig(const std::string_view Program, bool Global) {
fextl::string ConfigFile = GetConfigDirectory(Global);
if (!Global && !FHU::Filesystem::Exists(ConfigFile) && !FHU::Filesystem::CreateDirectories(ConfigFile)) {
LogMan::Msg::DFmt("Couldn't create config directory: '{}'", ConfigFile);
// Let's go local in this case
return fextl::fmt::format("./{}.json", Program);
}
ConfigFile += "AppConfig/";
// Attempt to create the local folder if it doesn't exist
if (!Global && !FHU::Filesystem::Exists(ConfigFile) && !FHU::Filesystem::CreateDirectories(ConfigFile)) {
// Let's go local in this case
return fextl::fmt::format("./{}.json", Program);
}
return fextl::fmt::format("{}{}.json", ConfigFile, Program);
}
void SetConfig(FEXCore::Context::Context* CTX, ConfigOption Option, uint64_t Config) {}
void SetConfig(FEXCore::Context::Context* CTX, ConfigOption Option, const fextl::string& Config) {}
uint64_t GetConfig(FEXCore::Context::Context* CTX, ConfigOption Option) {
return 0;
}
static fextl::map<FEXCore::Config::LayerType, fextl::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static FEXCore::Config::Layer* Meta {};
constexpr std::array<FEXCore::Config::LayerType, 10> LoadOrder = {
FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN, FEXCore::Config::LayerType::LAYER_MAIN,
FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP, FEXCore::Config::LayerType::LAYER_GLOBAL_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP, FEXCore::Config::LayerType::LAYER_LOCAL_APP,
FEXCore::Config::LayerType::LAYER_ARGUMENTS, FEXCore::Config::LayerType::LAYER_USER_OVERRIDE,
FEXCore::Config::LayerType::LAYER_ENVIRONMENT, FEXCore::Config::LayerType::LAYER_TOP};
Layer::Layer(const LayerType _Type)
: Type {_Type} {}
Layer::~Layer() {}
class MetaLayer final : public FEXCore::Config::Layer {
public:
MetaLayer(const LayerType _Type)
: FEXCore::Config::Layer(_Type) {}
~MetaLayer() {}
void Load();
private:
void MergeConfigMap(const LayerOptions& Options);
void MergeEnvironmentVariables(const ConfigOption& Option, const LayerValue& Value);
};
void MetaLayer::Load() {
OptionMap.clear();
for (auto CurrentLayer = LoadOrder.begin(); CurrentLayer != LoadOrder.end(); ++CurrentLayer) {
auto it = ConfigLayers.find(*CurrentLayer);
if (it != ConfigLayers.end() && *CurrentLayer != Type) {
// Merge this layer's options to this layer
MergeConfigMap(it->second->GetOptionMap());
}
}
}
return fextl::fmt::format("{}{}.json", ConfigFile, Program);
void MetaLayer::MergeEnvironmentVariables(const ConfigOption& Option, const LayerValue& Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
if (MetaEnvironment == OptionMap.end()) {
// Doesn't exist, just insert
OptionMap.insert_or_assign(Option, Value);
return;
}
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, uint64_t Config) {
}
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](const FEXCore::Config::LayerValue& Value) {
for (const auto& EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
// Broken environment variable
// Skip
continue;
}
auto Key = fextl::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = fextl::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
void SetConfig(FEXCore::Context::Context *CTX, ConfigOption Option, fextl::string const &Config) {
}
uint64_t GetConfig(FEXCore::Context::Context *CTX, ConfigOption Option) {
return 0;
}
static fextl::map<FEXCore::Config::LayerType, fextl::unique_ptr<FEXCore::Config::Layer>> ConfigLayers;
static FEXCore::Config::Layer *Meta{};
constexpr std::array<FEXCore::Config::LayerType, 9> LoadOrder = {
FEXCore::Config::LayerType::LAYER_GLOBAL_MAIN,
FEXCore::Config::LayerType::LAYER_MAIN,
FEXCore::Config::LayerType::LAYER_GLOBAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_GLOBAL_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_STEAM_APP,
FEXCore::Config::LayerType::LAYER_LOCAL_APP,
FEXCore::Config::LayerType::LAYER_ARGUMENTS,
FEXCore::Config::LayerType::LAYER_ENVIRONMENT,
FEXCore::Config::LayerType::LAYER_TOP
// Add the key to the map, overwriting whatever previous value was there
LookupMap.insert_or_assign(std::move(Key), std::move(Value));
}
};
Layer::Layer(const LayerType _Type)
: Type {_Type} {
AddToMap(MetaEnvironment->second);
AddToMap(Value);
// Now with the two layers merged in the map
// Add all the values to the option
Erase(Option);
for (auto& Val : LookupMap) {
// Set will emplace multiple options in to its list
Set(Option, Val.first + "=" + Val.second);
}
}
Layer::~Layer() {
}
class MetaLayer final : public FEXCore::Config::Layer {
public:
MetaLayer(const LayerType _Type)
: FEXCore::Config::Layer (_Type) {
}
~MetaLayer() {
}
void Load();
private:
void MergeConfigMap(const LayerOptions &Options);
void MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value);
};
void MetaLayer::Load() {
OptionMap.clear();
for (auto CurrentLayer = LoadOrder.begin(); CurrentLayer != LoadOrder.end(); ++CurrentLayer) {
auto it = ConfigLayers.find(*CurrentLayer);
if (it != ConfigLayers.end() && *CurrentLayer != Type) {
// Merge this layer's options to this layer
MergeConfigMap(it->second->GetOptionMap());
}
void MetaLayer::MergeConfigMap(const LayerOptions& Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto& it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV || it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
MergeEnvironmentVariables(it.first, it.second);
} else {
OptionMap.insert_or_assign(it.first, it.second);
}
}
}
void Initialize() {
AddLayer(fextl::make_unique<MetaLayer>(FEXCore::Config::LayerType::LAYER_TOP));
Meta = ConfigLayers.begin()->second.get();
}
void MetaLayer::MergeEnvironmentVariables(ConfigOption const &Option, LayerValue const &Value) {
// Environment variables need a bit of additional work
// We want to merge the arrays rather than overwrite entirely
auto MetaEnvironment = OptionMap.find(Option);
if (MetaEnvironment == OptionMap.end()) {
// Doesn't exist, just insert
OptionMap.insert_or_assign(Option, Value);
return;
}
void Shutdown() {
ConfigLayers.clear();
Meta = nullptr;
}
// If an environment variable exists in both current meta and in the incoming layer then the meta layer value is overwritten
fextl::unordered_map<fextl::string, fextl::string> LookupMap;
const auto AddToMap = [&LookupMap](FEXCore::Config::LayerValue const &Value) {
for (const auto &EnvVar : Value) {
const auto ItEq = EnvVar.find_first_of('=');
if (ItEq == fextl::string::npos) {
// Broken environment variable
// Skip
continue;
}
auto Key = fextl::string(EnvVar.begin(), EnvVar.begin() + ItEq);
auto Value = fextl::string(EnvVar.begin() + ItEq + 1, EnvVar.end());
// Add the key to the map, overwriting whatever previous value was there
LookupMap.insert_or_assign(std::move(Key), std::move(Value));
}
};
AddToMap(MetaEnvironment->second);
AddToMap(Value);
// Now with the two layers merged in the map
// Add all the values to the option
Erase(Option);
for (auto &Val : LookupMap) {
// Set will emplace multiple options in to its list
Set(Option, Val.first + "=" + Val.second);
void Load() {
for (auto CurrentLayer = LoadOrder.begin(); CurrentLayer != LoadOrder.end(); ++CurrentLayer) {
auto it = ConfigLayers.find(*CurrentLayer);
if (it != ConfigLayers.end()) {
it->second->Load();
}
}
}
void MetaLayer::MergeConfigMap(const LayerOptions &Options) {
// Insert this layer's options, overlaying previous options that exist here
for (auto &it : Options) {
if (it.first == FEXCore::Config::ConfigOption::CONFIG_ENV ||
it.first == FEXCore::Config::ConfigOption::CONFIG_HOSTENV) {
MergeEnvironmentVariables(it.first, it.second);
}
else {
OptionMap.insert_or_assign(it.first, it.second);
}
}
}
void Initialize() {
AddLayer(fextl::make_unique<MetaLayer>(FEXCore::Config::LayerType::LAYER_TOP));
Meta = ConfigLayers.begin()->second.get();
}
void Shutdown() {
ConfigLayers.clear();
Meta = nullptr;
}
void Load() {
for (auto CurrentLayer = LoadOrder.begin(); CurrentLayer != LoadOrder.end(); ++CurrentLayer) {
auto it = ConfigLayers.find(*CurrentLayer);
if (it != ConfigLayers.end()) {
it->second->Load();
}
}
}
fextl::string ExpandPath(fextl::string const &ContainerPrefix, fextl::string PathName) {
if (PathName.empty()) {
return {};
}
// Expand home if it exists
if (FHU::Filesystem::IsRelative(PathName)) {
fextl::string Home = getenv("HOME") ?: "";
// Home expansion only works if it is the first character
// This matches bash behaviour
if (PathName.at(0) == '~') {
PathName.replace(0, 1, Home);
return PathName;
}
// Expand relative path to absolute
char ExistsTempPath[PATH_MAX];
char *RealPath = FHU::Filesystem::Absolute(PathName.c_str(), ExistsTempPath);
if (RealPath) {
PathName = RealPath;
}
// Only return if it exists
if (FHU::Filesystem::Exists(PathName)) {
return PathName;
}
}
else {
// If the containerprefix and pathname isn't empty
// Then we check if the pathname exists in our current namespace
// If the path DOESN'T exist but DOES exist with the prefix applied
// then redirect to the prefix
//
// This might not be expected behaviour for some edge cases but since
// all paths aren't mounted inside the container, then it'll be fine
//
// Main catch case for this is the default thunk install folders
// HostThunks: $CMAKE_INSTALL_PREFIX/lib/fex-emu/HostThunks/
// GuestThunks: $CMAKE_INSTALL_PREFIX/share/fex-emu/GuestThunks/
if (!ContainerPrefix.empty() && !PathName.empty()) {
if (!FHU::Filesystem::Exists(PathName)) {
auto ContainerPath = ContainerPrefix + PathName;
if (FHU::Filesystem::Exists(ContainerPath)) {
return ContainerPath;
}
}
}
}
fextl::string ExpandPath(const fextl::string& ContainerPrefix, fextl::string PathName) {
if (PathName.empty()) {
return {};
}
constexpr char ContainerManager[] = "/run/host/container-manager";
fextl::string FindContainer() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
}
// Expand home if it exists
if (FHU::Filesystem::IsRelative(PathName)) {
fextl::string Home = getenv("HOME") ?: "";
// Home expansion only works if it is the first character
// This matches bash behaviour
if (PathName.at(0) == '~') {
PathName.replace(0, 1, Home);
return PathName;
}
return {};
}
fextl::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager{};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
return "/run/host/";
// Expand relative path to absolute
char ExistsTempPath[PATH_MAX];
char* RealPath = FHU::Filesystem::Absolute(PathName.c_str(), ExistsTempPath);
if (RealPath) {
PathName = RealPath;
}
// Only return if it exists
if (FHU::Filesystem::Exists(PathName)) {
return PathName;
}
} else {
// If the containerprefix and pathname isn't empty
// Then we check if the pathname exists in our current namespace
// If the path DOESN'T exist but DOES exist with the prefix applied
// then redirect to the prefix
//
// This might not be expected behaviour for some edge cases but since
// all paths aren't mounted inside the container, then it'll be fine
//
// Main catch case for this is the default thunk install folders
// HostThunks: $CMAKE_INSTALL_PREFIX/lib/fex-emu/HostThunks/
// GuestThunks: $CMAKE_INSTALL_PREFIX/share/fex-emu/GuestThunks/
if (!ContainerPrefix.empty() && !PathName.empty()) {
if (!FHU::Filesystem::Exists(PathName)) {
auto ContainerPath = ContainerPrefix + PathName;
if (FHU::Filesystem::Exists(ContainerPath)) {
return ContainerPath;
}
}
}
return {};
}
return {};
}
void ReloadMetaLayer() {
Meta->Load();
constexpr char ContainerManager[] = "/run/host/container-manager";
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CORE)) {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
fextl::string FindContainer() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager {};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
return ManagerStr;
}
}
return {};
}
fextl::string FindContainerPrefix() {
// We only support pressure-vessel at the moment
if (FHU::Filesystem::Exists(ContainerManager)) {
fextl::vector<char> Manager {};
if (FEXCore::FileLoading::LoadFile(Manager, ContainerManager)) {
// Trim the whitespace, may contain a newline
fextl::string ManagerStr = Manager.data();
ManagerStr = FEXCore::StringUtils::Trim(ManagerStr);
if (strncmp(ManagerStr.data(), "pressure-vessel", Manager.size()) == 0) {
// We are running inside of pressure vessel
// Our $CMAKE_INSTALL_PREFIX paths are now inside of /run/host/$CMAKE_INSTALL_PREFIX
return "/run/host/";
}
}
}
return {};
}
void ReloadMetaLayer() {
Meta->Load();
// Do configuration option fix ups after everything is reloaded
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CORE)) {
// Sanitize Core option
FEX_CONFIG_OPT(Core, CORE);
#if (_M_X86_64)
constexpr uint32_t MaxCoreNumber = 1;
constexpr uint32_t MaxCoreNumber = 1;
#else
constexpr uint32_t MaxCoreNumber = 0;
constexpr uint32_t MaxCoreNumber = 0;
#endif
if (Core > MaxCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, fextl::fmt::format("{}", static_cast<uint32_t>(FEXCore::Config::CONFIG_IRJIT)));
if (Core > MaxCoreNumber) {
// Sanitize the core option by setting the core to the JIT if invalid
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_CORE, fextl::fmt::format("{}", static_cast<uint32_t>(FEXCore::Config::CONFIG_IRJIT)));
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
}
fextl::string ContainerPrefix {FindContainerPrefix()};
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, fextl::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
}
};
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
} else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
if (FHU::Filesystem::Exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_CACHEOBJECTCODECOMPILATION)) {
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(Core, CORE);
}
fextl::string ContainerPrefix { FindContainerPrefix() };
auto ExpandPathIfExists = [&ContainerPrefix](FEXCore::Config::ConfigOption Config, fextl::string PathName) {
auto NewPath = ExpandPath(ContainerPrefix, PathName);
if (!NewPath.empty()) {
FEXCore::Config::EraseSet(Config, NewPath);
}
};
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_ROOTFS)) {
FEX_CONFIG_OPT(PathName, ROOTFS);
auto ExpandedString = ExpandPath(ContainerPrefix,PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, ExpandedString);
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedRootFS = GetDataDirectory() + "RootFS/" + PathName();
if (FHU::Filesystem::Exists(NamedRootFS)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_ROOTFS, NamedRootFS);
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKHOSTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKHOSTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKHOSTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKGUESTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKGUESTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKGUESTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKCONFIG)) {
FEX_CONFIG_OPT(PathName, THUNKCONFIG);
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, ExpandedString);
} else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedConfig = GetDataDirectory() + "ThunkConfigs/" + PathName();
if (FHU::Filesystem::Exists(NamedConfig)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, NamedConfig);
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKHOSTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKHOSTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKHOSTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKGUESTLIBS)) {
FEX_CONFIG_OPT(PathName, THUNKGUESTLIBS);
ExpandPathIfExists(FEXCore::Config::CONFIG_THUNKGUESTLIBS, PathName());
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_THUNKCONFIG)) {
FEX_CONFIG_OPT(PathName, THUNKCONFIG);
auto ExpandedString = ExpandPath(ContainerPrefix, PathName());
if (!ExpandedString.empty()) {
// Adjust the path if it ended up being relative
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, ExpandedString);
}
else if (!PathName().empty()) {
// If the filesystem doesn't exist then let's see if it exists in the fex-emu folder
fextl::string NamedConfig = GetDataDirectory() + "ThunkConfigs/" + PathName();
if (FHU::Filesystem::Exists(NamedConfig)) {
FEXCore::Config::EraseSet(FEXCore::Config::CONFIG_THUNKCONFIG, NamedConfig);
}
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_OUTPUTLOG)) {
FEX_CONFIG_OPT(PathName, OUTPUTLOG);
if (PathName() != "stdout" && PathName() != "stderr" && PathName() != "server") {
ExpandPathIfExists(FEXCore::Config::CONFIG_OUTPUTLOG, PathName());
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_DUMPIR) &&
!FEXCore::Config::Exists(FEXCore::Config::CONFIG_PASSMANAGERDUMPIR)) {
// If DumpIR is set but no PassManagerDumpIR configuration is set, then default to `afteropt`
FEX_CONFIG_OPT(PathName, DUMPIR);
if (PathName() != "no") {
EraseSet(FEXCore::Config::ConfigOption::CONFIG_PASSMANAGERDUMPIR, fextl::fmt::format("{}", static_cast<uint64_t>(FEXCore::Config::PassManagerDumpIR::AFTEROPT)));
}
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_SINGLESTEP)) {
// Single stepping also enforces single instruction size blocks
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, "1");
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_OUTPUTLOG)) {
FEX_CONFIG_OPT(PathName, OUTPUTLOG);
if (PathName() != "stdout" && PathName() != "stderr" && PathName() != "server") {
ExpandPathIfExists(FEXCore::Config::CONFIG_OUTPUTLOG, PathName());
}
}
void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer) {
ConfigLayers.emplace(_Layer->GetLayerType(), std::move(_Layer));
}
bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<LayerValue*> All(ConfigOption Option) {
return Meta->All(Option);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
return Meta->Get(Option);
}
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
void Erase(ConfigOption Option) {
Meta->Erase(Option);
}
void EraseSet(ConfigOption Option, std::string_view Data) {
Meta->EraseSet(Option, Data);
}
template<typename T>
T Value<T>::Get(FEXCore::Config::ConfigOption Option) {
T Result;
auto Value = FEXCore::Config::Get(Option);
if (!FEXCore::StrConv::Conv(**Value, &Result)) {
LOGMAN_MSG_A_FMT("Attempted to convert invalid value");
}
return Result;
}
template<typename T>
T Value<T>::GetIfExists(FEXCore::Config::ConfigOption Option, T Default) {
T Result;
auto Value = FEXCore::Config::Get(Option);
if (Value && FEXCore::StrConv::Conv(**Value, &Result)) {
return Result;
}
else {
return Default;
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_DUMPIR) && !FEXCore::Config::Exists(FEXCore::Config::CONFIG_PASSMANAGERDUMPIR)) {
// If DumpIR is set but no PassManagerDumpIR configuration is set, then default to `afteropt`
FEX_CONFIG_OPT(PathName, DUMPIR);
if (PathName() != "no") {
EraseSet(FEXCore::Config::ConfigOption::CONFIG_PASSMANAGERDUMPIR,
fextl::fmt::format("{}", static_cast<uint64_t>(FEXCore::Config::PassManagerDumpIR::AFTEROPT)));
}
}
template<>
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, fextl::string Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
}
else {
return Default;
}
if (FEXCore::Config::Exists(FEXCore::Config::CONFIG_SINGLESTEP)) {
// Single stepping also enforces single instruction size blocks
Set(FEXCore::Config::ConfigOption::CONFIG_MAXINST, "1");
}
template<>
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
}
else {
return fextl::string(Default);
}
}
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template int8_t Value<int8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int8_t Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template int16_t Value<int16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int16_t Default);
template uint16_t Value<uint16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint16_t Default);
template int32_t Value<int32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int32_t Default);
template uint32_t Value<uint32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint32_t Default);
template int64_t Value<int64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int64_t Default);
template uint64_t Value<uint64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint64_t Default);
// Constructor
template Value<fextl::string>::Value(FEXCore::Config::ConfigOption _Option, fextl::string Default);
template Value<bool>::Value(FEXCore::Config::ConfigOption _Option, bool Default);
template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t Default);
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string> *List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<fextl::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string> *List);
}
void AddLayer(fextl::unique_ptr<FEXCore::Config::Layer> _Layer) {
ConfigLayers.emplace(_Layer->GetLayerType(), std::move(_Layer));
}
bool Exists(ConfigOption Option) {
return Meta->OptionExists(Option);
}
std::optional<LayerValue*> All(ConfigOption Option) {
return Meta->All(Option);
}
std::optional<fextl::string*> Get(ConfigOption Option) {
return Meta->Get(Option);
}
void Set(ConfigOption Option, std::string_view Data) {
Meta->Set(Option, Data);
}
void Erase(ConfigOption Option) {
Meta->Erase(Option);
}
void EraseSet(ConfigOption Option, std::string_view Data) {
Meta->EraseSet(Option, Data);
}
template<typename T>
T Value<T>::Get(FEXCore::Config::ConfigOption Option) {
T Result;
auto Value = FEXCore::Config::Get(Option);
if (!FEXCore::StrConv::Conv(**Value, &Result)) {
LOGMAN_MSG_A_FMT("Attempted to convert invalid value");
}
return Result;
}
template<typename T>
T Value<T>::GetIfExists(FEXCore::Config::ConfigOption Option, T Default) {
T Result;
auto Value = FEXCore::Config::Get(Option);
if (Value && FEXCore::StrConv::Conv(**Value, &Result)) {
return Result;
} else {
return Default;
}
}
template<>
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, fextl::string Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
} else {
return Default;
}
}
template<>
fextl::string Value<fextl::string>::GetIfExists(FEXCore::Config::ConfigOption Option, std::string_view Default) {
auto Value = FEXCore::Config::Get(Option);
if (Value) {
return **Value;
} else {
return fextl::string(Default);
}
}
template bool Value<bool>::GetIfExists(FEXCore::Config::ConfigOption Option, bool Default);
template int8_t Value<int8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int8_t Default);
template uint8_t Value<uint8_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint8_t Default);
template int16_t Value<int16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int16_t Default);
template uint16_t Value<uint16_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint16_t Default);
template int32_t Value<int32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int32_t Default);
template uint32_t Value<uint32_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint32_t Default);
template int64_t Value<int64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, int64_t Default);
template uint64_t Value<uint64_t>::GetIfExists(FEXCore::Config::ConfigOption Option, uint64_t Default);
// Constructor
template Value<fextl::string>::Value(FEXCore::Config::ConfigOption _Option, fextl::string Default);
template Value<bool>::Value(FEXCore::Config::ConfigOption _Option, bool Default);
template Value<uint8_t>::Value(FEXCore::Config::ConfigOption _Option, uint8_t Default);
template Value<uint64_t>::Value(FEXCore::Config::ConfigOption _Option, uint64_t Default);
template<typename T>
void Value<T>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string>* List) {
auto Value = FEXCore::Config::All(Option);
List->clear();
if (Value) {
*List = **Value;
}
}
template void Value<fextl::string>::GetListIfExists(FEXCore::Config::ConfigOption Option, fextl::list<fextl::string>* List);
} // namespace FEXCore::Config
+41 -3
View File
@@ -368,6 +368,14 @@
"File to write FEX output to.",
"[stdout, stderr, server, <Filename>]"
]
},
"TelemetryDirectory": {
"Type": "str",
"Default": "",
"Desc": [
"Redirects the telemetry folder that FEX usually writes to.",
"By default telemetry data is stored in {$FEX_APP_DATA_LOCATION,{$XDG_DATA_HOME,$HOME}/.fex-emu/Telemetry/}"
]
}
},
"Hacks": {
@@ -379,9 +387,8 @@
"Desc": [
"Checks code for modification before execution.",
"\tnone: No checks",
"\tmtrack: Page tracking based invalidation",
"\tfull: Validate code before every run (slow)",
"\tmman: Invalidate on mmap, mprotect, munmap (deprecated, use mtrack)"
"\tmtrack: Page tracking based invalidation (default)",
"\tfull: Validate code before every run (slow)"
]
},
"TSOEnabled": {
@@ -392,6 +399,29 @@
"Highly likely to break any multithreaded application if disabled."
]
},
"VectorTSOEnabled": {
"Type": "bool",
"Default": "false",
"Desc": [
"When TSO emulation is enabled, controls if vector loadstores should also be atomic."
]
},
"MemcpySetTSOEnabled": {
"Type": "bool",
"Default": "false",
"Desc": [
"When TSO emulation is enabled, controls if memcpy and memset should also be atomic.",
"Only affects REP MOVS and REP STOS instructions"
]
},
"HalfBarrierTSOEnabled": {
"Type": "bool",
"Default": "true",
"Desc": [
"When TSO emulation is enabled, controls if unaligned loads and stores should be backpatched to half-barrier atomics.",
"Can be dangerous due to aligned loadstores through the same code now become non-atomic."
]
},
"TSOAutoMigration": {
"Type": "bool",
"Default": "true",
@@ -439,6 +469,14 @@
"Hides the hypervisor CPUID bit when set.",
"Should only be used for applications that have issues with this set."
]
},
"StartupSleep": {
"Type": "uint32",
"Default": "0",
"Desc": [
"Sleeps the process at startup for a duration of seconds.",
"Useful if an application crashes too quickly to attach a debugger."
]
}
},
"Misc": {
+57 -53
View File
@@ -13,57 +13,61 @@
#include <utility>
namespace FEXCore::Context {
void InitializeStaticTables(OperatingMode Mode) {
X86Tables::InitializeInfoTables(Mode);
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext() {
return fextl::make_unique<FEXCore::Context::ContextImpl>();
}
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
CustomExitHandler = std::move(handler);
}
ExitHandler FEXCore::Context::ContextImpl::GetExitHandler() const {
return CustomExitHandler;
}
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
void FEXCore::Context::ContextImpl::CompileRIPCount(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst) {
CompileBlock(Thread->CurrentFrame, GuestRIP, MaxInst);
}
void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) {
CustomCPUFactory = std::move(Factory);
}
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator *_SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
void FEXCore::Context::ContextImpl::SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) {
SyscallHandler = Handler;
SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunction(uint32_t Function, uint32_t Leaf) {
return CPUID.RunFunction(Function, Leaf);
}
FEXCore::CPUID::XCRResults FEXCore::Context::ContextImpl::RunXCRFunction(uint32_t Function) {
return CPUID.RunXCRFunction(Function);
}
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CPUID.RunFunctionName(Function, Leaf, CPU);
}
void InitializeStaticTables(OperatingMode Mode) {
X86Tables::InitializeInfoTables(Mode);
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext() {
return fextl::make_unique<FEXCore::Context::ContextImpl>();
}
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
CustomExitHandler = std::move(handler);
}
ExitHandler FEXCore::Context::ContextImpl::GetExitHandler() const {
return CustomExitHandler;
}
void FEXCore::Context::ContextImpl::CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
CompileBlock(Thread->CurrentFrame, GuestRIP);
}
void FEXCore::Context::ContextImpl::CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
CompileBlock(Thread->CurrentFrame, GuestRIP, MaxInst);
}
void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) {
CustomCPUFactory = std::move(Factory);
}
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator* _SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
void FEXCore::Context::ContextImpl::SetSyscallHandler(FEXCore::HLE::SyscallHandler* Handler) {
SyscallHandler = Handler;
SourcecodeResolver = Handler->GetSourcecodeResolver();
}
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunction(uint32_t Function, uint32_t Leaf) {
return CPUID.RunFunction(Function, Leaf);
}
FEXCore::CPUID::XCRResults FEXCore::Context::ContextImpl::RunXCRFunction(uint32_t Function) {
return CPUID.RunXCRFunction(Function);
}
FEXCore::CPUID::FunctionResults FEXCore::Context::ContextImpl::RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) {
return CPUID.RunFunctionName(Function, Leaf, CPU);
}
bool FEXCore::Context::ContextImpl::IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const {
return Thread->CPUBackend->IsAddressInCodeBuffer(Address);
}
} // namespace FEXCore::Context
+311 -297
View File
@@ -13,6 +13,7 @@
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/SignalScopeGuards.h>
@@ -45,359 +46,372 @@ namespace CodeSerialize {
namespace CPU {
class Arm64JITCore;
class Dispatcher;
}
} // namespace CPU
namespace HLE {
struct SyscallArguments;
class SyscallHandler;
class SourcecodeResolver;
struct SourcecodeMap;
}
}
struct SyscallArguments;
class SyscallHandler;
class SourcecodeResolver;
struct SourcecodeMap;
} // namespace HLE
} // namespace FEXCore
namespace FEXCore::IR {
class RegisterAllocationData;
class IRListView;
class RegisterAllocationData;
struct IRListCopy;
class IRListView;
namespace Validation {
class IRValidation;
}
}
} // namespace FEXCore::IR
namespace FEXCore::Context {
enum CoreRunningMode {
MODE_RUN = 0,
MODE_SINGLESTEP = 1,
};
enum CoreRunningMode {
MODE_RUN = 0,
MODE_SINGLESTEP = 1,
};
struct ExitFunctionLinkData {
uint64_t HostBranch;
uint64_t GuestRIP;
};
struct ExitFunctionLinkData {
uint64_t HostBranch;
uint64_t GuestRIP;
};
using BlockDelinkerFunc = void(*)(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record);
constexpr uint32_t TSC_SCALE = 128;
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
using BlockDelinkerFunc = void (*)(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record);
constexpr uint32_t TSC_SCALE = 128;
constexpr uint32_t TSC_SCALE_MAXIMUM = 1'000'000'000; ///< 1Ghz
class ContextImpl final : public FEXCore::Context::Context {
public:
// Context base class implementation.
bool InitCore() override;
class ContextImpl final : public FEXCore::Context::Context {
public:
// Context base class implementation.
bool InitCore() override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
void SetExitHandler(ExitHandler handler) override;
ExitHandler GetExitHandler() const override;
ExitReason RunUntilExit(FEXCore::Core::InternalThreadState *Thread) override;
ExitReason RunUntilExit(FEXCore::Core::InternalThreadState* Thread) override;
void ExecuteThread(FEXCore::Core::InternalThreadState *Thread) override;
void ExecuteThread(FEXCore::Core::InternalThreadState* Thread) override;
void CompileRIP(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
void CompileRIP(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) override;
void CompileRIPCount(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) override;
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
HostFeatures GetHostFeatures() const override;
HostFeatures GetHostFeatures() const override;
void HandleCallback(FEXCore::Core::InternalThreadState *Thread, uint64_t RIP) override;
void HandleCallback(FEXCore::Core::InternalThreadState* Thread, uint64_t RIP) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState *Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState *Thread, bool WasInJIT, uint64_t *HostGPRs, uint64_t PSTATE) override;
void SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState *Thread, uint32_t EFLAGS) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, uint64_t* HostGPRs, uint64_t PSTATE) override;
void SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) override;
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param InitialRIP The starting RIP of this thread
* @param StackPointer The starting RSP of this thread
* @param NewThreadState The initial thread state to setup for our state, if inheriting.
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* Parent thread Creation:
* - Thread = CreateThread(InitialRIP, InitialStack, nullptr, 0);
* - CTX->RunUntilExit(Thread);
* OS thread Creation:
* - Thread = CreateThread(0, 0, NewState, PPID);
* - Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
* - ThreadHandler calls `CTX->ExecutionThread(Thread)`
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(0, 0, CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(0, 0, NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
* @param InitialRIP The starting RIP of this thread
* @param StackPointer The starting RSP of this thread
* @param NewThreadState The initial thread state to setup for our state, if inheriting.
* @param ParentTID The PID that was the parent thread that created this
*
* @return The InternalThreadState object that tracks all of the emulated thread's state
*
* Usecases:
* Parent thread Creation:
* - Thread = CreateThread(InitialRIP, InitialStack, nullptr, 0);
* - CTX->RunUntilExit(Thread);
* OS thread Creation:
* - Thread = CreateThread(0, 0, NewState, PPID);
* - Thread->ExecutionThread = FEXCore::Threads::Thread::Create(ThreadHandler, Arg);
* - ThreadHandler calls `CTX->ExecutionThread(Thread)`
* OS fork (New thread created with a clone of thread state):
* - clone{2, 3}
* - Thread = CreateThread(0, 0, CopyOfThreadState, PPID);
* - ExecutionThread(Thread); // Starts executing without creating another host thread
* Thunk callback executing guest code from native host thread
* - Thread = CreateThread(0, 0, NewState, PPID);
* - InitializeThreadTLSData(Thread);
* - HandleCallback(Thread, RIP);
*/
FEXCore::Core::InternalThreadState* CreateThread(uint64_t InitialRIP, uint64_t StackPointer, FEXCore::Core::CPUState *NewThreadState, uint64_t ParentTID) override;
FEXCore::Core::InternalThreadState*
CreateThread(uint64_t InitialRIP, uint64_t StackPointer, FEXCore::Core::CPUState* NewThreadState, uint64_t ParentTID) override;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState *Thread) override;
// Public for threading
void ExecutionThread(FEXCore::Core::InternalThreadState* Thread) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState *Thread, bool NeedsTLSUninstall) override;
/**
* @brief Destroys this FEX thread object and stops tracking it internally
*
* @param Thread The internal FEX thread state object
*/
void DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool NeedsTLSUninstall) override;
#ifndef _WIN32
void LockBeforeFork(FEXCore::Core::InternalThreadState *Thread) override;
void UnlockAfterFork(FEXCore::Core::InternalThreadState *Thread, bool Child) override;
void LockBeforeFork(FEXCore::Core::InternalThreadState* Thread) override;
void UnlockAfterFork(FEXCore::Core::InternalThreadState* Thread, bool Child) override;
#endif
void SetSignalDelegator(FEXCore::SignalDelegator *SignalDelegation) override;
void SetSyscallHandler(FEXCore::HLE::SyscallHandler *Handler) override;
void SetSignalDelegator(FEXCore::SignalDelegator* SignalDelegation) override;
void SetSyscallHandler(FEXCore::HLE::SyscallHandler* Handler) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) override;
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunction(uint32_t Function, uint32_t Leaf) override;
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) override;
FEXCore::CPUID::FunctionResults RunCPUIDFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) override;
FEXCore::IR::AOTIRCacheEntry *LoadAOTIRCacheEntry(const fextl::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry *Entry) override;
FEXCore::IR::AOTIRCacheEntry* LoadAOTIRCacheEntry(const fextl::string& Name) override;
void UnloadAOTIRCacheEntry(FEXCore::IR::AOTIRCacheEntry* Entry) override;
void SetAOTIRLoader(AOTIRLoaderCBFn CacheReader) override {
IRCaptureCache.SetAOTIRLoader(std::move(CacheReader));
}
void SetAOTIRWriter(AOTIRWriterCBFn CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(std::move(CacheWriter));
}
void SetAOTIRRenamer(AOTIRRenamerCBFn CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(std::move(CacheRenamer));
}
void SetAOTIRLoader(AOTIRLoaderCBFn CacheReader) override {
IRCaptureCache.SetAOTIRLoader(std::move(CacheReader));
}
void SetAOTIRWriter(AOTIRWriterCBFn CacheWriter) override {
IRCaptureCache.SetAOTIRWriter(std::move(CacheWriter));
}
void SetAOTIRRenamer(AOTIRRenamerCBFn CacheRenamer) override {
IRCaptureCache.SetAOTIRRenamer(std::move(CacheRenamer));
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(AOTIRCodeFileWriterFn Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void FinalizeAOTIRCache() override {
IRCaptureCache.FinalizeAOTIRCache();
}
void WriteFilesWithCode(AOTIRCodeFileWriterFn Writer) override {
IRCaptureCache.WriteFilesWithCode(Writer);
}
void ClearCodeCache(FEXCore::Core::InternalThreadState *Thread) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState *Thread, uint64_t Start, uint64_t Length, CodeRangeInvalidationFn callback) override;
FEXCore::ForkableSharedMutex &GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
void ClearCodeCache(FEXCore::Core::InternalThreadState* Thread) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length) override;
void InvalidateGuestCodeRange(FEXCore::Core::InternalThreadState* Thread, uint64_t Start, uint64_t Length, CodeRangeInvalidationFn callback) override;
FEXCore::ForkableSharedMutex& GetCodeInvalidationMutex() override {
return CodeInvalidationMutex;
}
void MarkMemoryShared(FEXCore::Core::InternalThreadState *Thread) override;
void MarkMemoryShared(FEXCore::Core::InternalThreadState* Thread) override;
void ConfigureAOTGen(FEXCore::Core::InternalThreadState *Thread, fextl::set<uint64_t> *ExternalBranches, uint64_t SectionMaxAddress) override;
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void *Creator = nullptr, void *Data = nullptr);
void ConfigureAOTGen(FEXCore::Core::InternalThreadState* Thread, fextl::set<uint64_t>* ExternalBranches, uint64_t SectionMaxAddress) override;
void AppendThunkDefinitions(fextl::vector<FEXCore::IR::ThunkDefinition> const& Definitions) override;
bool IsAddressInCodeBuffer(FEXCore::Core::InternalThreadState* Thread, uintptr_t Address) const override;
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void* Creator = nullptr, void* Data = nullptr);
friend class FEXCore::IR::Validation::IRValidation;
void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) override;
struct {
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
uint64_t VirtualMemSize{1ULL << 36};
public:
friend class FEXCore::HLE::SyscallHandler;
#ifdef JIT_ARM64
friend class FEXCore::CPU::Arm64JITCore;
#endif
// this is for internal use
bool ValidateIRarser { false };
friend class FEXCore::IR::Validation::IRValidation;
// Used if the JIT needs to have its interrupt fault code emitted.
bool NeedsPendingInterruptFaultCheck { false };
struct {
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
uint64_t VirtualMemSize {1ULL << 36};
FEX_CONFIG_OPT(Multiblock, MULTIBLOCK);
FEX_CONFIG_OPT(SingleStepConfig, SINGLESTEP);
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(TSOAutoMigration, TSOAUTOMIGRATION);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRGenerate, AOTIRGENERATE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(Core, CORE);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(ThunkHostLibsPath32, THUNKHOSTLIBS32);
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
} Config;
// Used if the JIT needs to have its interrupt fault code emitted.
bool NeedsPendingInterruptFaultCheck {false};
FEX_CONFIG_OPT(Multiblock, MULTIBLOCK);
FEX_CONFIG_OPT(SingleStepConfig, SINGLESTEP);
FEX_CONFIG_OPT(GdbServer, GDBSERVER);
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
FEX_CONFIG_OPT(TSOEnabled, TSOENABLED);
FEX_CONFIG_OPT(TSOAutoMigration, TSOAUTOMIGRATION);
FEX_CONFIG_OPT(ABILocalFlags, ABILOCALFLAGS);
FEX_CONFIG_OPT(AOTIRCapture, AOTIRCAPTURE);
FEX_CONFIG_OPT(AOTIRGenerate, AOTIRGENERATE);
FEX_CONFIG_OPT(AOTIRLoad, AOTIRLOAD);
FEX_CONFIG_OPT(SMCChecks, SMCCHECKS);
FEX_CONFIG_OPT(Core, CORE);
FEX_CONFIG_OPT(MaxInstPerBlock, MAXINST);
FEX_CONFIG_OPT(RootFSPath, ROOTFS);
FEX_CONFIG_OPT(ThunkHostLibsPath, THUNKHOSTLIBS);
FEX_CONFIG_OPT(ThunkHostLibsPath32, THUNKHOSTLIBS32);
FEX_CONFIG_OPT(ThunkConfigFile, THUNKCONFIG);
FEX_CONFIG_OPT(GlobalJITNaming, GLOBALJITNAMING);
FEX_CONFIG_OPT(LibraryJITNaming, LIBRARYJITNAMING);
FEX_CONFIG_OPT(BlockJITNaming, BLOCKJITNAMING);
FEX_CONFIG_OPT(GDBSymbols, GDBSYMBOLS);
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(CacheObjectCodeCompilation, CACHEOBJECTCODECOMPILATION);
FEX_CONFIG_OPT(x87ReducedPrecision, X87REDUCEDPRECISION);
FEX_CONFIG_OPT(DisableTelemetry, DISABLETELEMETRY);
FEX_CONFIG_OPT(DisableVixlIndirectCalls, DISABLE_VIXL_INDIRECT_RUNTIME_CALLS);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
} Config;
std::atomic_bool CoreShuttingDown{false};
std::atomic_bool CoreShuttingDown {false};
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::ForkableSharedMutex CodeInvalidationMutex;
FEXCore::HostFeatures HostFeatures;
// CPUID depends on HostFeatures so needs to be initialized after that.
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler *SyscallHandler{};
FEXCore::HLE::SourcecodeResolver *SourcecodeResolver{};
fextl::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
FEXCore::HostFeatures HostFeatures;
// CPUID depends on HostFeatures so needs to be initialized after that.
FEXCore::CPUIDEmu CPUID;
FEXCore::HLE::SyscallHandler* SyscallHandler {};
FEXCore::HLE::SourcecodeResolver* SourcecodeResolver {};
fextl::unique_ptr<FEXCore::ThunkHandler> ThunkHandler;
fextl::unique_ptr<FEXCore::CPU::Dispatcher> Dispatcher;
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
CustomCPUFactoryType CustomCPUFactory;
FEXCore::Context::ExitHandler CustomExitHandler;
#ifdef BLOCKSTATS
fextl::unique_ptr<FEXCore::BlockSamplingData> BlockData;
fextl::unique_ptr<FEXCore::BlockSamplingData> BlockData;
#endif
SignalDelegator *SignalDelegation{};
X86GeneratedCode X86CodeGen;
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl();
~ContextImpl();
ContextImpl();
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData *HostLink, const BlockDelinkerFunc &delinker);
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
static void ThreadAddBlockLink(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestDestination,
FEXCore::Context::ExitFunctionLinkData* HostLink, const BlockDelinkerFunc& delinker);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, ExitFunctionLinkData *Record) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
template<auto Fn>
static uint64_t ThreadExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
auto lk = GuardSignalDeferringSection<std::shared_lock>(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
return Fn(Frame, Record);
}
return Fn(Frame, Record);
}
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
// Wrapper which takes CpuStateFrame instead of InternalThreadState and unique_locks CodeInvalidationMutex
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LogMan::Throw::AFmt(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}", Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
LOGMAN_THROW_A_FMT(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}",
Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
ThreadRemoveCodeEntry(Thread, GuestRIP);
}
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst);
struct CompileCodeResult {
void* CompiledCode;
FEXCore::IR::IRListView* IRData;
FEXCore::Core::DebugData* DebugData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
bool GeneratedIR;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]] CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState *Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame *Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
// Used for thread creation from syscalls
/**
* @brief Initializes TID, PID and TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState *Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState *ParentThread, FEXCore::Core::InternalThreadState *ChildThread);
uint8_t GetGPRSize() const { return Config.Is64BitMode ? 8 : 4; }
FEXCore::JITSymbols Symbols;
void GetVDSOSigReturn(VDSOSigReturn *VDSOPointers) override {
if (VDSOPointers->VDSO_kernel_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_sigreturn = reinterpret_cast<void*>(X86CodeGen.sigreturn_32);
}
if (VDSOPointers->VDSO_kernel_rt_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_rt_sigreturn = reinterpret_cast<void*>(X86CodeGen.rt_sigreturn_32);
}
}
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const { return AtomicTSOEmulationEnabled; }
void SetHardwareTSOSupport(bool HardwareTSOSupported) override {
SupportsHardwareTSO = HardwareTSOSupported;
UpdateAtomicTSOEmulationConfig();
}
// Returns if Software TSO emulation is required.
// NOTE: This doesn't necessary return if Atomic-based TSO is currently enabled.
// This will still return true if on a single thread and TSO is currently disabled.
//
// This is to ensure that if early initialization checks CPU features and TSO /could/ be enabled, that
// we return consistent results.
//
// To check if Atomic TSO is currently enabled in the JIT, use `IsAtomicTSOEnabled` instead.
bool SoftwareTSORequired() const {
if (SupportsHardwareTSO) return false;
return Config.TSOEnabled;
}
void EnableExitOnHLT() override { ExitOnHLT = true; }
bool ExitOnHLTEnabled() const { return ExitOnHLT; }
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
// If the hardware supports TSO then we don't need to emulate it through atomics.
AtomicTSOEmulationEnabled = false;
}
else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
}
}
private:
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void AddBlockMapping(FEXCore::Core::InternalThreadState *Thread, uint64_t Address, void *Ptr);
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
bool ExitOnHLT = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::atomic<bool> HasCustomIRHandlers{};
fextl::unordered_map<uint64_t, std::tuple<CustomIREntrypointHandler, void *, void *>> CustomIRHandlers;
struct GenerateIRResult {
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
uint64_t Length;
};
}
[[nodiscard]]
GenerateIRResult GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst);
struct CompileCodeResult {
void* CompiledCode;
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData;
bool GeneratedIR;
uint64_t StartAddr;
uint64_t Length;
};
[[nodiscard]]
CompileCodeResult CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst = 0);
uintptr_t CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP, uint64_t MaxInst = 0);
// Used for thread creation from syscalls
/**
* @brief Initializes TID, PID and TLS data for a thread
*
* @param Thread The internal FEX thread state object
*/
void InitializeThreadTLSData(FEXCore::Core::InternalThreadState* Thread);
void CopyMemoryMapping(FEXCore::Core::InternalThreadState* ParentThread, FEXCore::Core::InternalThreadState* ChildThread);
uint8_t GetGPRSize() const {
return Config.Is64BitMode ? 8 : 4;
}
FEXCore::JITSymbols Symbols;
void GetVDSOSigReturn(VDSOSigReturn* VDSOPointers) override {
if (VDSOPointers->VDSO_kernel_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_sigreturn = reinterpret_cast<void*>(X86CodeGen.sigreturn_32);
}
if (VDSOPointers->VDSO_kernel_rt_sigreturn == nullptr) {
VDSOPointers->VDSO_kernel_rt_sigreturn = reinterpret_cast<void*>(X86CodeGen.rt_sigreturn_32);
}
}
FEXCore::Utils::PooledAllocatorVirtual OpDispatcherAllocator;
FEXCore::Utils::PooledAllocatorVirtual FrontendAllocator;
// If Atomic-based TSO emulation is enabled or not.
bool IsAtomicTSOEnabled() const {
return AtomicTSOEmulationEnabled;
}
void SetHardwareTSOSupport(bool HardwareTSOSupported) override {
SupportsHardwareTSO = HardwareTSOSupported;
UpdateAtomicTSOEmulationConfig();
}
// Returns if Software TSO emulation is required.
// NOTE: This doesn't necessary return if Atomic-based TSO is currently enabled.
// This will still return true if on a single thread and TSO is currently disabled.
//
// This is to ensure that if early initialization checks CPU features and TSO /could/ be enabled, that
// we return consistent results.
//
// To check if Atomic TSO is currently enabled in the JIT, use `IsAtomicTSOEnabled` instead.
bool SoftwareTSORequired() const {
if (SupportsHardwareTSO) {
return false;
}
return Config.TSOEnabled;
}
void EnableExitOnHLT() override {
ExitOnHLT = true;
}
bool ExitOnHLTEnabled() const {
return ExitOnHLT;
}
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
// If the hardware supports TSO then we don't need to emulate it through atomics.
AtomicTSOEmulationEnabled = false;
} else {
// Atomic TSO emulation only enabled if the config option is enabled.
AtomicTSOEmulationEnabled = (IsMemoryShared || !Config.TSOAutoMigration) && Config.TSOEnabled;
}
}
private:
/**
* @brief Initializes the JIT compilers for the thread
*
* @param State The internal FEX thread state object
*
* InitializeCompiler is called inside of CreateThread, so you likely don't need this
*/
void InitializeCompiler(FEXCore::Core::InternalThreadState* Thread);
void AddBlockMapping(FEXCore::Core::InternalThreadState* Thread, uint64_t Address, void* Ptr);
IR::AOTIRCaptureCache IRCaptureCache;
fextl::unique_ptr<FEXCore::CodeSerialize::CodeObjectSerializeService> CodeObjectCacheService;
bool StartPaused = false;
bool IsMemoryShared = false;
bool SupportsHardwareTSO = false;
bool AtomicTSOEmulationEnabled = true;
bool ExitOnHLT = false;
FEX_CONFIG_OPT(AppFilename, APP_FILENAME);
std::shared_mutex CustomIRMutex;
std::atomic<bool> HasCustomIRHandlers {};
fextl::unordered_map<uint64_t, std::tuple<CustomIREntrypointHandler, void*, void*>> CustomIRHandlers;
};
} // namespace FEXCore::Context
File diff suppressed because it is too large. Load diff
@@ -2,9 +2,6 @@
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include <aarch64/assembler-aarch64.h>
@@ -22,6 +19,8 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#include <array>
#include <cstddef>
@@ -35,85 +34,74 @@ class ContextImpl;
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
constexpr auto STATE = ARMEmitter::XReg::x28;
#ifndef _M_ARM_64EC
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x0;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x1;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x2;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
constexpr auto TMP1 = ARMEmitter::XReg::x0;
constexpr auto TMP2 = ARMEmitter::XReg::x1;
constexpr auto TMP3 = ARMEmitter::XReg::x2;
constexpr auto TMP4 = ARMEmitter::XReg::x3;
constexpr bool TMP_ABIARGS = true;
// We pin r26/r27 as PF/AF respectively, this is internal FEX ABI.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r26;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r27;
constexpr auto REG_PF = ARMEmitter::Reg::r26;
constexpr auto REG_AF = ARMEmitter::Reg::r27;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
#else
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x10;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x11;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x12;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x13;
constexpr auto TMP1 = ARMEmitter::XReg::x10;
constexpr auto TMP2 = ARMEmitter::XReg::x11;
constexpr auto TMP3 = ARMEmitter::XReg::x12;
constexpr auto TMP4 = ARMEmitter::XReg::x13;
constexpr bool TMP_ABIARGS = false;
// We pin r11/r12 as PF/AF respectively for arm64ec, as r26/r27 are used for SRA.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r9;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r24;
constexpr auto REG_PF = ARMEmitter::Reg::r9;
constexpr auto REG_AF = ARMEmitter::Reg::r24;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v16;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v17;
constexpr auto VTMP1 = ARMEmitter::VReg::v16;
constexpr auto VTMP2 = ARMEmitter::VReg::v17;
// Entry/Exit ABI
constexpr auto EC_CALL_CHECKER_PC_REG = ARMEmitter::XReg::x9;
constexpr auto EC_ENTRY_CPUAREA_REG = ARMEmitter::XReg::x17;
#endif
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_16B = FEXCore::ARMEmitter::PReg::p6;
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PReg::p7;
constexpr ARMEmitter::PRegister PRED_TMP_16B = ARMEmitter::PReg::p6;
constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
Arm64Emitter(FEXCore::Context::ContextImpl *ctx, void* EmissionPtr = nullptr, size_t size = 0);
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
FEXCore::Context::ContextImpl *EmitterCTX;
FEXCore::Context::ContextImpl* EmitterCTX;
vixl::aarch64::CPU CPU;
std::span<const FEXCore::ARMEmitter::Register> ConfiguredDynamicRegisterBase{};
std::span<const FEXCore::ARMEmitter::Register> StaticRegisters{};
std::span<const FEXCore::ARMEmitter::Register> GeneralRegisters{};
std::span<const std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>> GeneralPairRegisters{};
std::span<const FEXCore::ARMEmitter::VRegister> StaticFPRegisters{};
std::span<const FEXCore::ARMEmitter::VRegister> GeneralFPRegisters{};
std::span<const ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
std::span<const ARMEmitter::Register> StaticRegisters {};
std::span<const ARMEmitter::Register> GeneralRegisters {};
std::span<const ARMEmitter::VRegister> StaticFPRegisters {};
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
/**
* @name Register Allocation
* @{ */
constexpr static uint32_t RegisterClasses = 6;
constexpr static uint64_t GPRBase = (0ULL << 32);
constexpr static uint64_t FPRBase = (1ULL << 32);
constexpr static uint64_t GPRPairBase = (2ULL << 32);
/** @} */
constexpr static uint8_t RA_32 = 0;
constexpr static uint8_t RA_64 = 1;
constexpr static uint8_t RA_FPR = 2;
void LoadConstant(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
// Register 0-18 + 29 + 30 are caller saved
@@ -124,13 +112,13 @@ protected:
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
// Generic push and pop vector registers.
void PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs);
void PushVectorRegisters(ARMEmitter::Register TmpReg, bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(ARMEmitter::Register TmpReg, std::span<const ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const ARMEmitter::Register> Regs);
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
void PushDynamicRegsAndLR(ARMEmitter::Register TmpReg);
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
@@ -146,14 +134,13 @@ protected:
// Callee Saved:
// - X9-X15, X19-X31
// - Low 128-bits of v8-v31
void SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true);
void SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, bool FPRs = true);
void FillForPreserveAllABICall(bool FPRs = true);
void SpillForABICall(bool SupportsPreserveAllABI, FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true) {
void SpillForABICall(bool SupportsPreserveAllABI, ARMEmitter::Register TmpReg, bool FPRs = true) {
if (SupportsPreserveAllABI) {
SpillForPreserveAllABICall(TmpReg, FPRs);
}
else {
} else {
SpillStaticRegs(TmpReg, FPRs);
PushDynamicRegsAndLR(TmpReg);
}
@@ -162,8 +149,7 @@ protected:
void FillForABICall(bool SupportsPreserveAllABI, bool FPRs = true) {
if (SupportsPreserveAllABI) {
FillForPreserveAllABICall(FPRs);
}
else {
} else {
PopDynamicRegsAndLR();
FillStaticRegs(FPRs);
}
@@ -185,8 +171,7 @@ protected:
template<typename R, typename... P>
void GenerateRuntimeCall(R (*Function)(P...)) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
uintptr_t FunctionAddress = reinterpret_cast<uintptr_t>(Function);
@@ -204,8 +189,7 @@ protected:
template<typename R, typename... P>
void GenerateIndirectRuntimeCall(ARMEmitter::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(&(vixl::aarch64::Simulator::RuntimeCallStructHelper<R, P...>::Wrapper));
hlt(vixl::aarch64::kIndirectRuntimeCallOpcode);
@@ -221,8 +205,8 @@ protected:
template<>
void GenerateIndirectRuntimeCall<float, __uint128_t>(ARMEmitter::Register Reg) {
uintptr_t SimulatorWrapperAddress = reinterpret_cast<uintptr_t>(
&(vixl::aarch64::Simulator::RuntimeCallStructHelper<float, __uint128_t>::Wrapper));
uintptr_t SimulatorWrapperAddress =
reinterpret_cast<uintptr_t>(&(vixl::aarch64::Simulator::RuntimeCallStructHelper<float, __uint128_t>::Wrapper));
hlt(vixl::aarch64::kIndirectRuntimeCallOpcode);
@@ -262,4 +246,4 @@ protected:
#endif
};
}
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
@@ -1,106 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef>
#include <cstdint>
#include <cstring>
namespace FEXCore::ARMEmitter {
class Buffer {
public:
Buffer() {
SetBuffer(nullptr, 0);
}
Buffer(uint8_t* Base, uint64_t BaseSize) {
SetBuffer(Base, BaseSize);
}
void SetBuffer(uint8_t* Base, uint64_t BaseSize) {
BufferBase = Base;
CurrentOffset = BufferBase;
Size = BaseSize;
}
void dc8(uint8_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc16(uint16_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc32(uint32_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void dc64(uint64_t Data) {
decltype(Data) *Memory = reinterpret_cast<decltype(Data)*>(CurrentOffset);
*Memory = Data;
CurrentOffset += sizeof(Data);
}
void EmitString(const char *String) {
const auto StringLength = strlen(String);
memcpy(CurrentOffset, String, StringLength);
CurrentOffset += StringLength;
}
void Align() {
// Align the buffer to instruction size
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & 0b11;
if (!CurrentAlignment) {
return;
}
CurrentOffset += 4 - CurrentAlignment;
}
template<typename T>
T GetCursorAddress() const {
return reinterpret_cast<T>(CurrentOffset);
}
static void ClearICache(void* Begin, std::size_t Length) {
__builtin___clear_cache(static_cast<char*>(Begin), static_cast<char*>(Begin) + Length);
}
size_t GetCursorOffset() const {
return static_cast<size_t>(CurrentOffset - BufferBase);
}
uint8_t *GetBufferBase() const {
return BufferBase;
}
void CursorIncrement(size_t Size) {
CurrentOffset += Size;
}
void SetCursorOffset(size_t Offset) {
CurrentOffset = BufferBase + Offset;
}
uint64_t GetBufferSize() const {
return Size;
}
template<typename T>
size_t GetCursorOffsetFromAddress(const T* Address) const {
return static_cast<size_t>(reinterpret_cast<const uint8_t*>(Address) - BufferBase);
}
protected:
void ResetBuffer() {
CurrentOffset = BufferBase;
}
uint8_t* BufferBase;
uint8_t* CurrentOffset;
uint64_t Size;
};
}
@@ -1,850 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <aarch64/assembler-aarch64.h>
#include <array>
#include <cstdint>
#include <utility>
#include <type_traits>
/*
* Welcome to FEX-Emu's custom AArch64 emitter.
* This was written specifically to avoid the performance cost of the vixl emitter.
*
* There are some specific design constraints in this design to target a couple features:
* - High performance
* - Low CPU cache performance hit
* - Significantly reduced code footprint
* - Low number of branches
*
* These requirements are mostly achieved by removing a bunch of developer conveniences
* that vixl provides. The developer needs to take a lot of care to not shoot themselves in the foot.
*
* Misc design decisions:
* - Registers are encoded as basic uint32_t enums.
* - Converting between different registers is zero-cost.
* - Passing around as arguments are as cheap as registers
* - Contrast to vixl where every register requires living on the stack.
* - Registers can get encoded in to instructions with a simple `BFM` instruction.
*
* - Instructions are very simply emitted, allowing direct inlining most of the time.
* - These are simple enough that multiple back-to-back instructions get optimized to 128-bit load-store operations.
* - Contrast to vixl where pretty much no instruction emitter gets inlined.
*
* - Instruction emitters are /mostly/ unsized. Most instructions take a size argument first, which gets encoded
* directly in to the instruction.
* - Contrast to vixl where the register arguments are how the instructions determine operating size.
* - Size argument allows FEX to use `CSEL` to select a size at runtime, instead of branching.
* - Some instructions are explicitly sized based on register type. Read comments in the respective `inl` files to
* see why.
* Some scalar/vector operations are an example of this.
*
* - Almost zero helper functions.
* - Primary exception to this rule is load-store operations. These will use a helper to make
* it easier to select the correct load-store instruction. Mostly because these are a nightmare selecting
* the right instruction.
*/
namespace FEXCore::ARMEmitter {
/*
* This `Size` enum is used for most ALU operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class Size : uint32_t {
i32Bit = 0,
i64Bit,
};
// This allows us to get the `Size` enum in bits.
[[nodiscard]]
constexpr size_t RegSizeInBits(Size size) {
return size_t{32} << FEXCore::ToUnderlying(size);
}
/* This `SubRegSize` enum is used for most ASIMD operations.
* These follow the AArch64 encoding style in most cases.
*/
enum class SubRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
i128Bit = 0b100,
};
// This allows us to get the `SubRegSize` in bits.
[[nodiscard]]
constexpr size_t SubRegSizeInBits(SubRegSize size) {
return size_t{8} << FEXCore::ToUnderlying(size);
}
/* This `ScalarRegSize` enum is used for most scalar float
* operations.
*
* This is specifically duplicated from `SubRegSize` to have strongly
* typed functions.
*
* `ScalarRegSize` specifically doesn't have `i128Bit` because scalar operations
* can't operate at 128-bit.
*/
enum class ScalarRegSize : uint32_t {
i8Bit = 0b00,
i16Bit = 0b01,
i32Bit = 0b10,
i64Bit = 0b11,
};
// This allows us to get the `ScalarRegSize` in bits.
[[nodiscard]]
constexpr size_t ScalarRegSizeInBits(ScalarRegSize size) {
return size_t{8} << FEXCore::ToUnderlying(size);
}
/* This `VectorRegSizePair` union allows us to have an overlapping type
* to select a scalar operation or a vector depending on which operation
* we pass in.
* Useful in FEX's vector operations that behave as scalar or vector
* depending on various factors. But since the operation will have the sa,e
* element size, we want to choose the operation more easily
*/
union VectorRegSizePair {
ScalarRegSize Scalar;
SubRegSize Vector;
};
// This allows us to create a `VectorRegSizePair` union.
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(SubRegSize size) {
return VectorRegSizePair {.Vector = size};
}
[[nodiscard]]
constexpr VectorRegSizePair ToVectorSizePair(ScalarRegSize size) {
return VectorRegSizePair {.Scalar = size};
}
// This `ShiftType` enum is used for ALU shift-register encoded instructions.
enum class ShiftType : uint32_t {
LSL = 0,
LSR,
ASR,
ROR,
};
// This `ExtendedType` enum is used for ALU extended-register encoded instructions.
enum class ExtendedType : uint32_t {
UXTB = 0b000,
UXTH = 0b001,
UXTW = 0b010,
UXTX = 0b011,
SXTB = 0b100,
SXTH = 0b101,
SXTW = 0b110,
SXTX = 0b111,
LSL_32 = UXTW,
LSL_64 = UXTX,
};
// This `Condition` enum is used for various conditional instructions.
enum class Condition : uint32_t {
// Meaning: Int - Float
CC_EQ = 0, // Equal - Equal
CC_NE, // Not Eq - Not Eq or unordered
CC_CS, // Carry set - Greater than, equal, or unordered
CC_CC, // Carry clear - Less than
CC_MI, // Minus/Negative - Less than
CC_PL, // Plus, positive or zero - GT, equal, or unordered
CC_VS, // Overflow - Unordered
CC_VC, // No Overflow - Ordered
CC_HI, // Unsigned higher - GT, or unordered
CC_LS, // Unsigned lower or same - LT or EQ
CC_GE, // Signed GT or EQ - GT or EQ
CC_LT, // Signed LT - LT or Unordered
CC_GT, // Signed GT - GT
CC_LE, // Signed LT or EQ - LT, EQ, or Unordered
CC_AL, // Always - Always
CC_NV, // Always - Always
// Aliases
CC_HS = CC_CS,
CC_LO = CC_CC,
};
/*
* This `StatusFlags` enum is used for conditional compare encoded instructions.
* These directly encode to the `nzcv` flags.
*/
enum class StatusFlags : uint32_t {
None = 0,
Flag_V = 0b0001,
Flag_C = 0b0010,
Flag_Z = 0b0100,
Flag_N = 0b1000,
Flag_NZCV = Flag_N | Flag_Z | Flag_C | Flag_V,
};
/*
* This `IndexType` enum is used for load-store instructions.
* Not all load-store instructions use this, so the user needs to be careful.
*/
enum class IndexType {
POST,
OFFSET,
PRE,
UNPRIVILEGED,
};
// Used with adr and scalar + vector load/store variants to denote
// a modifier operation.
enum class SVEModType : uint8_t {
MOD_UXTW,
MOD_SXTW,
MOD_LSL,
MOD_NONE,
};
/* This `SVEMemOperand` class is used for the helper SVE load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class SVEMemOperand final {
public:
enum class Type {
ScalarPlusScalar,
ScalarPlusImm,
ScalarPlusVector,
VectorPlusImm,
};
SVEMemOperand(XRegister rn, XRegister rm = XReg::zr)
: rn {rn}
, MemType{Type::ScalarPlusScalar}
, MetaType {
.ScalarScalarType {
.rm = rm,
}
} {}
SVEMemOperand(XRegister rn, int32_t imm = 0)
: rn {rn}
, MemType{Type::ScalarPlusImm}
, MetaType {
.ScalarImmType {
.Imm = imm,
}
} {}
SVEMemOperand(XRegister rn, ZRegister zm, SVEModType mod = SVEModType::MOD_NONE, uint8_t scale = 0)
: rn{rn}
, MemType{Type::ScalarPlusVector}
, MetaType {
.ScalarVectorType {
.zm = zm,
.mod = mod,
.scale = scale,
}
} {}
SVEMemOperand(ZRegister zn, uint32_t imm)
: rn{Register{zn.Idx()}}
, MemType{Type::VectorPlusImm}
, MetaType {
.VectorImmType{
.Imm = imm,
}
} {}
[[nodiscard]] bool IsScalarPlusScalar() const {
return MemType == Type::ScalarPlusScalar;
}
[[nodiscard]] bool IsScalarPlusImm() const {
return MemType == Type::ScalarPlusImm;
}
[[nodiscard]] bool IsScalarPlusVector() const {
return MemType == Type::ScalarPlusVector;
}
[[nodiscard]] bool IsVectorPlusImm() const {
return MemType == Type::VectorPlusImm;
}
union Data {
struct {
Register rm;
} ScalarScalarType;
struct {
int32_t Imm;
} ScalarImmType;
struct {
ZRegister zm;
SVEModType mod;
uint8_t scale;
} ScalarVectorType;
struct {
// rn will be a ZRegister
uint32_t Imm;
} VectorImmType;
};
Register rn;
Type MemType;
Data MetaType;
};
/* This `ExtendedMemOperand` class is used for the helper load-store instructions.
* Load-store instructions are quite expressive, so having a helper that handles these differences is worth it.
*/
class ExtendedMemOperand final {
public:
ExtendedMemOperand(XRegister rn, XRegister rm = XReg::zr, ExtendedType Option = ExtendedType::LSL_64, uint32_t Shift = 0)
: rn {rn}
, MetaType {
.ExtendedType {
.Header = { .MemType = TYPE_EXTENDED },
.rm = rm,
.Option = Option,
.Shift = Shift,
}
} {}
ExtendedMemOperand(XRegister rn, IndexType Index = IndexType::OFFSET, int32_t Imm = 0)
: rn {rn}
, MetaType {
.ImmType {
.Header = { .MemType = TYPE_IMM },
.Index = Index,
.Imm = Imm,
}
} {}
Register rn;
enum Type {
TYPE_EXTENDED,
TYPE_IMM,
};
struct HeaderStruct {
Type MemType;
};
union {
HeaderStruct Header;
struct {
HeaderStruct Header;
Register rm;
ExtendedType Option;
uint32_t Shift;
} ExtendedType;
struct {
HeaderStruct Header;
IndexType Index;
int32_t Imm;
} ImmType;
} MetaType;
};
template<uint32_t op0, uint32_t op1, uint32_t CRn, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenSystemReg() {
return op0 << 19 |
op1 << 16 |
CRn << 12 |
CRm << 8 |
op2 << 5;
};
// This `SystemRegister` enum is used for the mrs/msr instructions.
enum class SystemRegister : uint32_t {
CTR_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b001>(),
DCZID_EL0 = GenSystemReg<0b11, 0b011, 0b0000, 0b0000, 0b111>(),
TPIDR_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b010>(),
RNDR = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b000>(),
RNDRRS = GenSystemReg<0b11, 0b011, 0b0010, 0b0100, 0b001>(),
NZCV = GenSystemReg<0b11, 0b011, 0b0100, 0b0010, 0b000>(),
FPCR = GenSystemReg<0b11, 0b011, 0b0100, 0b0100, 0b000>(),
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>(),
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>(),
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
constexpr uint32_t GenDCReg() {
return op1 << 16 |
CRm << 8 |
op2 << 5;
};
// This `DataCacheOperation` enum is used for the dc instruction.
enum class DataCacheOperation : uint32_t {
IVAC = GenDCReg<0b000, 0b0110, 0b001>(),
ISW = GenDCReg<0b000, 0b0110, 0b010>(),
CSW = GenDCReg<0b000, 0b1010, 0b010>(),
CISW = GenDCReg<0b000, 0b1110, 0b010>(),
ZVA = GenDCReg<0b011, 0b0100, 0b001>(),
CVAC = GenDCReg<0b011, 0b1010, 0b001>(),
CVAU = GenDCReg<0b011, 0b1011, 0b001>(),
CIVAC = GenDCReg<0b011, 0b1110, 0b001>(),
// MTE2
IGVAC = GenDCReg<0b000, 0b0110, 0b011>(),
IGSW = GenDCReg<0b000, 0b0110, 0b100>(),
IGDVAC = GenDCReg<0b000, 0b0110, 0b101>(),
IGDSW = GenDCReg<0b000, 0b0110, 0b110>(),
CGSW = GenDCReg<0b000, 0b1010, 0b100>(),
CGDSW = GenDCReg<0b000, 0b1010, 0b110>(),
CIGSW = GenDCReg<0b000, 0b1110, 0b100>(),
CIGDSW = GenDCReg<0b000, 0b1110, 0b110>(),
// MTE
GVA = GenDCReg<0b011, 0b0100, 0b011>(),
GZVA = GenDCReg<0b011, 0b0100, 0b100>(),
CGVAC = GenDCReg<0b011, 0b1010, 0b011>(),
CGDVAC = GenDCReg<0b011, 0b1010, 0b101>(),
CGVAP = GenDCReg<0b011, 0b1100, 0b011>(),
CGDVAP = GenDCReg<0b011, 0b1100, 0b101>(),
CGVADP = GenDCReg<0b011, 0b1101, 0b011>(),
CGDVADP = GenDCReg<0b011, 0b1101, 0b101>(),
CIGVAC = GenDCReg<0b011, 0b1110, 0b011>(),
CIGDVAC = GenDCReg<0b011, 0b1110, 0b101>(),
// DPB
CVAP = GenDCReg<0b011, 0b1100, 0b001>(),
// DPB2
CVADP = GenDCReg<0b011, 0b1101, 0b001>(),
};
template<uint32_t CRm, uint32_t op2>
constexpr uint32_t GenHintBarrierReg() {
return CRm << 8 |
op2 << 5;
}
// This `HintRegister` enum is used for the hint instruction.
enum class HintRegister : uint32_t {
NOP = GenHintBarrierReg<0b0000, 0b000>(),
YIELD = GenHintBarrierReg<0b0000, 0b001>(),
WFE = GenHintBarrierReg<0b0000, 0b010>(),
WFI = GenHintBarrierReg<0b0000, 0b011>(),
SEV = GenHintBarrierReg<0b0000, 0b100>(),
SEVL = GenHintBarrierReg<0b0000, 0b101>(),
DGH = GenHintBarrierReg<0b0000, 0b110>(),
CSDB = GenHintBarrierReg<0b0010, 0b100>(),
};
// This `BarrierRegister` enum is used for the various barrier instructions.
enum class BarrierRegister : uint32_t {
CLREX = GenHintBarrierReg<0b0000, 0b010>(),
TCOMMIT = GenHintBarrierReg<0b0000, 0b011>(),
DSB = GenHintBarrierReg<0b0000, 0b100>(),
DMB = GenHintBarrierReg<0b0000, 0b101>(),
ISB = GenHintBarrierReg<0b0000, 0b110>(),
SB = GenHintBarrierReg<0b0000, 0b111>(),
};
// This `BarrierScope` enum is used for the dsb/dmb instructions.
enum class BarrierScope : uint32_t {
// Outer shareable
OSHLD = 0b0001,
OSHST = 0b0010,
OSH = 0b0011,
// Non shareable
NSHLD = 0b0101,
NSHST = 0b0110,
NSH = 0b0111,
// Inner shareable
ISHLD = 0b1001,
ISHST = 0b1010,
ISH = 0b1011,
// Full System visibility
LD = 0b1101,
ST = 0b1110,
SY = 0b1111,
};
// This `Prefetch` enum is used for prefetch instructions.
enum class Prefetch : uint32_t {
// Prefetch for load
PLDL1KEEP = 0b00000,
PLDL1STRM = 0b00001,
PLDL2KEEP = 0b00010,
PLDL2STRM = 0b00011,
PLDL3KEEP = 0b00100,
PLDL3STRM = 0b00101,
// Preload instructions
PLIL1KEEP = 0b01000,
PLIL1STRM = 0b01001,
PLIL2KEEP = 0b01010,
PLIL2STRM = 0b01011,
PLIL3KEEP = 0b01100,
PLIL3STRM = 0b01101,
// Preload for store
PSTL1KEEP = 0b10000,
PSTL1STRM = 0b10001,
PSTL2KEEP = 0b10010,
PSTL2STRM = 0b10011,
PSTL3KEEP = 0b10100,
PSTL3STRM = 0b10101,
};
// This `PredicatePattern` enun is used for some SVE instructions.
enum class PredicatePattern : uint32_t {
SVE_POW2 = 0b00000,
SVE_VL1 = 0b00001,
SVE_VL2 = 0b00010,
SVE_VL3 = 0b00011,
SVE_VL4 = 0b00100,
SVE_VL5 = 0b00101,
SVE_VL6 = 0b00110,
SVE_VL7 = 0b00111,
SVE_VL8 = 0b01000,
SVE_VL16 = 0b01001,
SVE_VL32 = 0b01010,
SVE_VL64 = 0b01011,
SVE_VL128 = 0b01100,
SVE_VL256 = 0b01101,
SVE_MUL4 = 0b11101,
SVE_MUL3 = 0b11110,
SVE_ALL = 0b11111,
};
// Used with SVE FP immediate arithmetic instructions
enum class SVEFAddSubImm : uint32_t {
_0_5,
_1_0,
};
enum class SVEFMulImm : uint32_t {
_0_5,
_2_0,
};
enum class SVEFMaxMinImm : uint32_t {
_0_0,
_1_0,
};
/* This `BackwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `below` an instruction that uses it.
* Which means that a branch would jump backwards.
*/
struct BackwardLabel {
uint8_t *Location{};
};
/* This `SingleUseForwardLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is logically `above` an instruction that uses it.
* Which means that a branch would jump forwards.
*
* The `ForwardLabel` struct can be bound to multiple instructions, so it needs a vector for each bind instruction type.
*/
struct SingleUseForwardLabel {
enum class InstType {
UNKNOWN,
ADR,
ADRP,
B,
BC,
TEST_BRANCH,
RELATIVE_LOAD,
LONG_ADDRESS_GEN,
};
uint8_t *Location{};
InstType Type = InstType::UNKNOWN;
};
struct ForwardLabel {
fextl::vector<SingleUseForwardLabel> Insts{};
};
/* This `BiDirectionalLabel` struct used for retaining a location for PC-Relative instructions.
* This is specifically a label for a target that is in either direction of an instruction that uses it.
* Which means a branch could jump backwards or forwards depending on situation.
*/
struct BiDirectionalLabel {
BackwardLabel Backward;
ForwardLabel Forward;
};
static inline void AddLocationToLabel(SingleUseForwardLabel *Label, SingleUseForwardLabel&& Location) {
LOGMAN_THROW_A_FMT(Label->Type == SingleUseForwardLabel::InstType::UNKNOWN, "Trying to bind a SingleUseForwardLabel to multiple locations. Use ForwardLabel instead.");
*Label = std::move(Location);
}
static inline void AddLocationToLabel(ForwardLabel *Label, SingleUseForwardLabel&& Location) {
Label->Insts.emplace_back(std::move(Location));
}
// Some FCMA ASIMD instructions support a rotation argument.
enum class Rotation : uint32_t {
ROTATE_0 = 0b00,
ROTATE_90 = 0b01,
ROTATE_180 = 0b10,
ROTATE_270 = 0b11,
};
// Concept for contraining some instructions to accept only an XRegister or WRegister.
// Particularly for operations that differ encodings depending on which one is used.
template <typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
// For example, a set of registers like:
//
// v1, v2, v3 and
// v31, v0, v1
//
// would both be considered sequential sequences, and some instructions in particular
// limit register lists to these kind of sequences.
//
template <typename T, typename... Args>
constexpr bool AreVectorsSequential(T first, const Args&... args) {
// Ensure we always have a pair of registers to compare against.
static_assert(sizeof...(args) >= 1, "Number of arguments must be greater than 1");
const auto fn = [](auto& lhs, const auto& rhs) {
const auto result = ((lhs.Idx() + 1) % 32) == rhs.Idx();
lhs = rhs;
return result;
};
return (fn(first, args) && ...);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
// Choices:
// - Size of ops passed as an argument rather than template to let the compiler use csel instead of branching.
// - Registers are unsized so they can be passed in a GPR and not need conversion operations
class Emitter : public FEXCore::ARMEmitter::Buffer {
public:
Emitter() = default;
Emitter(uint8_t* Base, uint64_t BaseSize)
: Buffer (Base, BaseSize) {
}
// Bind a backward label to an address.
// Address that is bound is the current emitter location.
void Bind(BackwardLabel *Label) {
LOGMAN_THROW_AA_FMT(Label->Location == nullptr, "Trying to bind a label twice");
Label->Location = GetCursorAddress<uint8_t*>();
}
void Bind(const SingleUseForwardLabel *Label) {
uint8_t *CurrentAddress = GetCursorAddress<uint8_t*>();
// Patch up the instructions
switch (Label->Type) {
case SingleUseForwardLabel::InstType::ADR: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRRange(Imm), "Unscaled offset too large");
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::ADRP: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(IsADRPRange(Imm) && IsADRPAligned(Imm), "Unscaled offset too large");
Imm >>= 12;
uint32_t InstMask = 0b11 << 29 | 0b1111'1111'1111'1111'111 << 5;
uint32_t Offset = static_cast<uint32_t>(Imm) & 0x3F'FFFF;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= (Offset & 0b11) << 29;
Inst |= (Offset >> 2) << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::B: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -134217728 && Imm <= 134217724 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FF'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~InstMask;
Inst |= Offset;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::TEST_BRANCH: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x3FFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::BC:
case SingleUseForwardLabel::InstType::RELATIVE_LOAD: {
uint32_t *Instruction = reinterpret_cast<uint32_t*>(Label->Location);
int64_t Imm = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(Instruction);
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
Imm >>= 2;
uint32_t InstMask = 0x7'FFFF;
uint32_t Offset = static_cast<uint32_t>(Imm) & InstMask;
uint32_t Inst = *Instruction & ~(InstMask << 5);
Inst |= Offset << 5;
*Instruction = Inst;
break;
}
case SingleUseForwardLabel::InstType::LONG_ADDRESS_GEN: {
uint32_t *Instructions = reinterpret_cast<uint32_t*>(Label->Location);
int64_t ImmInstOne = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[0]);
int64_t ImmInstTwo = reinterpret_cast<int64_t>(CurrentAddress) - reinterpret_cast<int64_t>(&Instructions[1]);
auto OriginalOffset = GetCursorOffset();
auto InstOffset = GetCursorOffsetFromAddress(Instructions);
SetCursorOffset(InstOffset);
// We encoded the destination register in to the first instruction space.
// Read it back.
ARMEmitter::Register DestReg(Instructions[0]);
if (IsADRRange(ImmInstTwo)) {
// If within ADR range from the second instruction, then we can emit NOP+ADR
nop();
adr(DestReg, static_cast<uint32_t>(ImmInstTwo) & 0x7FFF);
}
else if (IsADRPRange(ImmInstOne)) {
// If within ADRP range from the first instruction, then we are /definitely/ in range for the second instruction.
// First check if we are in non-offset range for second instruction.
if (IsADRPAligned(reinterpret_cast<uint64_t>(CurrentAddress))) {
// We can emit nop + adrp
nop();
adrp(DestReg, static_cast<uint32_t>(ImmInstTwo >> 12) & 0x7FFF);
}
else {
// Not aligned, need adrp + add
adrp(DestReg, static_cast<uint32_t>(ImmInstOne >> 12) & 0x7FFF);
add(ARMEmitter::Size::i64Bit, DestReg, DestReg, ImmInstOne & 0xFFF);
}
}
else {
LOGMAN_MSG_A_FMT("Unscaled offset is too large");
FEX_UNREACHABLE;
}
SetCursorOffset(OriginalOffset);
break;
}
default: LOGMAN_MSG_A_FMT("Unexpected inst type in label fixup");
}
}
// Bind a forward label to a location.
// This walks all the instructions in the label's vector.
// Then backpatching all instructions that have used the label.
template<bool WarnAboutEmpty = false>
void Bind(ForwardLabel *Label) {
if constexpr (WarnAboutEmpty) {
LOGMAN_THROW_A_FMT(Label->Insts.empty() == false, "Binding forward label that didn't have any instructions using it");
}
for (auto &Inst : Label->Insts) {
Bind(&Inst);
}
}
// Bind a bidirectional location to a location.
// Binds both forwards and backwards depending on how the label was used.
void Bind(BiDirectionalLabel *Label) {
if (!Label->Backward.Location) {
Bind(&Label->Backward);
}
Bind<false>(&Label->Forward);
}
public:
// TODO: Implement SME when it matters.
#include "Interface/Core/ArchHelpers/CodeEmitter/ALUOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/BranchOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/LoadstoreOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SystemOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ScalarOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ASIMDOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SVEOps.inl"
private:
template<typename T>
uint32_t Encode_ra(T Reg) const {
return Reg.Idx() << 10;
}
uint32_t Encode_ra(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rt2(T Reg) const {
return Reg.Idx() << 10;
}
template<>
uint32_t Encode_rt2(uint32_t Reg) const {
return Reg << 10;
}
template<typename T>
uint32_t Encode_rm(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rm(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rs(T Reg) const {
return Reg.Idx() << 16;
}
uint32_t Encode_rs(uint32_t Reg) const {
return Reg << 16;
}
template<typename T>
uint32_t Encode_rn(T Reg) const {
return Reg.Idx() << 5;
}
uint32_t Encode_rn(uint32_t Reg) const {
return Reg << 5;
}
template<typename T>
uint32_t Encode_rd(T Reg) const {
return Reg.Idx();
}
uint32_t Encode_rd(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_rt(T Reg) const {
return Reg.Idx();
}
template<>
uint32_t Encode_rt(Prefetch Reg) const {
return FEXCore::ToUnderlying(Reg);
}
uint32_t Encode_rt(uint32_t Reg) const {
return Reg;
}
template<typename T>
uint32_t Encode_pd(T Reg) const {
return FEXCore::ToUnderlying(Reg);
}
};
}
File diff suppressed because it is too large. Load diff
@@ -6,48 +6,46 @@
#include <utility>
namespace FEXCore {
void BlockSamplingData::DumpBlockData() {
std::fstream Output;
Output.open("output.csv", std::fstream::out | std::fstream::binary);
void BlockSamplingData::DumpBlockData() {
std::fstream Output;
Output.open("output.csv", std::fstream::out | std::fstream::binary);
if (!Output.is_open())
return;
Output << "Entry, Min, Max, Total, Calls, Average" << std::endl;
for (auto it : SamplingMap) {
if (!it.second->TotalCalls)
continue;
Output << "0x" << std::hex << it.first
<< ", " << std::dec << it.second->Min
<< ", " << std::dec << it.second->Max
<< ", " << std::dec << it.second->TotalTime
<< ", " << std::dec << it.second->TotalCalls
<< ", " << std::dec << ((double)it.second->TotalTime / (double)it.second->TotalCalls)
<< std::endl;
}
Output.close();
LogMan::Msg::DFmt("Dumped {} blocks of sampling data", SamplingMap.size());
if (!Output.is_open()) {
return;
}
BlockSamplingData::BlockData *BlockSamplingData::GetBlockData(uint64_t RIP) {
auto it = SamplingMap.find(RIP);
if (it != SamplingMap.end()) {
return it->second;
}
BlockData *NewData = new BlockData{};
memset(NewData, 0, sizeof(BlockData));
NewData->Min = ~0ULL;
SamplingMap[RIP] = NewData;
return NewData;
}
Output << "Entry, Min, Max, Total, Calls, Average" << std::endl;
BlockSamplingData::~BlockSamplingData() {
DumpBlockData();
for (auto it : SamplingMap) {
delete it.second;
for (auto it : SamplingMap) {
if (!it.second->TotalCalls) {
continue;
}
SamplingMap.clear();
Output << "0x" << std::hex << it.first << ", " << std::dec << it.second->Min << ", " << std::dec << it.second->Max << ", " << std::dec
<< it.second->TotalTime << ", " << std::dec << it.second->TotalCalls << ", " << std::dec
<< ((double)it.second->TotalTime / (double)it.second->TotalCalls) << std::endl;
}
Output.close();
LogMan::Msg::DFmt("Dumped {} blocks of sampling data", SamplingMap.size());
}
BlockSamplingData::BlockData* BlockSamplingData::GetBlockData(uint64_t RIP) {
auto it = SamplingMap.find(RIP);
if (it != SamplingMap.end()) {
return it->second;
}
BlockData* NewData = new BlockData {};
memset(NewData, 0, sizeof(BlockData));
NewData->Min = ~0ULL;
SamplingMap[RIP] = NewData;
return NewData;
}
BlockSamplingData::~BlockSamplingData() {
DumpBlockData();
for (auto it : SamplingMap) {
delete it.second;
}
SamplingMap.clear();
}
} // namespace FEXCore
@@ -14,7 +14,7 @@ public:
uint64_t TotalCalls;
};
BlockData *GetBlockData(uint64_t RIP);
BlockData* GetBlockData(uint64_t RIP);
~BlockSamplingData();
void DumpBlockData();
@@ -22,4 +22,4 @@ public:
private:
std::unordered_map<uint64_t, BlockData*> SamplingMap;
};
}
} // namespace FEXCore
+328 -350
View File
@@ -2,8 +2,8 @@
#include "FEXCore/IR/IR.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
#ifndef _WIN32
#include <sys/prctl.h>
@@ -12,347 +12,326 @@
namespace FEXCore {
namespace CPU {
constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX][2] = {
{0x0003'0002'0001'0000ULL, 0x0007'0006'0005'0004ULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX
{0x000B'000A'0009'0008ULL, 0x000F'000E'000D'000CULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX_UPPER
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT_UPPER
{0x0000'0001'0000'0000ULL, 0x0000'0003'0000'0002ULL}, // NAMED_VECTOR_MOVMSKPS_SHIFT
{0x040B'0E01'0B0E'0104ULL, 0x0C03'0609'0306'090CULL}, // NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE
{0x0706'0504'FFFF'FFFFULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0110B
{0x0706'0504'0302'0100ULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0111B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1001B
{0x0706'0504'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1011B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1101B
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
};
constexpr static auto PSHUFLW_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// Expectation for this LUT is to simulate PSHUFLW with ARM's TBL (single register) instruction
// PSHUFLW behaviour:
// 16-bit words in [63:48], [47:32], [31:16], [15:0] are selected using the 8-bit Index.
// For 128-bit PSHUFLW, bits [127:64] are identity copied.
constexpr uint64_t IdentityCopyUpper = 0x0f'0e'0d'0c'0b'0a'09'08;
std::array<LUTType, 256> TotalLUT{};
uint64_t WordSelection[4] = {
0x01'00,
0x03'02,
0x05'04,
0x07'06,
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] =
(WordSelection[Word0] << 0) |
(WordSelection[Word1] << 16) |
(WordSelection[Word2] << 32) |
(WordSelection[Word3] << 48);
LUT.Val[1] = IdentityCopyUpper;
}
return TotalLUT;
}()
};
constexpr static auto PSHUFHW_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// Expectation for this LUT is to simulate PSHUFHW with ARM's TBL (single register) instruction
// PSHUFHW behaviour:
// 16-bit words in [127:112], [111:96], [95:80], [79:64] are selected using the 8-bit Index.
// Incoming words come from bits [127:64] of the source.
// Bits [63:0] are identity copied.
constexpr uint64_t IdentityCopyLower = 0x07'06'05'04'03'02'01'00;
std::array<LUTType, 256> TotalLUT{};
uint64_t WordSelection[4] = {
0x09'08,
0x0b'0a,
0x0d'0c,
0x0f'0e,
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] = IdentityCopyLower;
LUT.Val[1] =
(WordSelection[Word0] << 0) |
(WordSelection[Word1] << 16) |
(WordSelection[Word2] << 32) |
(WordSelection[Word3] << 48);
}
return TotalLUT;
}()
};
constexpr static auto PSHUFD_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// Expectation for this LUT is to simulate PSHUFD with ARM's TBL (single register) instruction
// PSHUFD behaviour:
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
std::array<LUTType, 256> TotalLUT{};
uint64_t WordSelection[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] =
(WordSelection[Word0] << 0) |
(WordSelection[Word1] << 32);
LUT.Val[1] =
(WordSelection[Word2] << 0) |
(WordSelection[Word3] << 32);
}
return TotalLUT;
}()
};
constexpr static auto SHUFPS_LUT {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
// Expectation for this LUT is to simulate SHUFPS with ARM's TBL (two register) instruction.
// SHUFPS behaviour:
// Two 32-bits words from each source are selected from each source in the lower and upper halves of the 128-bit destination.
// Dest[31:0] = Src1[<Word0>]
// Dest[63:32] = Src1[<Word1>]
// Dest[95:64] = Src2[<Word2>]
// Dest[127:96] = Src2[<Word3>]
std::array<LUTType, 256> TotalLUT{};
const uint64_t WordSelectionSrc1[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
constexpr static uint64_t NamedVectorConstants[FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX][2] = {
{0x0003'0002'0001'0000ULL, 0x0007'0006'0005'0004ULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX
{0x000B'000A'0009'0008ULL, 0x000F'000E'000D'000CULL}, // NAMED_VECTOR_INCREMENTAL_U16_INDEX_UPPER
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT_UPPER
{0x0000'0001'0000'0000ULL, 0x0000'0003'0000'0002ULL}, // NAMED_VECTOR_MOVMSKPS_SHIFT
{0x040B'0E01'0B0E'0104ULL, 0x0C03'0609'0306'090CULL}, // NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE
{0x0706'0504'FFFF'FFFFULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0110B
{0x0706'0504'0302'0100ULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0111B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1001B
{0x0706'0504'0302'0100ULL, 0x0F0E'0D0C'FFFF'FFFFULL}, // NAMED_VECTOR_BLENDPS_1011B
{0xFFFF'FFFF'0302'0100ULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1101B
{0x0706'0504'FFFF'FFFFULL, 0x0F0E'0D0C'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_1110B
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB
{0x8040'2010'0804'0201ULL, 0x8040'2010'0804'0201ULL}, // NAMED_VECTOR_MOVMASKB_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_ONE
{0xD49A'784B'CD1B'8AFEULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_LOG2_10
{0xB8AA'3B29'5C17'F0BCULL, 0x0000'0000'0000'3FFFULL}, // NAMED_VECTOR_X87_LOG2_E
{0xC90F'DAA2'2168'C235ULL, 0x0000'0000'0000'4000ULL}, // NAMED_VECTOR_X87_PI
{0x9A20'9A84'FBCF'F799ULL, 0x0000'0000'0000'3FFDULL}, // NAMED_VECTOR_X87_LOG10_2
{0xB172'17F7'D1CF'79ACULL, 0x0000'0000'0000'3FFEULL}, // NAMED_VECTOR_X87_LOG_2
};
// Src2 needs to offset each byte index by 16-bytes to pull from the second source.
const uint64_t WordSelectionSrc2[4] = {
0x03'02'01'00 + (0x10101010),
0x07'06'05'04 + (0x10101010),
0x0b'0a'09'08 + (0x10101010),
0x0f'0e'0d'0c + (0x10101010),
};
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] =
(WordSelectionSrc1[Word0] << 0) |
(WordSelectionSrc1[Word1] << 32);
LUT.Val[1] =
(WordSelectionSrc2[Word2] << 0) |
(WordSelectionSrc2[Word3] << 32);
}
return TotalLUT;
}()
};
constexpr static auto DPPS_MASK {
[]() consteval {
struct LUTType {
uint32_t Val[4];
};
std::array<LUTType, 16> TotalLUT{};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto &LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1U;
}
return 0U;
constexpr static auto PSHUFLW_LUT {[]() consteval {
struct LUTType {
uint64_t Val[2];
};
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
LUT.Val[2] = GetLUT(i, 2);
LUT.Val[3] = GetLUT(i, 3);
}
return TotalLUT;
}()
};
constexpr static auto DPPD_MASK {
[]() consteval {
struct LUTType {
uint64_t Val[2];
};
std::array<LUTType, 4> TotalLUT{};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto &LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1ULL;
}
return 0ULL;
// Expectation for this LUT is to simulate PSHUFLW with ARM's TBL (single register) instruction
// PSHUFLW behaviour:
// 16-bit words in [63:48], [47:32], [31:16], [15:0] are selected using the 8-bit Index.
// For 128-bit PSHUFLW, bits [127:64] are identity copied.
constexpr uint64_t IdentityCopyUpper = 0x0f'0e'0d'0c'0b'0a'09'08;
std::array<LUTType, 256> TotalLUT {};
uint64_t WordSelection[4] = {
0x01'00,
0x03'02,
0x05'04,
0x07'06,
};
for (size_t i = 0; i < 256; ++i) {
auto& LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
}
return TotalLUT;
}()
};
LUT.Val[0] = (WordSelection[Word0] << 0) | (WordSelection[Word1] << 16) | (WordSelection[Word2] << 32) | (WordSelection[Word3] << 48);
constexpr static auto PBLENDW_LUT {
[]() consteval {
struct LUTType {
uint16_t Val[8];
};
// 16-bit words in [127:112], [111:96], [95:80], [79:64], [63:48], [47:32], [31:16], [15:0] are selected using 8-bit swizzle.
// Expectation for this LUT is to simulate PBLENDW with ARM's TBX (one register) instruction.
// PBLENDW behaviour:
// 16-bit words from the source is moved in to the destination based on the bit in the swizzle.
// Dest[15:0] = Swizzle[0] ? Src[15:0] : Dest[15:0]
// Dest[31:16] = Swizzle[1] ? Src[31:16] : Dest[31:16]
// Dest[47:32] = Swizzle[2] ? Src[47:32] : Dest[47:32]
// Dest[63:48] = Swizzle[3] ? Src[63:48] : Dest[63:48]
// Dest[79:64] = Swizzle[4] ? Src[79:64] : Dest[79:64]
// Dest[95:80] = Swizzle[5] ? Src[95:80] : Dest[95:80]
// Dest[111:96] = Swizzle[6] ? Src[111:96] : Dest[111:96]
// Dest[127:112] = Swizzle[7] ? Src[127:112] : Dest[127:112]
std::array<LUTType, 256> TotalLUT{};
const uint16_t WordSelectionSrc[8] = {
0x01'00,
0x03'02,
0x05'04,
0x07'06,
0x09'08,
0x0B'0A,
0x0D'0C,
0x0F'0E,
};
constexpr uint16_t OriginalDest = 0xFF'FF;
for (size_t i = 0; i < 256; ++i) {
auto &LUT = TotalLUT[i];
for (size_t j = 0; j < 8; ++j) {
LUT.Val[j] = ((i >> j) & 1) ? WordSelectionSrc[j] : OriginalDest;
LUT.Val[1] = IdentityCopyUpper;
}
}
return TotalLUT;
}()
};
return TotalLUT;
}()};
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {
constexpr static auto PSHUFHW_LUT {[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// Expectation for this LUT is to simulate PSHUFHW with ARM's TBL (single register) instruction
// PSHUFHW behaviour:
// 16-bit words in [127:112], [111:96], [95:80], [79:64] are selected using the 8-bit Index.
// Incoming words come from bits [127:64] of the source.
// Bits [63:0] are identity copied.
constexpr uint64_t IdentityCopyLower = 0x07'06'05'04'03'02'01'00;
std::array<LUTType, 256> TotalLUT {};
uint64_t WordSelection[4] = {
0x09'08,
0x0b'0a,
0x0d'0c,
0x0f'0e,
};
for (size_t i = 0; i < 256; ++i) {
auto& LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
LUT.Val[0] = IdentityCopyLower;
// Initialize named vector constants.
for (size_t i = 0; i < FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX; ++i) {
Common.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
}
LUT.Val[1] = (WordSelection[Word0] << 0) | (WordSelection[Word1] << 16) | (WordSelection[Word2] << 32) | (WordSelection[Word3] << 48);
}
return TotalLUT;
}()};
// Copy named vector constants.
memcpy(Common.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
constexpr static auto PSHUFD_LUT {[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// Expectation for this LUT is to simulate PSHUFD with ARM's TBL (single register) instruction
// PSHUFD behaviour:
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
std::array<LUTType, 256> TotalLUT {};
uint64_t WordSelection[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
};
for (size_t i = 0; i < 256; ++i) {
auto& LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
// Initialize Indexed named vector constants.
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] = reinterpret_cast<uint64_t>(PSHUFLW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] = reinterpret_cast<uint64_t>(PSHUFHW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] = reinterpret_cast<uint64_t>(PSHUFD_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] = reinterpret_cast<uint64_t>(SHUFPS_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPS_MASK] = reinterpret_cast<uint64_t>(DPPS_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPD_MASK] = reinterpret_cast<uint64_t>(DPPD_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW] = reinterpret_cast<uint64_t>(PBLENDW_LUT.data());
LUT.Val[0] = (WordSelection[Word0] << 0) | (WordSelection[Word1] << 32);
LUT.Val[1] = (WordSelection[Word2] << 0) | (WordSelection[Word3] << 32);
}
return TotalLUT;
}()};
constexpr static auto SHUFPS_LUT {[]() consteval {
struct LUTType {
uint64_t Val[2];
};
// 32-bit words in [127:96], [95:64], [63:32], [31:0] are selected using the 8-bit Index.
// Expectation for this LUT is to simulate SHUFPS with ARM's TBL (two register) instruction.
// SHUFPS behaviour:
// Two 32-bits words from each source are selected from each source in the lower and upper halves of the 128-bit destination.
// Dest[31:0] = Src1[<Word0>]
// Dest[63:32] = Src1[<Word1>]
// Dest[95:64] = Src2[<Word2>]
// Dest[127:96] = Src2[<Word3>]
std::array<LUTType, 256> TotalLUT {};
const uint64_t WordSelectionSrc1[4] = {
0x03'02'01'00,
0x07'06'05'04,
0x0b'0a'09'08,
0x0f'0e'0d'0c,
};
// Src2 needs to offset each byte index by 16-bytes to pull from the second source.
const uint64_t WordSelectionSrc2[4] = {
0x03'02'01'00 + (0x10101010),
0x07'06'05'04 + (0x10101010),
0x0b'0a'09'08 + (0x10101010),
0x0f'0e'0d'0c + (0x10101010),
};
for (size_t i = 0; i < 256; ++i) {
auto& LUT = TotalLUT[i];
const auto Word0 = (i >> 0) & 0b11;
const auto Word1 = (i >> 2) & 0b11;
const auto Word2 = (i >> 4) & 0b11;
const auto Word3 = (i >> 6) & 0b11;
LUT.Val[0] = (WordSelectionSrc1[Word0] << 0) | (WordSelectionSrc1[Word1] << 32);
LUT.Val[1] = (WordSelectionSrc2[Word2] << 0) | (WordSelectionSrc2[Word3] << 32);
}
return TotalLUT;
}()};
constexpr static auto DPPS_MASK {[]() consteval {
struct LUTType {
uint32_t Val[4];
};
std::array<LUTType, 16> TotalLUT {};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto& LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1U;
}
return 0U;
};
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
LUT.Val[2] = GetLUT(i, 2);
LUT.Val[3] = GetLUT(i, 3);
}
return TotalLUT;
}()};
constexpr static auto DPPD_MASK {[]() consteval {
struct LUTType {
uint64_t Val[2];
};
std::array<LUTType, 4> TotalLUT {};
for (size_t i = 0; i < TotalLUT.size(); ++i) {
auto& LUT = TotalLUT[i];
constexpr auto GetLUT = [](size_t i, size_t Index) {
if (i & (1U << Index)) {
return -1ULL;
}
return 0ULL;
};
LUT.Val[0] = GetLUT(i, 0);
LUT.Val[1] = GetLUT(i, 1);
}
return TotalLUT;
}()};
constexpr static auto PBLENDW_LUT {[]() consteval {
struct LUTType {
uint16_t Val[8];
};
// 16-bit words in [127:112], [111:96], [95:80], [79:64], [63:48], [47:32], [31:16], [15:0] are selected using 8-bit swizzle.
// Expectation for this LUT is to simulate PBLENDW with ARM's TBX (one register) instruction.
// PBLENDW behaviour:
// 16-bit words from the source is moved in to the destination based on the bit in the swizzle.
// Dest[15:0] = Swizzle[0] ? Src[15:0] : Dest[15:0]
// Dest[31:16] = Swizzle[1] ? Src[31:16] : Dest[31:16]
// Dest[47:32] = Swizzle[2] ? Src[47:32] : Dest[47:32]
// Dest[63:48] = Swizzle[3] ? Src[63:48] : Dest[63:48]
// Dest[79:64] = Swizzle[4] ? Src[79:64] : Dest[79:64]
// Dest[95:80] = Swizzle[5] ? Src[95:80] : Dest[95:80]
// Dest[111:96] = Swizzle[6] ? Src[111:96] : Dest[111:96]
// Dest[127:112] = Swizzle[7] ? Src[127:112] : Dest[127:112]
std::array<LUTType, 256> TotalLUT {};
const uint16_t WordSelectionSrc[8] = {
0x01'00, 0x03'02, 0x05'04, 0x07'06, 0x09'08, 0x0B'0A, 0x0D'0C, 0x0F'0E,
};
constexpr uint16_t OriginalDest = 0xFF'FF;
for (size_t i = 0; i < 256; ++i) {
auto& LUT = TotalLUT[i];
for (size_t j = 0; j < 8; ++j) {
LUT.Val[j] = ((i >> j) & 1) ? WordSelectionSrc[j] : OriginalDest;
}
}
return TotalLUT;
}()};
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState* ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState)
, InitialCodeSize(InitialCodeSize)
, MaxCodeSize(MaxCodeSize) {
auto& Common = ThreadState->CurrentFrame->Pointers.Common;
// Initialize named vector constants.
for (size_t i = 0; i < FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_CONST_POOL_MAX; ++i) {
Common.NamedVectorConstantPointers[i] = reinterpret_cast<uint64_t>(NamedVectorConstants[i]);
}
// Copy named vector constants.
memcpy(Common.NamedVectorConstants, NamedVectorConstants, sizeof(NamedVectorConstants));
// Initialize Indexed named vector constants.
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFLW] =
reinterpret_cast<uint64_t>(PSHUFLW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFHW] =
reinterpret_cast<uint64_t>(PSHUFHW_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PSHUFD] =
reinterpret_cast<uint64_t>(PSHUFD_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_SHUFPS] =
reinterpret_cast<uint64_t>(SHUFPS_LUT.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPS_MASK] =
reinterpret_cast<uint64_t>(DPPS_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_DPPD_MASK] =
reinterpret_cast<uint64_t>(DPPD_MASK.data());
Common.IndexedNamedVectorConstantPointers[FEXCore::IR::IndexNamedVectorConstant::INDEXED_NAMED_VECTOR_PBLENDW] =
reinterpret_cast<uint64_t>(PBLENDW_LUT.data());
#ifndef FEX_DISABLE_TELEMETRY
// Fill in telemetry values
for (size_t i = 0; i < FEXCore::Telemetry::TYPE_LAST; ++i) {
auto &Telem = FEXCore::Telemetry::GetTelemetryValue(static_cast<FEXCore::Telemetry::TelemetryType>(i));
Common.TelemetryValueAddresses[i] = reinterpret_cast<uint64_t>(Telem.GetAddr());
}
// Fill in telemetry values
for (size_t i = 0; i < FEXCore::Telemetry::TYPE_LAST; ++i) {
auto& Telem = FEXCore::Telemetry::GetTelemetryValue(static_cast<FEXCore::Telemetry::TelemetryType>(i));
Common.TelemetryValueAddresses[i] = reinterpret_cast<uint64_t>(Telem.GetAddr());
}
#endif
}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer* {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
return CurrentCodeBuffer;
}
return CurrentCodeBuffer;
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
#ifndef _WIN32
// MDWE (Memory-Deny-Write-Execute) is a new Linux 6.3 feature.
// It's equivalent to systemd's `MemoryDenyWriteExecute` but implemented entirely in the kernel.
@@ -372,40 +351,39 @@ auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
#ifndef PR_GET_MDWE
#define PR_GET_MDWE 66
#endif
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
int MDWE = ::prctl(PR_GET_MDWE, 0, 0, 0, 0);
if (MDWE != -1 && MDWE != 0) {
LogMan::Msg::EFmt("MDWE was set to 0x{:x} which means FEX can't allocate executable memory", MDWE);
}
#endif
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t*>(FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto &Buffer: CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
return false;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
}
}
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto& Buffer : CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
}
}
return false;
}
} // namespace CPU
} // namespace FEXCore
@@ -20,14 +20,14 @@ namespace FEXCore {
namespace IR {
class IRListView;
class RegisterAllocationData;
}
} // namespace IR
namespace Core {
struct DebugData;
struct ThreadState;
struct CpuStateFrame;
struct InternalThreadState;
}
} // namespace Core
namespace CodeSerialize {
struct CodeObjectFileSection;
@@ -43,21 +43,22 @@ namespace CPU {
class CPUBackend {
public:
struct CodeBuffer {
uint8_t *Ptr;
uint8_t* Ptr;
size_t Size;
};
/**
* @param InitialCodeSize - Initial size for the code buffers
* @param MaxCodeSize - Max size for the code buffers
*/
CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize);
*/
CPUBackend(FEXCore::Core::InternalThreadState* ThreadState, size_t InitialCodeSize, size_t MaxCodeSize);
virtual ~CPUBackend();
/**
* @return The name of this backend
*/
[[nodiscard]] virtual fextl::string GetName() = 0;
[[nodiscard]]
virtual fextl::string GetName() = 0;
struct CompiledCode {
// Where this code block begins.
@@ -137,10 +138,9 @@ namespace CPU {
*
* @return Information about the compiled code block.
*/
[[nodiscard]] virtual CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) = 0;
[[nodiscard]]
virtual CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
@@ -150,14 +150,18 @@ namespace CPU {
*
* @return An executable function pointer relocated from the cache object
*/
[[nodiscard]] virtual void *RelocateJITObjectCode(uint64_t Entry, CodeSerialize::CodeObjectFileSection const *SerializationData) { return nullptr; }
[[nodiscard]]
virtual void* RelocateJITObjectCode(uint64_t Entry, const CodeSerialize::CodeObjectFileSection* SerializationData) {
return nullptr;
}
/**
* @brief Function for mapping memory in to the CPUBackend's visible space. Allows setting up virtual mappings if required
*
* @return Currently unused
*/
[[nodiscard]] virtual void *MapRegion(void *HostPtr, uint64_t GuestPtr, uint64_t Size) = 0;
[[nodiscard]]
virtual void* MapRegion(void* HostPtr, uint64_t GuestPtr, uint64_t Size) = 0;
/**
* @brief Lets FEXCore know if this CPUBackend needs IR and DebugData for CompileCode
@@ -168,7 +172,8 @@ namespace CPU {
*
* @return true if it needs the IR
*/
[[nodiscard]] virtual bool NeedsOpDispatch() = 0;
[[nodiscard]]
virtual bool NeedsOpDispatch() = 0;
virtual void ClearCache() {}
@@ -184,13 +189,14 @@ namespace CPU {
// to be able to handle a 256-bit vector store to a slot.
constexpr static uint32_t MaxSpillSlotSize = 32;
FEXCore::Core::InternalThreadState *ThreadState;
FEXCore::Core::InternalThreadState* ThreadState;
size_t InitialCodeSize, MaxCodeSize;
[[nodiscard]] CodeBuffer *GetEmptyCodeBuffer();
[[nodiscard]]
CodeBuffer* GetEmptyCodeBuffer();
// This is the current code buffer that we are tracking
CodeBuffer *CurrentCodeBuffer{};
CodeBuffer* CurrentCodeBuffer {};
private:
CodeBuffer AllocateNewCodeBuffer(size_t Size);
@@ -202,8 +208,8 @@ namespace CPU {
// This is the array of code buffers. Unless signals force us to keep more than
// buffer, there will be only one entry here
fextl::vector<CodeBuffer> CodeBuffers{};
fextl::vector<CodeBuffer> CodeBuffers {};
};
}
}
} // namespace CPU
} // namespace FEXCore
File diff suppressed because it is too large. Load diff
+81 -82
View File
@@ -30,7 +30,7 @@ private:
constexpr static uint32_t CPUID_VENDOR_AMD3 = 0x444D4163; // "cAMD"
public:
CPUIDEmu(FEXCore::Context::ContextImpl const *ctx);
CPUIDEmu(const FEXCore::Context::ContextImpl* ctx);
// X86 cacheline size effectively has to be hardcoded to 64
// if we report anything differently then applications are likely to break
@@ -58,12 +58,13 @@ public:
}
FEXCore::CPUID::FunctionResults RunFunctionName(uint32_t Function, uint32_t Leaf, uint32_t CPU) const {
if (Function == 0x8000'0002U)
if (Function == 0x8000'0002U) {
return Function_8000_0002h(Leaf, CPU % PerCPUData.size());
else if (Function == 0x8000'0003U)
} else if (Function == 0x8000'0003U) {
return Function_8000_0003h(Leaf, CPU % PerCPUData.size());
else
} else {
return Function_8000_0004h(Leaf, CPU % PerCPUData.size());
}
}
FEXCore::CPUID::XCRResults RunXCRFunction(uint32_t Function) const {
@@ -113,11 +114,12 @@ public:
}
private:
FEXCore::Context::ContextImpl const *CTX;
bool Hybrid{};
uint32_t Cores{};
const FEXCore::Context::ContextImpl* CTX;
bool Hybrid {};
uint32_t Cores {};
FEX_CONFIG_OPT(HideHypervisorBit, HIDEHYPERVISORBIT);
FEX_CONFIG_OPT(SmallTSCScale, SMALLTSCSCALE);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
// XFEATURE_ENABLED_MASK
// Mask that configures what features are enabled on the CPU.
@@ -148,10 +150,7 @@ private:
.SHA = 1,
};
uint64_t XCR0 {
XCR0_X87 |
XCR0_SSE
};
uint64_t XCR0 {XCR0_X87 | XCR0_SSE};
uint32_t SupportsAVX() const {
return (XCR0 & XCR0_AVX) ? 1 : 0;
@@ -160,13 +159,13 @@ private:
using FunctionHandler = FEXCore::CPUID::FunctionResults (CPUIDEmu::*)(uint32_t Leaf) const;
struct CPUData {
const char *ProductName{};
const char* ProductName {};
#ifdef _M_ARM_64
uint32_t MIDR{};
uint32_t MIDR {};
#endif
bool IsBig{};
bool IsBig {};
};
fextl::vector<CPUData> PerCPUData{};
fextl::vector<CPUData> PerCPUData {};
// Functions
FEXCore::CPUID::FunctionResults Function_0h(uint32_t Leaf) const;
@@ -277,74 +276,74 @@ private:
static constexpr std::array<FunctionConstant, PRIMARY_FUNCTION_COUNT> Primary_Constant = {{
// 0: Highest function parameter and ID
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 1: Processor info
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 2: Cache and TLB info
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 3: Serial Number(previously), now reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#ifndef CPUID_AMD
// 4: Deterministic cache parameters for each level
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 4: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
// 5: Monitor/mwait
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 6: Thermal and power management
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 7: Extended feature flags
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
// 0x08: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 9: Direct Cache Access information
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x0A: Architectural performance monitoring
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x0B: Extended topology enumeration
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x0C: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x0D: Processor extended state enumeration
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
// 0x0E: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x0F: Intel RDT monitoring
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x10: Intel RDT allocation enumeration
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x12: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x12: Intel SGX capability enumeration
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x13: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x14: Intel Processor trace
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#ifndef CPUID_AMD
// 0x15: Timestamp counter information
// Doesn't exist on AMD hardware
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#else
// 0x15: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
// 0x16: Processor frequency information
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x17: SoC vendor attribute enumeration
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x18: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x19: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#ifndef CPUID_AMD
// 0x1A: Hybrid Information Sub-leaf
{ SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#else
// 0x1A: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
}};
@@ -357,9 +356,9 @@ private:
static constexpr std::array<FunctionConstant, HYPERVISOR_FUNCTION_COUNT> Hypervisor_Constant = {{
// Hypervisor CPUID information leaf
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// FEX-Emu specific leaf
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
}};
static constexpr std::array<FunctionHandler, EXTENDED_FUNCTION_COUNT> Extended = {
@@ -439,79 +438,79 @@ private:
static constexpr std::array<FunctionConstant, EXTENDED_FUNCTION_COUNT> Extended_Constant = {{
// Largest extended function number
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// Processor vendor
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// Processor brand string
{ SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// Processor brand string continued
{ SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// Processor brand string continued
{ SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::NONCONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#ifdef CPUID_AMD
// 0x8000'0005: L1 Cache and TLB identifiers
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#else
// 0x8000'0005: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
// 0x8000'0006: L2 Cache identifiers
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0007: Advanced power management information
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0008: Virtual and physical address sizes
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0009: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000A: SVM Revision
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000B: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000C: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000D: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000E: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'000F: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0010: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0011: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0012: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0013: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0014: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0015: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0016: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0017: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0018: Reserved?
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'0019: TLB 1GB page identifiers
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'001A: Performance optimization identifiers
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'001B: Instruction based sampling identifiers
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'001C: Lightweight profiling capabilities
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#ifdef CPUID_AMD
// 0x8000'001D: Cache properties
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NEEDSLEAFCONSTANT},
#else
// 0x8000'001D: Reserved
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
#endif
// 0x8000'001E: Extended APIC ID
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
// 0x8000'001F: AMD Secure Encryption
{ SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT },
{SupportsConstant::CONSTANT, NeedsLeafConstant::NOLEAFCONSTANT},
}};
};
}
} // namespace FEXCore
File diff suppressed because it is too large. Load diff
@@ -1,6 +0,0 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <stdint.h>
namespace FEXCore::CPU {
}
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
@@ -17,6 +16,8 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <CodeEmitter/Emitter.h>
#include <atomic>
#include <condition_variable>
#include <csignal>
@@ -25,13 +26,13 @@
namespace FEXCore::CPU {
static void SleepThread(FEXCore::Context::ContextImpl *CTX, FEXCore::Core::CpuStateFrame *Frame) {
static void SleepThread(FEXCore::Context::ContextImpl* CTX, FEXCore::Core::CpuStateFrame* Frame) {
CTX->SyscallHandler->SleepThread(CTX, Frame);
}
constexpr size_t MAX_DISPATCHER_CODE_SIZE = 4096 * 2;
Dispatcher::Dispatcher(FEXCore::Context::ContextImpl *ctx)
Dispatcher::Dispatcher(FEXCore::Context::ContextImpl* ctx)
: Arm64Emitter(ctx, FEXCore::Allocator::VirtualAlloc(MAX_DISPATCHER_CODE_SIZE, true), MAX_DISPATCHER_CODE_SIZE)
, CTX {ctx} {
EmitDispatcher();
@@ -61,6 +62,14 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::SingleUseForwardLabel l_Sleep;
#ifdef _M_ARM_64EC
// These structures are not included in the standard Windows headers, define them here
static constexpr size_t TEBCPUAreaOffset = 0x1788;
static constexpr size_t CPUAreaInSyscallCallbackOffset = 0x1;
static constexpr size_t CPUAreaEmulatorStackLimitOffset = 0x8;
static constexpr size_t CPUAreaEmulatorDataOffset = 0x30;
ARMEmitter::SingleUseForwardLabel ExitEC;
#endif
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
// Push all the register we need to save
@@ -79,19 +88,45 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
FillStaticRegs();
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
FillStaticRegs();
// Enter JIT
b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
str(EC_CALL_CHECKER_PC_REG, STATE_PTR(CpuStateFrame, State.rip));
// Swap stacks to the emulator stack
ldr(TMP1, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorStackLimitOffset);
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
if (EmitterCTX->HostFeatures.SupportsSVE) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
// Enter JIT
#endif
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
ARMEmitter::BiDirectionalLabel FullLookup{};
ARMEmitter::BiDirectionalLabel CallBlock{};
ARMEmitter::BackwardLabel LoopTop{};
ARMEmitter::BiDirectionalLabel FullLookup {};
ARMEmitter::BiDirectionalLabel CallBlock {};
Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
auto RipReg = TMP3;
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
@@ -99,7 +134,7 @@ void Dispatcher::EmitDispatcher() {
ldr(TMP1, STATE_PTR(CpuStateFrame, Pointers.Common.L1Pointer));
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL , 4);
add(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP4, ARMEmitter::ShiftType::LSL, 4);
ldp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP1, TMP1, 0);
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, RipReg);
cbnz(ARMEmitter::Size::i64Bit, TMP1, &FullLookup);
@@ -117,8 +152,7 @@ void Dispatcher::EmitDispatcher() {
uint64_t VirtualMemorySize = CTX->Config.VirtualMemSize;
if (std::popcount(VirtualMemorySize) == 1) {
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), VirtualMemorySize - 1);
}
else {
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, VirtualMemorySize);
and_(ARMEmitter::Size::i64Bit, TMP4, RipReg.R(), TMP4);
}
@@ -134,6 +168,10 @@ void Dispatcher::EmitDispatcher() {
// If page pointer is zero then we have no block
cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
#ifdef _M_ARM_64EC
// The LSB of an L2 page entry indicates if this page contains EC code
tbnz(TMP1, 0, &ExitEC);
#endif
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
@@ -167,6 +205,16 @@ void Dispatcher::EmitDispatcher() {
}
}
#ifdef _M_ARM_64EC
{
Bind(&ExitEC);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
}
#endif
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
@@ -188,14 +236,19 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
#endif
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionLink));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void *, void *>(ARMEmitter::Reg::r2);
}
else {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
@@ -205,13 +258,18 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
strb(ARMEmitter::WReg::zr, TMP2, CPUAreaInSyscallCallbackOffset);
#endif
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP2, 0);
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
br(TMP1);
}
@@ -230,6 +288,12 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
#endif
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
// x2 contains guest RIP
@@ -237,21 +301,25 @@ void Dispatcher::EmitDispatcher() {
ldr(ARMEmitter::XReg::x4, &l_CompileBlock);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uintptr_t, void *, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
}
else {
GenerateIndirectRuntimeCall<uintptr_t, void*, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r4);
} else {
blr(ARMEmitter::Reg::r4); // { CTX, Frame, RIP, MaxInst }
}
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP1, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
strb(ARMEmitter::WReg::zr, TMP1, CPUAreaInSyscallCallbackOffset);
#endif
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP1, 0);
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
b(&LoopTop);
}
@@ -285,7 +353,7 @@ void Dispatcher::EmitDispatcher() {
{
// Guest SIGTRAP handler
// Needs to be distinct from the SignalHandlerReturnAddress
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
GuestSignal_SIGTRAP = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
@@ -308,8 +376,7 @@ void Dispatcher::EmitDispatcher() {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::r0, 0);
PopCalleeSavedRegisters();
ret();
}
else {
} else {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, 0);
ldr(ARMEmitter::XReg::x1, ARMEmitter::Reg::r1);
}
@@ -328,9 +395,8 @@ void Dispatcher::EmitDispatcher() {
mov(ARMEmitter::XReg::x1, STATE);
ldr(ARMEmitter::XReg::x2, &l_Sleep);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void *, void *>(ARMEmitter::Reg::r2);
}
else {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
@@ -405,8 +471,7 @@ void Dispatcher::EmitDispatcher() {
ldr(ARMEmitter::XReg::x3, R, Offset);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
} else {
blr(ARMEmitter::Reg::r3);
}
// Result is now in x0
@@ -463,26 +528,28 @@ void Dispatcher::EmitDispatcher() {
}
#ifdef VIXL_SIMULATOR
void Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
void Dispatcher::ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.RunFrom(reinterpret_cast<vixl::aarch64::Instruction const*>(DispatchPtr));
Simulator.RunFrom(reinterpret_cast< const vixl::aarch64::Instruction*>(DispatchPtr));
}
void Dispatcher::ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
void Dispatcher::ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
Simulator.WriteXRegister(0, reinterpret_cast<int64_t>(Frame));
Simulator.WriteXRegister(1, RIP);
Simulator.RunFrom(reinterpret_cast<vixl::aarch64::Instruction const*>(CallbackPtr));
Simulator.RunFrom(reinterpret_cast< const vixl::aarch64::Instruction*>(CallbackPtr));
}
#endif
void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread) {
void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread) {
// Setup dispatcher specific pointers that need to be accessed from JIT code
{
auto &Common = Thread->CurrentFrame->Pointers.Common;
auto& Common = Thread->CurrentFrame->Pointers.Common;
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.DispatcherLoopTopEnterEC = AbsoluteLoopTopAddressEnterEC;
Common.DispatcherLoopTopEnterECFillSRA = AbsoluteLoopTopAddressEnterECFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
@@ -492,7 +559,7 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread)
Common.SignalReturnHandler = SignalHandlerReturnAddress;
Common.SignalReturnHandlerRT = SignalHandlerReturnAddressRT;
auto &AArch64 = Thread->CurrentFrame->Pointers.AArch64;
auto& AArch64 = Thread->CurrentFrame->Pointers.AArch64;
AArch64.LUDIVHandler = LUDIVHandlerAddress;
AArch64.LDIVHandler = LDIVHandlerAddress;
AArch64.LUREMHandler = LUREMHandlerAddress;
@@ -500,8 +567,8 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState *Thread)
}
}
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl *CTX) {
fextl::unique_ptr<Dispatcher> Dispatcher::Create(FEXCore::Context::ContextImpl* CTX) {
return fextl::make_unique<Dispatcher>(CTX);
}
}
} // namespace FEXCore::CPU
@@ -2,8 +2,8 @@
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/fextl/memory.h>
#ifdef VIXL_SIMULATOR
@@ -23,7 +23,7 @@ struct GuestSigAction;
namespace FEXCore::Core {
struct CpuStateFrame;
struct InternalThreadState;
}
} // namespace FEXCore::Core
namespace FEXCore::Context {
class ContextImpl;
@@ -31,51 +31,52 @@ class ContextImpl;
namespace FEXCore::CPU {
#define STATE_PTR(STATE_TYPE, FIELD) \
STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
#define STATE_PTR(STATE_TYPE, FIELD) STATE.R(), offsetof(FEXCore::Core::STATE_TYPE, FIELD)
class Dispatcher final : public Arm64Emitter {
public:
static fextl::unique_ptr<Dispatcher> Create(FEXCore::Context::ContextImpl *CTX);
static fextl::unique_ptr<Dispatcher> Create(FEXCore::Context::ContextImpl* CTX);
Dispatcher(FEXCore::Context::ContextImpl *ctx);
Dispatcher(FEXCore::Context::ContextImpl* ctx);
~Dispatcher();
/**
* @name Dispatch Helper functions
* @{ */
uint64_t ThreadStopHandlerAddress{};
uint64_t ThreadStopHandlerAddressSpillSRA{};
uint64_t AbsoluteLoopTopAddress{};
uint64_t AbsoluteLoopTopAddressFillSRA{};
uint64_t ThreadPauseHandlerAddress{};
uint64_t ThreadPauseHandlerAddressSpillSRA{};
uint64_t ExitFunctionLinkerAddress{};
uint64_t SignalHandlerReturnAddress{};
uint64_t SignalHandlerReturnAddressRT{};
uint64_t GuestSignal_SIGILL{};
uint64_t GuestSignal_SIGTRAP{};
uint64_t GuestSignal_SIGSEGV{};
uint64_t IntCallbackReturnAddress{};
uint64_t ThreadStopHandlerAddress {};
uint64_t ThreadStopHandlerAddressSpillSRA {};
uint64_t AbsoluteLoopTopAddress {};
uint64_t AbsoluteLoopTopAddressFillSRA {};
uint64_t AbsoluteLoopTopAddressEnterEC {};
uint64_t AbsoluteLoopTopAddressEnterECFillSRA {};
uint64_t ThreadPauseHandlerAddress {};
uint64_t ThreadPauseHandlerAddressSpillSRA {};
uint64_t ExitFunctionLinkerAddress {};
uint64_t SignalHandlerReturnAddress {};
uint64_t SignalHandlerReturnAddressRT {};
uint64_t GuestSignal_SIGILL {};
uint64_t GuestSignal_SIGTRAP {};
uint64_t GuestSignal_SIGSEGV {};
uint64_t IntCallbackReturnAddress {};
uint64_t PauseReturnInstruction{};
uint64_t PauseReturnInstruction {};
/** @} */
uint64_t Start{};
uint64_t End{};
uint64_t Start {};
uint64_t End {};
void InitThreadPointers(FEXCore::Core::InternalThreadState *Thread);
void InitThreadPointers(FEXCore::Core::InternalThreadState* Thread);
#ifdef VIXL_SIMULATOR
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) ;
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame);
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
#else
void ExecuteDispatch(FEXCore::Core::CpuStateFrame *Frame) {
void ExecuteDispatch(FEXCore::Core::CpuStateFrame* Frame) {
DispatchPtr(Frame);
}
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP) {
void ExecuteJITCallback(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP) {
CallbackPtr(Frame, RIP);
}
#endif
@@ -103,21 +104,21 @@ public:
}
protected:
FEXCore::Context::ContextImpl *CTX;
FEXCore::Context::ContextImpl* CTX;
using AsmDispatch = void(*)(FEXCore::Core::CpuStateFrame *Frame);
using JITCallback = void(*)(FEXCore::Core::CpuStateFrame *Frame, uint64_t RIP);
using AsmDispatch = void (*)(FEXCore::Core::CpuStateFrame* Frame);
using JITCallback = void (*)(FEXCore::Core::CpuStateFrame* Frame, uint64_t RIP);
AsmDispatch DispatchPtr;
JITCallback CallbackPtr;
private:
// Long division helpers
uint64_t LUDIVHandlerAddress{};
uint64_t LDIVHandlerAddress{};
uint64_t LUREMHandlerAddress{};
uint64_t LREMHandlerAddress{};
uint64_t LUDIVHandlerAddress {};
uint64_t LDIVHandlerAddress {};
uint64_t LUREMHandlerAddress {};
uint64_t LREMHandlerAddress {};
void EmitDispatcher();
};
}
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
+31 -24
View File
@@ -21,10 +21,10 @@ class Decoder final {
public:
// New Frontend decoding
struct DecodedBlocks final {
uint64_t Entry{};
uint64_t NumInstructions{};
FEXCore::X86Tables::DecodedInst *DecodedInstructions;
bool HasInvalidInstruction{};
uint64_t Entry {};
uint64_t NumInstructions {};
FEXCore::X86Tables::DecodedInst* DecodedInstructions;
bool HasInvalidInstruction {};
};
struct DecodedBlockInformation final {
@@ -32,19 +32,24 @@ public:
fextl::vector<DecodedBlocks> Blocks;
};
Decoder(FEXCore::Context::ContextImpl *ctx);
Decoder(FEXCore::Context::ContextImpl* ctx);
~Decoder();
void DecodeInstructionsAtEntry(uint8_t const* InstStream, uint64_t PC, uint64_t MaxInst, std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
void DecodeInstructionsAtEntry(const uint8_t* InstStream, uint64_t PC, uint64_t MaxInst,
std::function<void(uint64_t BlockEntry, uint64_t Start, uint64_t Length)> AddContainedCodePage);
DecodedBlockInformation const *GetDecodedBlockInfo() const {
const DecodedBlockInformation* GetDecodedBlockInfo() const {
return &BlockInfo;
}
uint64_t DecodedMinAddress {};
uint64_t DecodedMaxAddress {~0ULL};
void SetSectionMaxAddress(uint64_t v) { SectionMaxAddress = v; }
void SetExternalBranches(fextl::set<uint64_t> *v) { ExternalBranches = v; }
void SetSectionMaxAddress(uint64_t v) {
SectionMaxAddress = v;
}
void SetExternalBranches(fextl::set<uint64_t>* v) {
ExternalBranches = v;
}
void DelayedDisownBuffer() {
PoolObject.DelayedDisownBuffer();
@@ -59,8 +64,8 @@ private:
bool L; // VEX.L bit (if set then 256 bit operation, if unset then scalar or 128-bit operation)
};
FEXCore::Context::ContextImpl *CTX;
const FEXCore::HLE::SyscallOSABI OSABI{};
FEXCore::Context::ContextImpl* CTX;
const FEXCore::HLE::SyscallOSABI OSABI {};
bool DecodeInstruction(uint64_t PC);
@@ -70,22 +75,24 @@ private:
uint8_t ReadByte();
uint8_t PeekByte(uint8_t Offset) const;
uint64_t ReadData(uint8_t Size);
void SkipBytes(uint8_t Size) { InstructionSize += Size; }
void SkipBytes(uint8_t Size) {
InstructionSize += Size;
}
bool NormalOp(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(FEXCore::X86Tables::X86InstInfo const *Info, uint16_t Op);
bool NormalOp(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op, DecodedHeader Options = {});
bool NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16_t Op);
static constexpr size_t DefaultDecodedBufferSize = 0x10000;
FEXCore::X86Tables::DecodedInst *DecodedBuffer{};
FEXCore::X86Tables::DecodedInst* DecodedBuffer {};
Utils::FixedSizePooledAllocation<FEXCore::X86Tables::DecodedInst*, 5000, 500> PoolObject;
size_t DecodedSize {};
uint8_t const *InstStream;
const uint8_t* InstStream;
static constexpr size_t MAX_INST_SIZE = 15;
uint8_t InstructionSize;
std::array<uint8_t, MAX_INST_SIZE> Instruction;
FEXCore::X86Tables::DecodedInst *DecodeInst;
FEXCore::X86Tables::DecodedInst* DecodeInst;
// This is for multiblock data tracking
bool SymbolAvailable {false};
@@ -99,21 +106,21 @@ private:
DecodedBlockInformation BlockInfo;
fextl::set<uint64_t> BlocksToDecode;
fextl::set<uint64_t> HasBlocks;
fextl::set<uint64_t> *ExternalBranches {nullptr};
fextl::set<uint64_t>* ExternalBranches {nullptr};
// ModRM rm decoding
using DecodeModRMPtr = void (FEXCore::Frontend::Decoder::*)(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_16(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_64(X86Tables::DecodedOperand *Operand, X86Tables::ModRMDecoded ModRM);
using DecodeModRMPtr = void (FEXCore::Frontend::Decoder::*)(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_16(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM);
void DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModRMDecoded ModRM);
static constexpr std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp{
static constexpr std::array<DecodeModRMPtr, 2> DecodeModRMs_Disp {
&FEXCore::Frontend::Decoder::DecodeModRM_64,
&FEXCore::Frontend::Decoder::DecodeModRM_16,
};
const uint8_t *AdjustAddrForSpecialRegion(uint8_t const* _InstStream, uint64_t EntryPoint, uint64_t RIP);
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
FEXCORE_TELEMETRY_INIT(VEXOpTelem, TYPE_USES_VEX_OPS);
FEXCORE_TELEMETRY_INIT(EVEXOpTelem, TYPE_USES_EVEX_OPS);
};
}
} // namespace FEXCore::Frontend
+33 -45
View File
@@ -28,23 +28,21 @@ namespace FEXCore {
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
[[maybe_unused]] static uint32_t GetDCZID() {
uint64_t Result{};
__asm("mrs %[Res], DCZID_EL0"
: [Res] "=r" (Result));
[[maybe_unused]]
static uint32_t GetDCZID() {
uint64_t Result {};
__asm("mrs %[Res], DCZID_EL0" : [Res] "=r"(Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result{};
__asm ("mrs %[Res], FPCR"
: [Res] "=r" (Result));
uint64_t Result {};
__asm("mrs %[Res], FPCR" : [Res] "=r"(Result));
return Result;
}
static void SetFPCR(uint64_t Value) {
__asm ("msr FPCR, %[Value]"
:: [Value] "r" (Value));
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
}
#else
static uint32_t GetDCZID() {
@@ -53,7 +51,7 @@ static uint32_t GetDCZID() {
}
#endif
static void OverrideFeatures(HostFeatures *Features) {
static void OverrideFeatures(HostFeatures* Features) {
// Override features if the user has specifically called for it.
FEX_CONFIG_OPT(HostFeatures, HOSTFEATURES);
if (!HostFeatures()) {
@@ -61,20 +59,20 @@ static void OverrideFeatures(HostFeatures *Features) {
return;
}
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features->FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features->FeatureName = Result; \
} while (0)
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features->FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features->FeatureName = Result; \
} while (0)
#define GET_SINGLE_OPTION(name, enum_name) \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive");
#define GET_SINGLE_OPTION(name, enum_name) \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive");
ENABLE_DISABLE_OPTION(SupportsAVX, AVX, AVX);
ENABLE_DISABLE_OPTION(SupportsAVX2, AVX2, AVX2);
@@ -102,8 +100,7 @@ static void OverrideFeatures(HostFeatures *Features) {
Features->SupportsCRC = true;
Features->SupportsSHA = true;
Features->SupportsPMULL_128Bit = true;
}
else if (DisableCrypto) {
} else if (DisableCrypto) {
Features->SupportsAES = false;
Features->SupportsCRC = false;
Features->SupportsSHA = false;
@@ -127,8 +124,7 @@ HostFeatures::HostFeatures() {
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsSHA = Features.Has(vixl::CPUFeatures::Feature::kSHA1) &&
Features.Has(vixl::CPUFeatures::Feature::kSHA2);
SupportsSHA = Features.Has(vixl::CPUFeatures::Feature::kSHA1) && Features.Has(vixl::CPUFeatures::Feature::kSHA2);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
SupportsRAND = Features.Has(vixl::CPUFeatures::Feature::kRNG);
@@ -151,8 +147,7 @@ HostFeatures::HostFeatures() {
SupportsAVX = true;
#else
SupportsSVE = Features.Has(vixl::CPUFeatures::Feature::kSVE);
SupportsAVX = Features.Has(vixl::CPUFeatures::Feature::kSVE2) &&
vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
SupportsAVX = Features.Has(vixl::CPUFeatures::Feature::kSVE2) && vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
#endif
// TODO: AVX2 is currently unsupported. Disable until the remaining features are implemented.
SupportsAVX2 = false;
@@ -160,11 +155,6 @@ HostFeatures::HostFeatures() {
SupportsBMI2 = true;
SupportsCLWB = true;
// TODO: AFP is disabled until the scalar usage in the codebase can be audited to be working as expected.
SupportsAFP = false;
// RPRES has a dependency on AFP. Disable it until AFP is enabled.
SupportsRPRES = false;
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
@@ -173,21 +163,19 @@ HostFeatures::HostFeatures() {
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile ("mrs %[ctr], ctr_el0"
: [ctr] "=r"(CTR));
__asm volatile("mrs %[ctr], ctr_el0" : [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
constexpr uint32_t ExceptionEnableTraps =
(1U << 8) | // Invalid Operation float exception trap enable
(1U << 9) | // Divide by zero float exception trap enable
(1U << 10) | // Overflow float exception trap enable
(1U << 11) | // Underflow float exception trap enable
(1U << 12) | // Inexact float exception trap enable
(1U << 15); // Input Denormal float exception trap enable
constexpr uint32_t ExceptionEnableTraps = (1U << 8) | // Invalid Operation float exception trap enable
(1U << 9) | // Divide by zero float exception trap enable
(1U << 10) | // Overflow float exception trap enable
(1U << 11) | // Underflow float exception trap enable
(1U << 12) | // Inexact float exception trap enable
(1U << 15); // Input Denormal float exception trap enable
uint32_t OriginalFPCR = GetFPCR();
uint32_t FPCR = OriginalFPCR | ExceptionEnableTraps;
@@ -222,7 +210,7 @@ HostFeatures::HostFeatures() {
ICacheLineSize = 64U;
#if !defined(VIXL_SIMULATOR)
Xbyak::util::Cpu X86Features{};
Xbyak::util::Cpu X86Features {};
SupportsAES = X86Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = X86Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = X86Features.has(Xbyak::util::Cpu::tRDRAND) && X86Features.has(Xbyak::util::Cpu::tRDSEED);
@@ -256,4 +244,4 @@ HostFeatures::HostFeatures() {
SupportsPreserveAllABI = FEXCORE_HAS_PRESERVE_ALL_ATTR;
OverrideFeatures(this);
}
}
} // namespace FEXCore
@@ -2,48 +2,36 @@
#include "Common/SoftFloat.h"
#include "Common/SoftFloat-3e/softfloat.h"
#include <FEXCore/IR/IR.h>
#include "Interface/Core/Interpreter/Fallbacks/FallbackOpHandler.h"
#include "Interface/IR/IR.h"
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR
static void LoadDeferredFCW(uint16_t NewFCW) {
FEXCORE_PRESERVE_ALL_ATTR static void LoadDeferredFCW(uint16_t NewFCW) {
auto PC = (NewFCW >> 8) & 3;
switch(PC) {
case 0: extF80_roundingPrecision = 32; break;
case 2: extF80_roundingPrecision = 64; break;
case 3: extF80_roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
switch (PC) {
case 0: extF80_roundingPrecision = 32; break;
case 2: extF80_roundingPrecision = 64; break;
case 3: extF80_roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
}
auto RC = (NewFCW >> 10) & 3;
switch(RC) {
case 0:
softfloat_roundingMode = softfloat_round_near_even;
break;
case 1:
softfloat_roundingMode = softfloat_round_min;
break;
case 2:
softfloat_roundingMode = softfloat_round_max;
break;
case 3:
softfloat_roundingMode = softfloat_round_minMag;
break;
switch (RC) {
case 0: softfloat_roundingMode = softfloat_round_near_even; break;
case 1: softfloat_roundingMode = softfloat_round_min; break;
case 2: softfloat_roundingMode = softfloat_round_max; break;
case 3: softfloat_roundingMode = softfloat_round_minMag; break;
}
}
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle4(uint16_t NewFCW, float src) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, float src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle8(uint16_t NewFCW, double src) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
return src;
}
@@ -52,24 +40,20 @@ struct OpHandlers<IR::OP_F80CVTTO> {
template<>
struct OpHandlers<IR::OP_F80CMP> {
template<uint32_t Flags>
FEXCORE_PRESERVE_ALL_ATTR
static uint64_t handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Flags & (1 << IR::FCMP_FLAG_LT) &&
lt) {
if (Flags & (1 << IR::FCMP_FLAG_LT) && lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
if (Flags & (1 << IR::FCMP_FLAG_UNORDERED) &&
nan) {
if (Flags & (1 << IR::FCMP_FLAG_UNORDERED) && nan) {
ResultFlags |= (1 << IR::FCMP_FLAG_UNORDERED);
}
if (Flags & (1 << IR::FCMP_FLAG_EQ) &&
eq) {
if (Flags & (1 << IR::FCMP_FLAG_EQ) && eq) {
ResultFlags |= (1 << IR::FCMP_FLAG_EQ);
}
return ResultFlags;
@@ -78,14 +62,12 @@ struct OpHandlers<IR::OP_F80CMP> {
template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR
static float handle4(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static double handle8(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
}
@@ -93,26 +75,22 @@ struct OpHandlers<IR::OP_F80CVT> {
template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR
static int16_t handle2(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static int32_t handle4(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static int64_t handle8(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static int16_t handle2t(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
auto rv = extF80_to_i32(src, softfloat_round_minMag, false);
@@ -125,14 +103,12 @@ struct OpHandlers<IR::OP_F80CVTINT> {
}
}
FEXCORE_PRESERVE_ALL_ATTR
static int32_t handle4t(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i32(src, softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR
static int64_t handle8t(uint16_t NewFCW, X80SoftFloat src) {
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i64(src, softfloat_round_minMag, false);
}
@@ -140,14 +116,12 @@ struct OpHandlers<IR::OP_F80CVTINT> {
template<>
struct OpHandlers<IR::OP_F80CVTTOINT> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle2(uint16_t NewFCW, int16_t src) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t NewFCW, int16_t src) {
LoadDeferredFCW(NewFCW);
return src;
}
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle4(uint16_t NewFCW, int32_t src) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, int32_t src) {
LoadDeferredFCW(NewFCW);
return src;
}
@@ -155,8 +129,7 @@ struct OpHandlers<IR::OP_F80CVTTOINT> {
template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FRNDINT(Src1);
}
@@ -164,8 +137,7 @@ struct OpHandlers<IR::OP_F80ROUND> {
template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::F2XM1(Src1);
}
@@ -173,8 +145,7 @@ struct OpHandlers<IR::OP_F80F2XM1> {
template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FTAN(Src1);
}
@@ -182,8 +153,7 @@ struct OpHandlers<IR::OP_F80TAN> {
template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSQRT(Src1);
}
@@ -191,8 +161,7 @@ struct OpHandlers<IR::OP_F80SQRT> {
template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSIN(Src1);
}
@@ -200,8 +169,7 @@ struct OpHandlers<IR::OP_F80SIN> {
template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FCOS(Src1);
}
@@ -209,8 +177,7 @@ struct OpHandlers<IR::OP_F80COS> {
template<>
struct OpHandlers<IR::OP_F80XTRACT_EXP> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FXTRACT_EXP(Src1);
}
@@ -218,8 +185,7 @@ struct OpHandlers<IR::OP_F80XTRACT_EXP> {
template<>
struct OpHandlers<IR::OP_F80XTRACT_SIG> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FXTRACT_SIG(Src1);
}
@@ -227,8 +193,7 @@ struct OpHandlers<IR::OP_F80XTRACT_SIG> {
template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FADD(Src1, Src2);
}
@@ -236,8 +201,7 @@ struct OpHandlers<IR::OP_F80ADD> {
template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSUB(Src1, Src2);
}
@@ -245,8 +209,7 @@ struct OpHandlers<IR::OP_F80SUB> {
template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FMUL(Src1, Src2);
}
@@ -254,8 +217,7 @@ struct OpHandlers<IR::OP_F80MUL> {
template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FDIV(Src1, Src2);
}
@@ -263,8 +225,7 @@ struct OpHandlers<IR::OP_F80DIV> {
template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FYL2X(Src1, Src2);
}
@@ -272,8 +233,7 @@ struct OpHandlers<IR::OP_F80FYL2X> {
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FATAN(Src1, Src2);
}
@@ -281,8 +241,7 @@ struct OpHandlers<IR::OP_F80ATAN> {
template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM1(Src1, Src2);
}
@@ -290,8 +249,7 @@ struct OpHandlers<IR::OP_F80FPREM1> {
template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM(Src1, Src2);
}
@@ -299,8 +257,7 @@ struct OpHandlers<IR::OP_F80FPREM> {
template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSCALE(Src1, Src2);
}
@@ -374,15 +331,14 @@ template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
double trunc = (double)(int64_t)(src2); //truncate
double trunc = (double)(int64_t)(src2); // truncate
return src1 * exp2(trunc);
}
};
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
bool Negative = Src1.Sign;
@@ -393,7 +349,7 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
uint64_t Tmp = Src1;
X80SoftFloat Rv;
uint8_t *BCD = reinterpret_cast<uint8_t*>(&Rv);
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
for (size_t i = 0; i < 9; ++i) {
@@ -423,11 +379,10 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
template<>
struct OpHandlers<IR::OP_F80BCDLOAD> {
FEXCORE_PRESERVE_ALL_ATTR
static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src) {
LoadDeferredFCW(NewFCW);
uint8_t *Src1 = reinterpret_cast<uint8_t *>(&Src);
uint64_t BCD{};
uint8_t* Src1 = reinterpret_cast<uint8_t*>(&Src);
uint64_t BCD {};
// We walk through each uint8_t and pull out the BCD encoding
// Each 4bit split is a digit
// Only 0-9 is supported, A-F results in undefined data
@@ -68,8 +68,7 @@ namespace FEXCore::CPU {
//
// 5. Done.
//
template <IR::IROps Op>
struct OpHandlers {
};
template<IR::IROps Op>
struct OpHandlers {};
} // namespace FEXCore::CPU
@@ -10,23 +10,23 @@
namespace FEXCore::CPU {
template<typename R, typename... Args>
static FallbackInfo GetFallbackInfo(R(*fn)(Args...), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
static FallbackInfo GetFallbackInfo(R (*fn)(Args...), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_UNKNOWN, (void*)fn, HandlerIndex, false};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(uint16_t, double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
FallbackInfo GetFallbackInfo(double (*fn)(uint16_t, double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64, (void*)fn, HandlerIndex, false};
}
template<>
FallbackInfo GetFallbackInfo(double(*fn)(uint16_t, double,double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
FallbackInfo GetFallbackInfo(double (*fn)(uint16_t, double, double), FEXCore::Core::FallbackHandlerIndex HandlerIndex) {
return {FABI_F64_I16_F64_F64, (void*)fn, HandlerIndex, false};
}
void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80CVTTO_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
Info[Core::OPINDEX_F80CVTTO_8] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
void InterpreterOps::FillFallbackIndexPointers(uint64_t* Info) {
Info[Core::OPINDEX_F80CVTTO_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4);
Info[Core::OPINDEX_F80CVTTO_8] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8);
Info[Core::OPINDEX_F80CVT_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4);
Info[Core::OPINDEX_F80CVT_8] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8);
Info[Core::OPINDEX_F80CVTINT_2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2);
@@ -55,8 +55,8 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_F80COS] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80COS>::handle);
Info[Core::OPINDEX_F80XTRACT_EXP] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_EXP>::handle);
Info[Core::OPINDEX_F80XTRACT_SIG] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80XTRACT_SIG>::handle);
Info[Core::OPINDEX_F80BCDSTORE] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDSTORE>::handle);
Info[Core::OPINDEX_F80BCDLOAD] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDLOAD>::handle);
Info[Core::OPINDEX_F80BCDSTORE] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDSTORE>::handle);
Info[Core::OPINDEX_F80BCDLOAD] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80BCDLOAD>::handle);
// Binary
Info[Core::OPINDEX_F80ADD] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80ADD>::handle);
@@ -85,126 +85,123 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t *Info) {
Info[Core::OPINDEX_VPCMPISTRX] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle);
}
bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, IR::IROp_Header const *IROp, FallbackInfo *Info) {
bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::IROp_Header* IROp, FallbackInfo* Info) {
uint8_t OpSize = IROp->Size;
switch(IROp->Op) {
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (IROp->Op) {
case IR::OP_F80CVTTO: {
auto Op = IROp->C<IR::IROp_F80CVTTo>();
switch (Op->SrcSize) {
case 4: {
*Info = {FABI_F80_I16_F32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F80_I16_F64, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = {FABI_F32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
switch (OpSize) {
case 2: {
if (Op->Truncate) {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2, SupportsPreserveAllABI};
}
return true;
}
case 4: {
if (Op->Truncate) {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4, SupportsPreserveAllABI};
}
return true;
}
case 8: {
if (Op->Truncate) {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8, SupportsPreserveAllABI};
}
else {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8, SupportsPreserveAllABI};
}
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers{
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags), SupportsPreserveAllABI};
switch (Op->SrcSize) {
case 4: {
*Info = {FABI_F80_I16_F32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle4, Core::OPINDEX_F80CVTTO_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F80_I16_F64, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTO>::handle8, Core::OPINDEX_F80CVTTO_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVT: {
switch (OpSize) {
case 4: {
*Info = {FABI_F32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle4, Core::OPINDEX_F80CVT_4, SupportsPreserveAllABI};
return true;
}
case 8: {
*Info = {FABI_F64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVT>::handle8, Core::OPINDEX_F80CVT_8, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CVTINT: {
auto Op = IROp->C<IR::IROp_F80CVTInt>();
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->SrcSize) {
case 2: {
*Info = {FABI_F80_I16_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2, SupportsPreserveAllABI};
return true;
}
case 4: {
*Info = {FABI_F80_I16_I32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
switch (OpSize) {
case 2: {
if (Op->Truncate) {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t, Core::OPINDEX_F80CVTINT_TRUNC2,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I16_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2, Core::OPINDEX_F80CVTINT_2, SupportsPreserveAllABI};
}
break;
return true;
}
case 4: {
if (Op->Truncate) {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t, Core::OPINDEX_F80CVTINT_TRUNC4,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I32_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4, Core::OPINDEX_F80CVTINT_4, SupportsPreserveAllABI};
}
return true;
}
case 8: {
if (Op->Truncate) {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t, Core::OPINDEX_F80CVTINT_TRUNC8,
SupportsPreserveAllABI};
} else {
*Info = {FABI_I64_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8, Core::OPINDEX_F80CVTINT_8, SupportsPreserveAllABI};
}
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
static constexpr std::array handlers {
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags),
SupportsPreserveAllABI};
return true;
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
case IR::OP_F80CVTTOINT: {
auto Op = IROp->C<IR::IROp_F80CVTToInt>();
switch (Op->SrcSize) {
case 2: {
*Info = {FABI_F80_I16_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2, Core::OPINDEX_F80CVTTOINT_2, SupportsPreserveAllABI};
return true;
}
case 4: {
*Info = {FABI_F80_I16_I32, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4, Core::OPINDEX_F80CVTTOINT_4, SupportsPreserveAllABI};
return true;
}
default: LogMan::Msg::DFmt("Unhandled size: {}", OpSize);
}
break;
}
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
}
// Unary
COMMON_UNARY_X87_OP(ROUND)
@@ -242,20 +239,20 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, IR::IROp_He
COMMON_F64_OP(FPREM)
COMMON_F64_OP(SCALE)
// SSE4.2 Fallbacks
case IR::OP_VPCMPESTRX:
*Info = {FABI_I32_I64_I64_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX, SupportsPreserveAllABI};
return true;
case IR::OP_VPCMPISTRX:
*Info = {FABI_I32_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, SupportsPreserveAllABI};
return true;
// SSE4.2 Fallbacks
case IR::OP_VPCMPESTRX:
*Info = {FABI_I32_I64_I64_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPESTRX>::handle, Core::OPINDEX_VPCMPESTRX,
SupportsPreserveAllABI};
return true;
case IR::OP_VPCMPISTRX:
*Info = {FABI_I32_I128_I128_I16, (void*)&FEXCore::CPU::OpHandlers<IR::OP_VPCMPISTRX>::handle, Core::OPINDEX_VPCMPISTRX, SupportsPreserveAllABI};
return true;
default:
break;
default: break;
}
return false;
}
}
} // namespace FEXCore::CPU
@@ -15,9 +15,9 @@ namespace FEXCore::CPU {
template<>
struct OpHandlers<IR::OP_VPCMPESTRX> {
enum class AggregationOp {
EqualAny = 0b00,
Ranges = 0b01,
EqualEach = 0b10,
EqualAny = 0b00,
Ranges = 0b01,
EqualEach = 0b10,
EqualOrdered = 0b11,
};
@@ -35,8 +35,7 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
NegativeMasked,
};
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t handle(uint64_t RAX, uint64_t RDX, __uint128_t lhs, __uint128_t rhs, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t handle(uint64_t RAX, uint64_t RDX, __uint128_t lhs, __uint128_t rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetExplicitLength(RAX, control) - 1;
const auto valid_rhs = GetExplicitLength(RDX, control) - 1;
@@ -45,33 +44,31 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
}
// Main PCMPXSTRX algorithm body. Allows for reuse with both implicit and explicit length variants.
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t MainBody(const __uint128_t& lhs, int valid_lhs, const __uint128_t& rhs, int valid_rhs, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t MainBody(const __uint128_t& lhs, int valid_lhs, const __uint128_t& rhs, int valid_rhs, uint16_t control) {
const uint32_t aggregation = PerformAggregation(lhs, valid_lhs, rhs, valid_rhs, control);
const int32_t upper_limit = (16 >> (control & 1)) - 1;
// Bits are arranged as:
// Bit #: 3 2 1 0
// [OF | CF | SF | ZF]
// [SF | ZF | CF | OF]
uint32_t flags = 0;
flags |= (valid_rhs < upper_limit) ? 0b01 : 0b00;
flags |= (valid_lhs < upper_limit) ? 0b10 : 0b00;
flags |= (valid_rhs < upper_limit) ? 0b0100 : 0b0000;
flags |= (valid_lhs < upper_limit) ? 0b1000 : 0b0000;
const uint32_t result = HandlePolarity(aggregation, control, upper_limit, valid_rhs);
if (result != 0) {
flags |= 0b0100;
flags |= 0b0010;
}
if ((result & 1) != 0) {
flags |= 0b1000;
flags |= 0b0001;
}
// We tack the flags on top of the result to avoid needing to handle
// multiple return values in the JITs.
return result | (flags << 16);
// We track the flags in the usual NZCV bit position so we can msr them
// later. Avoids handling flags natively in JIT.
return result | (flags << 28);
}
FEXCORE_PRESERVE_ALL_ATTR
static int32_t GetExplicitLength(uint64_t reg, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetExplicitLength(uint64_t reg, uint16_t control) {
// Bit 8 controls whether or not the reg value is 64-bit or 32-bit.
int64_t value = 0;
if (((control >> 8) & 1) != 0) {
@@ -94,62 +91,50 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
return std::abs(static_cast<int>(value));
}
FEXCORE_PRESERVE_ALL_ATTR
static int32_t GetElement(const __uint128_t& vec, int32_t index, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetElement(const __uint128_t& vec, int32_t index, uint16_t control) {
const auto* vec_ptr = reinterpret_cast<const uint8_t*>(&vec);
// Control bits [1:0] define the data type being dealt with.
switch (static_cast<SourceData>(control & 0b11)) {
case SourceData::U8:
return static_cast<int32_t>(vec_ptr[index]);
case SourceData::U8: return static_cast<int32_t>(vec_ptr[index]);
case SourceData::U16: {
uint16_t value{};
uint16_t value {};
std::memcpy(&value, vec_ptr + (sizeof(uint16_t) * static_cast<size_t>(index)), sizeof(value));
return value;
}
case SourceData::S8:
return static_cast<int8_t>(vec_ptr[index]);
case SourceData::S8: return static_cast<int8_t>(vec_ptr[index]);
case SourceData::S16:
default: {
int16_t value{};
int16_t value {};
std::memcpy(&value, vec_ptr + (sizeof(int16_t) * static_cast<size_t>(index)), sizeof(value));
return value;
}
}
}
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t PerformAggregation(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t
PerformAggregation(const __uint128_t& lhs, int32_t valid_lhs, const __uint128_t& rhs, int32_t valid_rhs, uint16_t control) {
switch (static_cast<AggregationOp>((control >> 2) & 0b11)) {
case AggregationOp::EqualAny:
return HandleEqualAny(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::Ranges:
return HandleRanges(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualEach:
return HandleEqualEach(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualAny: return HandleEqualAny(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::Ranges: return HandleRanges(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualEach: return HandleEqualEach(lhs, valid_lhs, rhs, valid_rhs, control);
case AggregationOp::EqualOrdered:
default:
return HandleEqualOrdered(lhs, valid_lhs, rhs, valid_rhs, control);
default: return HandleEqualOrdered(lhs, valid_lhs, rhs, valid_rhs, control);
}
}
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t HandlePolarity(uint32_t value, uint16_t control, int upper_limit, int valid_rhs) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t HandlePolarity(uint32_t value, uint16_t control, int upper_limit, int valid_rhs) {
switch (static_cast<Polarity>((control >> 4) & 0b11)) {
case Polarity::Negative:
return value ^ ((2U << upper_limit) - 1);
case Polarity::NegativeMasked:
return value ^ ((1U << (valid_rhs + 1)) - 1);
case Polarity::Positive:
case Polarity::PositiveMasked:
default:
// Both positive masking and positive polarity are documented
// as both being equivalent to "IntRes2 = IntRes1", where IntRes1
// is our 'value' parameter, so we don't need to do anything in
// these cases except return the same value.
return value;
case Polarity::Negative: return value ^ ((2U << upper_limit) - 1);
case Polarity::NegativeMasked: return value ^ ((1U << (valid_rhs + 1)) - 1);
case Polarity::Positive:
case Polarity::PositiveMasked:
default:
// Both positive masking and positive polarity are documented
// as both being equivalent to "IntRes2 = IntRes1", where IntRes1
// is our 'value' parameter, so we don't need to do anything in
// these cases except return the same value.
return value;
}
}
@@ -175,10 +160,8 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
// │
// 'c' match ────────┘
//
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t HandleEqualAny(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t
HandleEqualAny(const __uint128_t& lhs, int32_t valid_lhs, const __uint128_t& rhs, int32_t valid_rhs, uint16_t control) {
uint32_t result = 0;
for (int j = valid_rhs; j >= 0; j--) {
@@ -222,10 +205,8 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
// │
// 'Z' >= 'z' && 'A' <= 'z' ──────────┘
//
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t HandleRanges(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t
HandleRanges(const __uint128_t& lhs, int32_t valid_lhs, const __uint128_t& rhs, int32_t valid_rhs, uint16_t control) {
uint32_t result = 0;
for (int j = valid_rhs; j >= 0; j--) {
@@ -275,10 +256,8 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
// │
// 'a' == 'a' ──────────┘
//
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t HandleEqualEach(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t
HandleEqualEach(const __uint128_t& lhs, int32_t valid_lhs, const __uint128_t& rhs, int32_t valid_rhs, uint16_t control) {
const auto upper_limit = (16 >> (control & 1)) - 1;
const auto max_valid = std::max(valid_lhs, valid_rhs);
const auto min_valid = std::min(valid_lhs, valid_rhs);
@@ -330,10 +309,8 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
// │
// At index 0 ──────────┘
//
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t HandleEqualOrdered(const __uint128_t& lhs, int32_t valid_lhs,
const __uint128_t& rhs, int32_t valid_rhs,
uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t
HandleEqualOrdered(const __uint128_t& lhs, int32_t valid_lhs, const __uint128_t& rhs, int32_t valid_rhs, uint16_t control) {
const auto upper_limit = (16 >> (control & 1)) - 1;
// Edge case!
@@ -345,8 +322,7 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
}
uint32_t result = 0;
const int initial = valid_rhs == upper_limit ? valid_rhs
: valid_rhs - valid_lhs;
const int initial = valid_rhs == upper_limit ? valid_rhs : valid_rhs - valid_lhs;
for (int j = initial; j >= 0; j--) {
result <<= 1;
@@ -379,8 +355,7 @@ struct OpHandlers<IR::OP_VPCMPISTRX> {
// to be the max length possible for the given character size specified
// in the control flags (16 characters for 8-bit, and 8 characters for 16-bit).
//
FEXCORE_PRESERVE_ALL_ATTR
static uint32_t handle(__uint128_t lhs, __uint128_t rhs, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static uint32_t handle(__uint128_t lhs, __uint128_t rhs, uint16_t control) {
// Subtract by 1 in order to make validity limits 0-based
const auto valid_lhs = GetImplicitLength(lhs, control) - 1;
const auto valid_rhs = GetImplicitLength(rhs, control) - 1;
@@ -388,8 +363,7 @@ struct OpHandlers<IR::OP_VPCMPISTRX> {
return OpHandlers<IR::OP_VPCMPESTRX>::MainBody(lhs, valid_lhs, rhs, valid_rhs, control);
}
FEXCORE_PRESERVE_ALL_ATTR
static int32_t GetImplicitLength(const __uint128_t& data, uint16_t control) {
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetImplicitLength(const __uint128_t& data, uint16_t control) {
const auto* data_u8 = reinterpret_cast<const uint8_t*>(&data);
const auto is_using_words = (control & 1) != 0;
@@ -399,7 +373,7 @@ struct OpHandlers<IR::OP_VPCMPISTRX> {
const auto get_word = [data_u8](int32_t index) {
const auto* src = data_u8 + (index * sizeof(uint16_t));
uint16_t element{};
uint16_t element {};
std::memcpy(&element, src, sizeof(uint16_t));
return element;
};
@@ -7,44 +7,43 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
namespace FEXCore::IR {
class IRListView;
struct IROp_Header;
}
class IRListView;
struct IROp_Header;
} // namespace FEXCore::IR
namespace FEXCore::CPU {
enum FallbackABI {
FABI_UNKNOWN,
FABI_F80_I16_F32,
FABI_F80_I16_F64,
FABI_F80_I16_I16,
FABI_F80_I16_I32,
FABI_F32_I16_F80,
FABI_F64_I16_F80,
FABI_F64_I16_F64,
FABI_F64_I16_F64_F64,
FABI_I16_I16_F80,
FABI_I32_I16_F80,
FABI_I64_I16_F80,
FABI_I64_I16_F80_F80,
FABI_F80_I16_F80,
FABI_F80_I16_F80_F80,
FABI_I32_I64_I64_I128_I128_I16,
FABI_I32_I128_I128_I16,
};
enum FallbackABI {
FABI_UNKNOWN,
FABI_F80_I16_F32,
FABI_F80_I16_F64,
FABI_F80_I16_I16,
FABI_F80_I16_I32,
FABI_F32_I16_F80,
FABI_F64_I16_F80,
FABI_F64_I16_F64,
FABI_F64_I16_F64_F64,
FABI_I16_I16_F80,
FABI_I32_I16_F80,
FABI_I64_I16_F80,
FABI_I64_I16_F80_F80,
FABI_F80_I16_F80,
FABI_F80_I16_F80_F80,
FABI_I32_I64_I64_I128_I128_I16,
FABI_I32_I128_I128_I16,
};
struct FallbackInfo {
FallbackABI ABI;
void *fn;
FEXCore::Core::FallbackHandlerIndex HandlerIndex;
bool SupportsPreserveAllABI;
};
struct FallbackInfo {
FallbackABI ABI;
void* fn;
FEXCore::Core::FallbackHandlerIndex HandlerIndex;
bool SupportsPreserveAllABI;
};
class InterpreterOps {
public:
static void FillFallbackIndexPointers(uint64_t *Info);
static bool GetFallbackHandler(bool SupportsPreserveAllABI, IR::IROp_Header const *IROp, FallbackInfo *Info);
};
class InterpreterOps {
public:
static void FillFallbackIndexPointers(uint64_t* Info);
static bool GetFallbackHandler(bool SupportsPreserveAllABI, const IR::IROp_Header* IROp, FallbackInfo* Info);
};
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
@@ -13,21 +13,19 @@ namespace FEXCore::CPU {
uint64_t Arm64JITCore::GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op) {
switch (Op) {
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default:
ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op));
case FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol::SYMBOL_LITERAL_EXITFUNCTION_LINKER:
return ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker;
break;
default: ERROR_AND_DIE_FMT("Unknown named symbol literal: {}", static_cast<uint32_t>(Op)); break;
}
return ~0ULL;
}
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum &Sum) {
Relocation MoveABI{};
void Arm64JITCore::InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum) {
Relocation MoveABI {};
MoveABI.NamedThunkMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.NamedThunkMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.NamedThunkMove.Symbol = Sum;
MoveABI.NamedThunkMove.RegisterIndex = Reg.Idx();
@@ -43,22 +41,25 @@ Arm64JITCore::NamedSymbolLiteralPair Arm64JITCore::InsertNamedSymbolLiteral(FEXC
Arm64JITCore::NamedSymbolLiteralPair Lit {
.Lit = Pointer,
.MoveABI = {
.NamedSymbolLiteral = {
.Header = {
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
.MoveABI =
{
.NamedSymbolLiteral =
{
.Header =
{
.Type = FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL,
},
.Symbol = Op,
.Offset = 0,
},
},
},
};
return Lit;
}
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit) {
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
auto CurrentCursor = GetCursorAddress<uint8_t*>();
Lit.MoveABI.NamedSymbolLiteral.Offset = CurrentCursor - CodeData.BlockBegin;
Bind(&Lit.Loc);
@@ -67,10 +68,10 @@ void Arm64JITCore::PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit) {
}
void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant) {
Relocation MoveABI{};
Relocation MoveABI {};
MoveABI.GuestRIPMove.Header.Type = FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE;
// Offset is the offset from the entrypoint of the block
auto CurrentCursor = GetCursorAddress<uint8_t *>();
auto CurrentCursor = GetCursorAddress<uint8_t*>();
MoveABI.GuestRIPMove.Offset = CurrentCursor - CodeData.BlockBegin;
MoveABI.GuestRIPMove.GuestRIP = Constant;
MoveABI.GuestRIPMove.RegisterIndex = Reg.Idx();
@@ -79,54 +80,54 @@ void Arm64JITCore::InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constan
Relocations.emplace_back(MoveABI);
}
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations) {
size_t DataIndex{};
bool Arm64JITCore::ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations,
const char* EntryRelocations) {
size_t DataIndex {};
for (size_t j = 0; j < NumRelocations; ++j) {
const FEXCore::CPU::Relocation *Reloc = reinterpret_cast<const FEXCore::CPU::Relocation *>(&EntryRelocations[DataIndex]);
const FEXCore::CPU::Relocation* Reloc = reinterpret_cast<const FEXCore::CPU::Relocation*>(&EntryRelocations[DataIndex]);
LOGMAN_THROW_AA_FMT((DataIndex % alignof(Relocation)) == 0, "Alignment of relocation wasn't adhered to");
switch (Reloc->Header.Type) {
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_SYMBOL_LITERAL: {
uint64_t Pointer = GetNamedSymbolLiteral(Reloc->NamedSymbolLiteral.Symbol);
// Relocation occurs at the cursorEntry + offset relative to that cursor
SetCursorOffset(CursorEntry + Reloc->NamedSymbolLiteral.Offset);
// Generate a literal so we can place it
dc64(Pointer);
// Generate a literal so we can place it
dc64(Pointer);
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
DataIndex += sizeof(Reloc->NamedSymbolLiteral);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
case FEXCore::CPU::RelocationTypes::RELOC_NAMED_THUNK_MOVE: {
uint64_t Pointer = reinterpret_cast<uint64_t>(EmitterCTX->ThunkHandler->LookupThunk(Reloc->NamedThunkMove.Symbol));
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->NamedThunkMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->NamedThunkMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->NamedThunkMove);
break;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
case FEXCore::CPU::RelocationTypes::RELOC_GUEST_RIP_MOVE: {
// XXX: Reenable once the JIT Object Cache is upstream
// XXX: Should spin the relocation list, create a list of guest RIP moves, and ask for them all once, reduces lock contention.
uint64_t Pointer = ~0ULL; // EmitterCTX->JITObjectCache->FindRelocatedRIP(Reloc->GuestRIPMove.GuestRIP);
if (Pointer == ~0ULL) {
return false;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
// Relocation occurs at the cursorEntry + offset relative to that cursor.
SetCursorOffset(CursorEntry + Reloc->GuestRIPMove.Offset);
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Register(Reloc->GuestRIPMove.RegisterIndex), Pointer, true);
DataIndex += sizeof(Reloc->GuestRIPMove);
break;
}
}
}
return true;
}
}
} // namespace FEXCore::CPU
@@ -6,12 +6,11 @@ $end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(CASPair) {
auto Op = IROp->C<IR::IROp_CASPair>();
LOGMAN_THROW_AA_FMT(IROp->ElementSize == 4 || IROp->ElementSize == 8, "Wrong element size");
@@ -29,13 +28,20 @@ DEF_OP(CASPair) {
caspal(EmitSize, TMP3, TMP4, Desired.first, Desired.second, MemSrc);
mov(EmitSize, Dst.first, TMP3.R());
mov(EmitSize, Dst.second, TMP4.R());
}
else {
} else {
// Save NZCV so we don't have to mark this op as clobbering NZCV (the
// SupportsAtomics does not clobber atomics and this !SupportsAtomics path
// is so slow it's not worth the complexity of splitting the IR op.). We
// clobber NZCV inside the hot loop and we can't replace cmp/ccmp/b.ne with
// something NZCV-preserving without requiring an extra instruction.
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
Bind(&LoopTop);
// This instruction sequence must be synced with HandleCASPAL_Armv8.
ldaxp(EmitSize, TMP2, TMP3, MemSrc);
cmp(EmitSize, TMP2, Expected.first);
ccmp(EmitSize, TMP3, Expected.second, ARMEmitter::StatusFlags::None, ARMEmitter::Condition::CC_EQ);
@@ -47,20 +53,23 @@ DEF_OP(CASPair) {
b(&LoopExpected);
Bind(&LoopNotExpected);
mov(EmitSize, Dst.first, TMP2.R());
mov(EmitSize, Dst.second, TMP3.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopNotExpected);
mov(EmitSize, Dst.first, TMP2.R());
mov(EmitSize, Dst.second, TMP3.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopExpected);
// Restore
msr(ARMEmitter::SystemRegister::NZCV, TMP1);
}
}
DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
@@ -69,30 +78,21 @@ DEF_OP(CAS) {
auto Desired = GetReg(Op->Desired.ID());
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mov(EmitSize, TMP2, Expected);
casal(SubEmitSize, TMP2, Desired, MemSrc);
mov(EmitSize, GetReg(Node), TMP2.R());
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
ARMEmitter::SingleUseForwardLabel LoopNotExpected;
ARMEmitter::SingleUseForwardLabel LoopExpected;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
if (OpSize == 1) {
if (IROp->Size == 1) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTB, 0);
}
else if (OpSize == 2) {
} else if (IROp->Size == 2) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTH, 0);
}
else {
} else {
cmp(EmitSize, TMP2, Expected);
}
b(ARMEmitter::Condition::CC_NE, &LoopNotExpected);
@@ -101,33 +101,26 @@ DEF_OP(CAS) {
mov(EmitSize, GetReg(Node), Expected);
b(&LoopExpected);
Bind(&LoopNotExpected);
mov(EmitSize, GetReg(Node), TMP2.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopNotExpected);
mov(EmitSize, GetReg(Node), TMP2.R());
// exclusive monitor needs to be cleared here
// Might have hit the case where ldaxr was hit but stlxr wasn't
clrex();
Bind(&LoopExpected);
}
}
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
staddl(SubEmitSize, Src, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -139,23 +132,16 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
staddl(SubEmitSize, TMP2, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -167,23 +153,16 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
stclrl(SubEmitSize, TMP2, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -195,22 +174,15 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicCLR) {
auto Op = IROp->C<IR::IROp_AtomicCLR>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
stclrl(SubEmitSize, Src, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -222,22 +194,15 @@ DEF_OP(AtomicCLR) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
stsetl(SubEmitSize, Src, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -249,22 +214,15 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -276,17 +234,11 @@ DEF_OP(AtomicXor) {
DEF_OP(AtomicNeg) {
auto Op = IROp->C<IR::IROp_AtomicNeg>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -303,16 +255,16 @@ DEF_OP(AtomicSwap) {
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldswpal(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -324,22 +276,15 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldaddal(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -352,23 +297,16 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
ldaddal(SubEmitSize, TMP2, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -381,23 +319,16 @@ DEF_OP(AtomicFetchSub) {
DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
ldclral(SubEmitSize, TMP2, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -410,22 +341,15 @@ DEF_OP(AtomicFetchAnd) {
DEF_OP(AtomicFetchCLR) {
auto Op = IROp->C<IR::IROp_AtomicFetchCLR>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldclral(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -438,22 +362,15 @@ DEF_OP(AtomicFetchCLR) {
DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldsetal(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -466,22 +383,15 @@ DEF_OP(AtomicFetchOr) {
DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldeoral(SubEmitSize, Src, GetReg(Node), MemSrc);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -494,17 +404,11 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -527,8 +431,7 @@ DEF_OP(TelemetrySetValue) {
if (CTX->HostFeatures.SupportsAtomics) {
stsetl(ARMEmitter::SubRegSize::i64Bit, TMP1, TMP2);
}
else {
} else {
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(ARMEmitter::SubRegSize::i64Bit, TMP3, TMP2);
@@ -540,5 +443,4 @@ DEF_OP(TelemetrySetValue) {
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
@@ -7,7 +7,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -19,7 +18,7 @@ $end_info$
#include <Interface/HLE/Thunks/Thunks.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(CallbackReturn) {
// spill back to CTX
@@ -53,14 +52,24 @@ DEF_OP(ExitFunction) {
uint64_t NewRIP;
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
ARMEmitter::SingleUseForwardLabel l_BranchHost;
#ifdef _M_ARM_64EC
if (RtlIsEcCode(NewRIP)) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
} else {
#endif
ARMEmitter::SingleUseForwardLabel l_BranchHost;
ldr(TMP1, &l_BranchHost);
blr(TMP1);
ldr(TMP1, &l_BranchHost);
blr(TMP1);
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
dc64(NewRIP);
Bind(&l_BranchHost);
dc64(ThreadState->CurrentFrame->Pointers.Common.ExitFunctionLinker);
dc64(NewRIP);
#ifdef _M_ARM_64EC
}
#endif
} else {
ARMEmitter::SingleUseForwardLabel FullLookup;
@@ -92,41 +101,13 @@ DEF_OP(Jump) {
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
static ARMEmitter::Condition MapBranchCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU:return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
default:
LOGMAN_MSG_A_FMT("Unsupported compare type");
return ARMEmitter::Condition::CC_NV;
}
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
if (Op->FromNZCV) {
b(MapBranchCC(Op->Cond), TrueTargetLabel);
b(MapCC(Op->Cond), TrueTargetLabel);
} else {
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
@@ -135,13 +116,12 @@ DEF_OP(CondJump) {
LOGMAN_THROW_A_FMT(IsGPR(Op->Cmp1.ID()), "CondJump: Expected GPR");
LOGMAN_THROW_A_FMT(isConst && Const == 0, "CondJump: Expected 0 source");
LOGMAN_THROW_A_FMT(Op->Cond.Val == FEXCore::IR::COND_EQ ||
Op->Cond.Val == FEXCore::IR::COND_NEQ,
"CondJump: Expected simple condition");
LOGMAN_THROW_A_FMT(Op->Cond.Val == FEXCore::IR::COND_EQ || Op->Cond.Val == FEXCore::IR::COND_NEQ, "CondJump: Expected simple "
"condition");
if (Op->Cond.Val == FEXCore::IR::COND_EQ) {
cbz(Size, GetReg(Op->Cmp1.ID()), TrueTargetLabel);
} else {
} else {
cbnz(Size, GetReg(Op->Cmp1.ID()), TrueTargetLabel);
}
@@ -181,7 +161,9 @@ DEF_OP(Syscall) {
uint64_t SPOffset = AlignUp(FEXCore::HLE::SyscallArguments::MAX_ARGS * 8, 16);
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, SPOffset);
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++i) {
if (Op->Header.Args[i].IsInvalid()) continue;
if (Op->Header.Args[i].IsInvalid()) {
continue;
}
str(GetReg(Op->Header.Args[i].ID()).X(), ARMEmitter::Reg::rsp, i * 8);
}
@@ -193,8 +175,7 @@ DEF_OP(Syscall) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, ARMEmitter::Reg::rsp, 0);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, void*, void*, void*>(ARMEmitter::Reg::r3);
}
else {
} else {
blr(ARMEmitter::Reg::r3);
}
@@ -232,20 +213,19 @@ DEF_OP(InlineSyscall) {
// X6: Arg6 - Doesn't exist in x86-64 land. RA INTERSECT
// One argument is removed from the SyscallArguments::MAX_ARGS since the first argument was syscall number
const static std::array<ARMEmitter::XRegister, FEXCore::HLE::SyscallArguments::MAX_ARGS-1> RegArgs = {{
ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::XReg::x2, ARMEmitter::XReg::x3, ARMEmitter::XReg::x4, ARMEmitter::XReg::x5
}};
const static std::array<ARMEmitter::XRegister, FEXCore::HLE::SyscallArguments::MAX_ARGS - 1> RegArgs = {
{ARMEmitter::XReg::x0, ARMEmitter::XReg::x1, ARMEmitter::XReg::x2, ARMEmitter::XReg::x3, ARMEmitter::XReg::x4, ARMEmitter::XReg::x5}};
bool Intersects{};
bool Intersects {};
// We always need to spill x8 since we can't know if it is live at this SSA location
uint32_t SpillMask = 1U << 8;
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
auto Reg = GetReg(Op->Header.Args[i].ID());
if (Reg == ARMEmitter::Reg::r8 ||
Reg == ARMEmitter::Reg::r4 ||
Reg == ARMEmitter::Reg::r5) {
if (Reg == ARMEmitter::Reg::r8 || Reg == ARMEmitter::Reg::r4 || Reg == ARMEmitter::Reg::r5) {
SpillMask |= (1U << Reg.Idx());
Intersects = true;
@@ -269,8 +249,10 @@ DEF_OP(InlineSyscall) {
const auto EmitSize = CTX->Config.Is64BitMode() ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSubSize = CTX->Config.Is64BitMode() ? ARMEmitter::SubRegSize::i64Bit : ARMEmitter::SubRegSize::i32Bit;
if (Intersects) {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
auto Reg = GetReg(Op->Header.Args[i].ID());
// In the case of intersection with x4, x5, or x8 then these are currently SRA
@@ -278,21 +260,19 @@ DEF_OP(InlineSyscall) {
// Just load back from the context. Could be slightly smarter but this is fairly uncommon
if (Reg == ARMEmitter::Reg::r8) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RSP]));
}
else if (Reg == ARMEmitter::Reg::r4) {
} else if (Reg == ARMEmitter::Reg::r4) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RAX]));
}
else if (Reg == ARMEmitter::Reg::r5) {
} else if (Reg == ARMEmitter::Reg::r5) {
ldr(EmitSubSize, RegArgs[i].R(), STATE, offsetof(FEXCore::Core::CpuStateFrame, State.gregs[X86State::REG_RCX]));
}
else {
} else {
mov(EmitSize, RegArgs[i].R(), Reg);
}
}
}
else {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS-1; ++i) {
if (Op->Header.Args[i].IsInvalid()) break;
} else {
for (uint32_t i = 0; i < FEXCore::HLE::SyscallArguments::MAX_ARGS - 1; ++i) {
if (Op->Header.Args[i].IsInvalid()) {
break;
}
mov(EmitSize, RegArgs[i].R(), GetReg(Op->Header.Args[i].ID()));
}
@@ -333,8 +313,7 @@ DEF_OP(Thunk) {
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r2, (uintptr_t)thunkFn);
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
}
else {
} else {
blr(ARMEmitter::Reg::r2);
}
@@ -345,7 +324,7 @@ DEF_OP(Thunk) {
DEF_OP(ValidateCode) {
auto Op = IROp->C<IR::IROp_ValidateCode>();
const auto *OldCode = (const uint8_t *)&Op->CodeOriginalLow;
const auto* OldCode = (const uint8_t*)&Op->CodeOriginalLow;
int len = Op->CodeLength;
int idx = 0;
@@ -355,37 +334,33 @@ DEF_OP(ValidateCode) {
const auto Dst = GetReg(Node);
while (len >= 8)
{
while (len >= 8) {
ldr(ARMEmitter::XReg::x2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t *)(OldCode + idx));
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i64Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 8;
idx += 8;
}
while (len >= 4)
{
while (len >= 4) {
ldr(ARMEmitter::WReg::w2, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t *)(OldCode + idx));
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint32_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 4;
idx += 4;
}
while (len >= 2)
{
while (len >= 2) {
ldrh(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint16_t *)(OldCode + idx));
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint16_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 2;
idx += 2;
}
while (len >= 1)
{
while (len >= 1) {
ldrb(TMP3, TMP1, idx);
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint8_t *)(OldCode + idx));
LoadConstant(ARMEmitter::Size::i64Bit, TMP4, *(const uint8_t*)(OldCode + idx));
cmp(ARMEmitter::Size::i32Bit, TMP3, TMP4);
csel(ARMEmitter::Size::i64Bit, Dst, Dst, TMP2, ARMEmitter::Condition::CC_EQ);
len -= 1;
@@ -407,8 +382,7 @@ DEF_OP(ThreadRemoveCodeEntry) {
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.ThreadRemoveCodeEntryFromJIT));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<void, void*, void*>(ARMEmitter::Reg::r2);
}
else {
} else {
blr(ARMEmitter::Reg::r2);
}
FillStaticRegs();
@@ -439,8 +413,7 @@ DEF_OP(CPUID) {
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, void*, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
} else {
blr(ARMEmitter::Reg::r3);
}
@@ -456,7 +429,7 @@ DEF_OP(CPUID) {
// Results are in x0, x1
// Results want to be in a i64v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i64Bit, Dst.first, TMP1);
mov(ARMEmitter::Size::i64Bit, Dst.first, TMP1);
mov(ARMEmitter::Size::i64Bit, Dst.second, TMP2);
}
@@ -474,8 +447,7 @@ DEF_OP(XGetBV) {
ldr(ARMEmitter::XReg::x2, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.XCRFunction));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, void*, uint32_t>(ARMEmitter::Reg::r2);
}
else {
} else {
blr(ARMEmitter::Reg::r2);
}
@@ -490,10 +462,9 @@ DEF_OP(XGetBV) {
// Results are in x0
// Results want to be in a i32v2 vector
auto Dst = GetRegPair(Node);
mov(ARMEmitter::Size::i32Bit, Dst.first, TMP1);
mov(ARMEmitter::Size::i32Bit, Dst.first, TMP1);
lsr(ARMEmitter::Size::i64Bit, Dst.second, TMP1, 32);
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
@@ -5,11 +5,10 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(VInsGPR) {
const auto Op = IROp->C<IR::IROp_VInsGPR>();
const auto OpSize = IROp->Size;
@@ -18,11 +17,7 @@ DEF_OP(VInsGPR) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
const auto SubEmitSize = ConvertSubRegSize8(IROp);
const auto ElementsPer128Bit = 16 / ElementSize;
const auto Dst = GetVReg(Node);
@@ -65,7 +60,7 @@ DEF_OP(VInsGPR) {
// Inserts the GPR value into the given V register.
// Also automatically adjusts the index in the case of using the
// moved upper lane.
const auto Insert = [&](const FEXCore::ARMEmitter::VRegister& reg, int index) {
const auto Insert = [&](const ARMEmitter::VRegister& reg, int index) {
if (InUpperLane) {
index -= ElementsPer128Bit;
}
@@ -94,21 +89,17 @@ DEF_OP(VCastFromGPR) {
auto Src = GetReg(Op->Src.ID());
switch (Op->Header.ElementSize) {
case 1:
uxtb(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 2:
uxth(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 4:
fmov(ARMEmitter::Size::i32Bit, Dst.S(), Src);
break;
case 8:
fmov(ARMEmitter::Size::i64Bit, Dst.D(), Src);
break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
case 1:
uxtb(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 2:
uxth(ARMEmitter::Size::i32Bit, TMP1, Src);
fmov(ARMEmitter::Size::i32Bit, Dst.S(), TMP1);
break;
case 4: fmov(ARMEmitter::Size::i32Bit, Dst.S(), Src); break;
case 8: fmov(ARMEmitter::Size::i64Bit, Dst.D(), Src); break;
default: LOGMAN_MSG_A_FMT("Unknown castGPR element size: {}", Op->Header.ElementSize);
}
}
@@ -120,16 +111,7 @@ DEF_OP(VDupFromGPR) {
const auto Src = GetReg(Op->Src.ID());
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1,
"Unexpected {} element size: {}", __func__, ElementSize);
const auto SubEmitSize =
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit : ARMEmitter::SubRegSize::i8Bit;
const auto SubEmitSize = ConvertSubRegSize8(IROp);
if (HostSupportsSVE256 && Is256Bit) {
dup(SubEmitSize, Dst.Z(), Src);
@@ -148,34 +130,33 @@ DEF_OP(Float_FromGPR_S) {
auto Src = GetReg(Op->Src.ID());
switch (Conv) {
case 0x0204: { // Half <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.H(), Src);
break;
}
case 0x0208: { // Half <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.H(), Src);
break;
}
case 0x0404: { // Float <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.S(), Src);
break;
}
case 0x0408: { // Float <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.S(), Src);
break;
}
case 0x0804: { // Double <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.D(), Src);
break;
}
case 0x0808: { // Double <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.D(), Src);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled conversion mask: Mask=0x{:04x}, ElementSize={}, SrcElementSize={}",
Conv, ElementSize, Op->SrcElementSize);
break;
case 0x0204: { // Half <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.H(), Src);
break;
}
case 0x0208: { // Half <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.H(), Src);
break;
}
case 0x0404: { // Float <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.S(), Src);
break;
}
case 0x0408: { // Float <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.S(), Src);
break;
}
case 0x0804: { // Double <- int32_t
scvtf(ARMEmitter::Size::i32Bit, Dst.D(), Src);
break;
}
case 0x0808: { // Double <- int64_t
scvtf(ARMEmitter::Size::i64Bit, Dst.D(), Src);
break;
}
default:
LOGMAN_MSG_A_FMT("Unhandled conversion mask: Mask=0x{:04x}, ElementSize={}, SrcElementSize={}", Conv, ElementSize, Op->SrcElementSize);
break;
}
}
@@ -187,31 +168,31 @@ DEF_OP(Float_FToF) {
auto Src = GetVReg(Op->Scalar.ID());
switch (Conv) {
case 0x0204: { // Half <- Float
fcvt(Dst.H(), Src.S());
break;
}
case 0x0208: { // Half <- Double
fcvt(Dst.H(), Src.D());
break;
}
case 0x0402: { // Float <- Half
fcvt(Dst.S(), Src.H());
break;
}
case 0x0802: { // Double <- Half
fcvt(Dst.D(), Src.H());
break;
}
case 0x0804: { // Double <- Float
fcvt(Dst.D(), Src.S());
break;
}
case 0x0408: { // Float <- Double
fcvt(Dst.S(), Src.D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
case 0x0204: { // Half <- Float
fcvt(Dst.H(), Src.S());
break;
}
case 0x0208: { // Half <- Double
fcvt(Dst.H(), Src.D());
break;
}
case 0x0402: { // Float <- Half
fcvt(Dst.S(), Src.H());
break;
}
case 0x0802: { // Double <- Half
fcvt(Dst.D(), Src.H());
break;
}
case 0x0804: { // Double <- Float
fcvt(Dst.D(), Src.S());
break;
}
case 0x0408: { // Float <- Double
fcvt(Dst.S(), Src.D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown FCVT sizes: 0x{:x}", Conv);
}
}
@@ -220,13 +201,9 @@ DEF_OP(Vector_SToF) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE256 && Is256Bit) {
@@ -236,19 +213,15 @@ DEF_OP(Vector_SToF) {
if (OpSize == ElementSize) {
if (ElementSize == 8) {
scvtf(ARMEmitter::ScalarRegSize::i64Bit, Dst.D(), Vector.D());
}
else if (ElementSize == 4) {
} else if (ElementSize == 4) {
scvtf(ARMEmitter::ScalarRegSize::i32Bit, Dst.S(), Vector.S());
}
else {
} else {
scvtf(ARMEmitter::ScalarRegSize::i16Bit, Dst.H(), Vector.H());
}
}
else {
} else {
if (OpSize == 8) {
scvtf(SubEmitSize, Dst.D(), Vector.D());
}
else {
} else {
scvtf(SubEmitSize, Dst.Q(), Vector.Q());
}
}
@@ -260,13 +233,9 @@ DEF_OP(Vector_FToZS) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE256 && Is256Bit) {
@@ -276,19 +245,15 @@ DEF_OP(Vector_FToZS) {
if (OpSize == ElementSize) {
if (ElementSize == 8) {
fcvtzs(ARMEmitter::ScalarRegSize::i64Bit, Dst.D(), Vector.D());
}
else if (ElementSize == 4) {
} else if (ElementSize == 4) {
fcvtzs(ARMEmitter::ScalarRegSize::i32Bit, Dst.S(), Vector.S());
}
else {
} else {
fcvtzs(ARMEmitter::ScalarRegSize::i16Bit, Dst.H(), Vector.H());
}
}
else {
} else {
if (OpSize == 8) {
fcvtzs(SubEmitSize, Dst.D(), Vector.D());
}
else {
} else {
fcvtzs(SubEmitSize, Dst.Q(), Vector.Q());
}
}
@@ -299,13 +264,8 @@ DEF_OP(Vector_FToS) {
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -320,8 +280,7 @@ DEF_OP(Vector_FToS) {
if (OpSize == 8) {
frinti(SubEmitSize, Dst.D(), Vector.D());
fcvtzs(SubEmitSize, Dst.D(), Dst.D());
}
else {
} else {
frinti(SubEmitSize, Dst.Q(), Vector.Q());
fcvtzs(SubEmitSize, Dst.Q(), Dst.Q());
}
@@ -333,14 +292,10 @@ DEF_OP(Vector_FToF) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -361,45 +316,41 @@ DEF_OP(Vector_FToF) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Conv) {
case 0x0402: { // Float <- Half
zip1(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0804: { // Double <- Float
zip1(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0204: { // Half <- Float
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv);
break;
case 0x0402: { // Float <- Half
zip1(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0804: { // Double <- Float
zip1(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0204: { // Half <- Float
fcvtnt(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
}
} else {
switch (Conv) {
case 0x0402: // Float <- Half
case 0x0804: { // Double <- Float
fcvtl(SubEmitSize, Dst.D(), Vector.D());
break;
}
case 0x0204: // Half <- Float
case 0x0408: { // Float <- Double
fcvtn(SubEmitSize, Dst.D(), Vector.D());
break;
}
default:
LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv);
break;
case 0x0402: // Float <- Half
case 0x0804: { // Double <- Float
fcvtl(SubEmitSize, Dst.D(), Vector.D());
break;
}
case 0x0204: // Half <- Float
case 0x0408: { // Float <- Double
fcvtn(SubEmitSize, Dst.D(), Vector.D());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
}
}
}
@@ -409,12 +360,8 @@ DEF_OP(Vector_FToI) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit : ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -423,82 +370,51 @@ DEF_OP(Vector_FToI) {
const auto Mask = PRED_TMP_32B.Merging();
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Towards_Zero.Val:
frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Host.Val:
frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z());
break;
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Z(), Mask, Vector.Z()); break;
}
} else {
const auto IsScalar = ElementSize == OpSize;
if (IsScalar) {
// Since we have multiple overloads of the same name (e.g.
// frinti having AdvSIMD, AdvSIMD scalar, and an SVE version),
// we can't just use a lambda without some seriously ugly casting.
// This is fairly self-contained otherwise.
#define ROUNDING_FN(name) \
if (ElementSize == 2) { \
name(Dst.H(), Vector.H()); \
} else if (ElementSize == 4) { \
name(Dst.S(), Vector.S()); \
} else if (ElementSize == 8) { \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
}
// Since we have multiple overloads of the same name (e.g.
// frinti having AdvSIMD, AdvSIMD scalar, and an SVE version),
// we can't just use a lambda without some seriously ugly casting.
// This is fairly self-contained otherwise.
#define ROUNDING_FN(name) \
if (ElementSize == 2) { \
name(Dst.H(), Vector.H()); \
} else if (ElementSize == 4) { \
name(Dst.S(), Vector.S()); \
} else if (ElementSize == 8) { \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
}
switch (Op->Round) {
case IR::Round_Nearest.Val:
ROUNDING_FN(frintn);
break;
case IR::Round_Negative_Infinity.Val:
ROUNDING_FN(frintm);
break;
case IR::Round_Positive_Infinity.Val:
ROUNDING_FN(frintp);
break;
case IR::Round_Towards_Zero.Val:
ROUNDING_FN(frintz);
break;
case IR::Round_Host.Val:
ROUNDING_FN(frinti);
break;
case IR::Round_Nearest.Val: ROUNDING_FN(frintn); break;
case IR::Round_Negative_Infinity.Val: ROUNDING_FN(frintm); break;
case IR::Round_Positive_Infinity.Val: ROUNDING_FN(frintp); break;
case IR::Round_Towards_Zero.Val: ROUNDING_FN(frintz); break;
case IR::Round_Host.Val: ROUNDING_FN(frinti); break;
}
#undef ROUNDING_FN
#undef ROUNDING_FN
} else {
switch (Op->Round) {
case FEXCore::IR::Round_Nearest.Val:
frintn(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Negative_Infinity.Val:
frintm(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Positive_Infinity.Val:
frintp(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Towards_Zero.Val:
frintz(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Host.Val:
frinti(SubEmitSize, Dst.Q(), Vector.Q());
break;
case FEXCore::IR::Round_Nearest.Val: frintn(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(SubEmitSize, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(SubEmitSize, Dst.Q(), Vector.Q()); break;
}
}
}
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
@@ -5,12 +5,10 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(VAESImc) {
auto Op = IROp->C<IR::IROp_VAESImc>();
@@ -26,8 +24,7 @@ DEF_OP(VAESEnc) {
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -35,8 +32,7 @@ DEF_OP(VAESEnc) {
aese(Dst.Q(), ZeroReg.Q());
aesmc(Dst.Q(), Dst.Q());
eor(Dst.Q(), Dst.Q(), Key.Q());
}
else {
} else {
mov(VTMP1.Q(), State.Q());
aese(VTMP1, ZeroReg.Q());
aesmc(VTMP1, VTMP1);
@@ -53,16 +49,14 @@ DEF_OP(VAESEncLast) {
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
// This matches the common case of XMM AES.
aese(Dst.Q(), ZeroReg.Q());
eor(Dst.Q(), Dst.Q(), Key.Q());
}
else {
} else {
mov(VTMP1.Q(), State.Q());
aese(VTMP1, ZeroReg.Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
@@ -78,8 +72,7 @@ DEF_OP(VAESDec) {
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
@@ -87,8 +80,7 @@ DEF_OP(VAESDec) {
aesd(Dst.Q(), ZeroReg.Q());
aesimc(Dst.Q(), Dst.Q());
eor(Dst.Q(), Dst.Q(), Key.Q());
}
else {
} else {
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, ZeroReg.Q());
aesimc(VTMP1, VTMP1);
@@ -105,16 +97,14 @@ DEF_OP(VAESDecLast) {
const auto State = GetVReg(Op->State.ID());
const auto ZeroReg = GetVReg(Op->ZeroReg.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
if (Dst == State && Dst != Key) {
// Optimal case in which Dst already contains the starting state.
// This matches the common case of XMM AES.
aesd(Dst.Q(), ZeroReg.Q());
eor(Dst.Q(), Dst.Q(), Key.Q());
}
else {
} else {
mov(VTMP1.Q(), State.Q());
aesd(VTMP1, ZeroReg.Q());
eor(Dst.Q(), VTMP1.Q(), Key.Q());
@@ -149,8 +139,7 @@ DEF_OP(VAESKeyGenAssist) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, static_cast<uint64_t>(Op->RCON) << 32);
dup(ARMEmitter::SubRegSize::i64Bit, VTMP2.Q(), TMP1);
eor(Dst.Q(), Dst.Q(), VTMP2.Q());
}
else {
} else {
tbl(Dst.Q(), Dst.Q(), Swizzle.Q());
}
}
@@ -163,19 +152,11 @@ DEF_OP(CRC32) {
const auto Src2 = GetReg(Op->Src2.ID());
switch (Op->SrcSize) {
case 1:
crc32cb(Dst.W(), Src1.W(), Src2.W());
break;
case 2:
crc32ch(Dst.W(), Src1.W(), Src2.W());
break;
case 4:
crc32cw(Dst.W(), Src1.W(), Src2.W());
break;
case 8:
crc32cx(Dst.X(), Src1.X(), Src2.X());
break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
case 1: crc32cb(Dst.W(), Src1.W(), Src2.W()); break;
case 2: crc32ch(Dst.W(), Src1.W(), Src2.W()); break;
case 4: crc32cw(Dst.W(), Src1.W(), Src2.W()); break;
case 8: crc32cx(Dst.X(), Src1.X(), Src2.X()); break;
default: LOGMAN_MSG_A_FMT("Unknown CRC32 size: {}", Op->SrcSize);
}
}
@@ -197,11 +178,10 @@ DEF_OP(VSha256U0) {
if (Dst == Src1) {
sha256su0(Dst, Src2);
}
else {
} else {
mov(VTMP1.Q(), Src1.Q());
sha256su0(VTMP1, Src2);
mov(Dst.Q(), Src1.Q());
mov(Dst.Q(), VTMP1.Q());
}
}
@@ -209,17 +189,14 @@ DEF_OP(PCLMUL) {
const auto Op = IROp->C<IR::IROp_PCLMUL>();
const auto OpSize = IROp->Size;
const auto Dst = GetVReg(Node);
const auto Dst = GetVReg(Node);
const auto Src1 = GetVReg(Op->Src1.ID());
const auto Src2 = GetVReg(Op->Src2.ID());
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE,
"Currently only supports 128-bit operations.");
LOGMAN_THROW_AA_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "Currently only supports 128-bit operations.");
switch (Op->Selector) {
case 0b00000000:
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D());
break;
case 0b00000000: pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), Src1.D(), Src2.D()); break;
case 0b00000001:
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src1.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src2.D());
@@ -228,14 +205,10 @@ DEF_OP(PCLMUL) {
dup(ARMEmitter::SubRegSize::i64Bit, VTMP1.Q(), Src2.Q(), 1);
pmull(ARMEmitter::SubRegSize::i128Bit, Dst.D(), VTMP1.D(), Src1.D());
break;
case 0b00010001:
pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q());
break;
default:
LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector);
break;
case 0b00010001: pmull2(ARMEmitter::SubRegSize::i128Bit, Dst.Q(), Src1.Q(), Src2.Q()); break;
default: LOGMAN_MSG_A_FMT("Unknown PCLMUL selector: {}", Op->Selector); break;
}
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
@@ -8,12 +8,11 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
+332 -383
View File
@@ -13,7 +13,6 @@ $end_info$
#include "FEXCore/Utils/Telemetry.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -73,11 +72,11 @@ static void PrintValue(uint64_t Value) {
static void PrintVectorValue(uint64_t Value, uint64_t ValueUpper) {
LogMan::Msg::DFmt("Value: 0x{:016x}'{:016x}", ValueUpper, Value);
}
}
} // namespace
namespace FEXCore::CPU {
void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
FallbackInfo Info;
if (!InterpreterOps::GetFallbackHandler(CTX->HostFeatures.SupportsPreserveAllABI, IROp, &Info)) {
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
@@ -118,379 +117,347 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
mov(Dst.W(), TMP1.W());
};
switch(Info.ABI) {
case FABI_F80_I16_F32:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
switch (Info.ABI) {
case FABI_F80_I16_F32: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
fmov(ARMEmitter::SReg::s0, Src1.S());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, float>(ARMEmitter::Reg::r1);
}
else {
blr(ARMEmitter::Reg::r1);
}
FillF80Result();
const auto Src1 = GetVReg(IROp->Args[0].ID());
fmov(ARMEmitter::SReg::s0, Src1.S());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, float>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
break;
case FABI_F80_I16_F64:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF80Result();
} break;
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, double>(ARMEmitter::Reg::r1);
}
else {
blr(ARMEmitter::Reg::r1);
}
case FABI_F80_I16_F64: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF80Result();
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, double>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
break;
case FABI_F80_I16_I16:
case FABI_F80_I16_I32: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF80Result();
} break;
const auto Src1 = GetReg(IROp->Args[0].ID());
if (Info.ABI == FABI_F80_I16_I16) {
sxth(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
else {
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint32_t>(ARMEmitter::Reg::r2);
}
else {
blr(ARMEmitter::Reg::r2);
}
case FABI_F80_I16_I16:
case FABI_F80_I16_I32: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF80Result();
const auto Src1 = GetReg(IROp->Args[0].ID());
if (Info.ABI == FABI_F80_I16_I16) {
sxth(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
} else {
mov(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, Src1);
}
break;
case FABI_F32_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<float, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
fmov(VTMP1.S(), ARMEmitter::SReg::s0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetVReg(Node);
fmov(Dst.S(), VTMP1.S());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x2, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint32_t>(ARMEmitter::Reg::r2);
} else {
blr(ARMEmitter::Reg::r2);
}
break;
case FABI_F64_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF80Result();
} break;
const auto Src1 = GetVReg(IROp->Args[0].ID());
case FABI_F32_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
FillF64Result();
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<float, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
break;
case FABI_F64_I16_F64: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double>(ARMEmitter::Reg::r1);
}
else {
blr(ARMEmitter::Reg::r1);
}
FillF64Result();
if (!TMP_ABIARGS) {
fmov(VTMP1.S(), ARMEmitter::SReg::s0);
}
break;
FillForABICall(Info.SupportsPreserveAllABI, true);
case FABI_F64_I16_F64_F64: {
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
const auto Dst = GetVReg(Node);
fmov(Dst.S(), VTMP1.S());
} break;
mov(VTMP1.D(), Src1.D());
mov(VTMP2.D(), Src2.D());
case FABI_F64_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
if (!TMP_ABIARGS) {
mov(ARMEmitter::DReg::d0, VTMP1.D());
mov(ARMEmitter::DReg::d1, VTMP2.D());
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, double>(ARMEmitter::Reg::r1);
}
else {
blr(ARMEmitter::Reg::r1);
}
FillF64Result();
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
break;
case FABI_I16_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillF64Result();
} break;
const auto Src1 = GetVReg(IROp->Args[0].ID());
case FABI_F64_I16_F64: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
sxth(ARMEmitter::Size::i64Bit, Dst, TMP1);
mov(ARMEmitter::DReg::d0, Src1.D());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
break;
case FABI_I32_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
FillF64Result();
} break;
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
case FABI_F64_I16_F64_F64: {
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
mov(VTMP1.D(), Src1.D());
mov(VTMP2.D(), Src2.D());
FillI32Result();
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
if (!TMP_ABIARGS) {
mov(ARMEmitter::DReg::d0, VTMP1.D());
mov(ARMEmitter::DReg::d1, VTMP2.D());
}
break;
case FABI_I64_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
ldr(ARMEmitter::XReg::x1, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<double, uint16_t, double, double>(ARMEmitter::Reg::r1);
} else {
blr(ARMEmitter::Reg::r1);
}
break;
case FABI_I64_I16_F80_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
FillF64Result();
} break;
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
case FABI_I16_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
}
else {
blr(ARMEmitter::Reg::r5);
}
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
break;
case FABI_F80_I16_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
else {
blr(ARMEmitter::Reg::r3);
}
FillF80Result();
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
break;
case FABI_F80_I16_F80_F80:{
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
const auto Dst = GetReg(Node);
sxth(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_I32_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
const auto Src1 = GetVReg(IROp->Args[0].ID());
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
}
else {
blr(ARMEmitter::Reg::r5);
}
FillF80Result();
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
break;
case FABI_I32_I64_I64_I128_I128_I16: {
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
mov(TMP1, SrcRAX.X());
mov(TMP2, SrcRDX.X());
FillI32Result();
} break;
case FABI_I64_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
SpillForABICall(Info.SupportsPreserveAllABI, TMP3, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Control = Op->Control;
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x0, TMP1);
mov(ARMEmitter::XReg::x1, TMP2);
}
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r4, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r5, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r6, Control);
ldr(ARMEmitter::XReg::x7, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r7);
}
else {
blr(ARMEmitter::Reg::r7);
}
FillI32Result();
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
break;
case FABI_I32_I128_I128_I16: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Control = Op->Control;
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r0, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r4, Control);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r5);
}
else {
blr(ARMEmitter::Reg::r5);
}
FillI32Result();
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
break;
case FABI_UNKNOWN:
default:
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_I64_I16_F80_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint64_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
} else {
blr(ARMEmitter::Reg::r5);
}
if (!TMP_ABIARGS) {
mov(TMP1, ARMEmitter::XReg::x0);
}
FillForABICall(Info.SupportsPreserveAllABI, true);
const auto Dst = GetReg(Node);
mov(ARMEmitter::Size::i64Bit, Dst, TMP1);
} break;
case FABI_F80_I16_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
ldr(ARMEmitter::XReg::x3, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
} else {
blr(ARMEmitter::Reg::r3);
}
FillF80Result();
} break;
case FABI_F80_I16_F80_F80: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Src1 = GetVReg(IROp->Args[0].ID());
const auto Src2 = GetVReg(IROp->Args[1].ID());
ldrh(ARMEmitter::WReg::w0, STATE, offsetof(FEXCore::Core::CPUState, FCW));
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r2, Src1, 4);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 0);
umov<ARMEmitter::SubRegSize::i16Bit>(ARMEmitter::Reg::r4, Src2, 4);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<__uint128_t, uint16_t, uint64_t, uint64_t, uint64_t, uint64_t>(ARMEmitter::Reg::r5);
} else {
blr(ARMEmitter::Reg::r5);
}
FillF80Result();
} break;
case FABI_I32_I64_I64_I128_I128_I16: {
const auto Op = IROp->C<IR::IROp_VPCMPESTRX>();
const auto SrcRAX = GetReg(Op->RAX.ID());
const auto SrcRDX = GetReg(Op->RDX.ID());
mov(TMP1, SrcRAX.X());
mov(TMP2, SrcRDX.X());
SpillForABICall(Info.SupportsPreserveAllABI, TMP3, true);
const auto Control = Op->Control;
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x0, TMP1);
mov(ARMEmitter::XReg::x1, TMP2);
}
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r4, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r5, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r6, Control);
ldr(ARMEmitter::XReg::x7, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r7);
} else {
blr(ARMEmitter::Reg::r7);
}
FillI32Result();
} break;
case FABI_I32_I128_I128_I16: {
SpillForABICall(Info.SupportsPreserveAllABI, TMP1, true);
const auto Op = IROp->C<IR::IROp_VPCMPISTRX>();
const auto Src1 = GetVReg(Op->LHS.ID());
const auto Src2 = GetVReg(Op->RHS.ID());
const auto Control = Op->Control;
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r0, Src1, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r1, Src1, 1);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r2, Src2, 0);
umov<ARMEmitter::SubRegSize::i64Bit>(ARMEmitter::Reg::r3, Src2, 1);
movz(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r4, Control);
ldr(ARMEmitter::XReg::x5, STATE_PTR(CpuStateFrame, Pointers.Common.FallbackHandlerPointers[Info.HandlerIndex]));
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
GenerateIndirectRuntimeCall<uint32_t, uint64_t, uint64_t, uint64_t, uint64_t, uint16_t>(ARMEmitter::Reg::r5);
} else {
blr(ARMEmitter::Reg::r5);
}
FillI32Result();
} break;
case FABI_UNKNOWN:
default:
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}",
FEXCore::IR::GetName(IROp->Op), ToUnderlying(Info.ABI));
LOGMAN_MSG_A_FMT("Unhandled IR Fallback ABI: {} {}", FEXCore::IR::GetName(IROp->Op), ToUnderlying(Info.ABI));
#endif
break;
}
@@ -498,24 +465,24 @@ void Arm64JITCore::Op_Unhandled(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
uintptr_t branch = (uintptr_t)(Record) - 8;
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
FEXCore::ARMEmitter::SingleUseForwardLabel l_BranchHost;
uintptr_t branch = (uintptr_t)(Record)-8;
ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
ARMEmitter::SingleUseForwardLabel l_BranchHost;
emit.ldr(TMP1, &l_BranchHost);
emit.blr(TMP1);
emit.Bind(&l_BranchHost);
emit.dc64(LinkerAddress);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 8);
ARMEmitter::Emitter::ClearICache((void*)branch, 8);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
Record->HostBranch = LinkerAddress;
}
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Frame, FEXCore::Context::ExitFunctionLinkData *Record) {
static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto Thread = Frame->Thread;
auto GuestRip = Record->GuestRIP;
@@ -526,15 +493,15 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
return Frame->Pointers.Common.DispatcherLoopTop;
}
uintptr_t branch = (uintptr_t)(Record) - 8;
uintptr_t branch = (uintptr_t)(Record)-8;
auto offset = HostCode/4 - branch/4;
auto offset = HostCode / 4 - branch / 4;
if (vixl::IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
emit.b(offset);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 4);
ARMEmitter::Emitter::ClearICache((void*)branch, 4);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
@@ -549,39 +516,30 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame *Fram
return HostCode;
}
void Arm64JITCore::Op_NoOp(IR::IROp_Header const *IROp, IR::NodeID Node) {
}
void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::NodeID Node) {}
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread)
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx)
, HostSupportsSVE128{ctx->HostFeatures.SupportsSVE}
, HostSupportsSVE256{ctx->HostFeatures.SupportsAVX}
, HostSupportsRPRES{ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP{ctx->HostFeatures.SupportsAFP}
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsAVX}
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AllocateRegisterSet(RegisterClasses);
RAPass->AddRegisters(FEXCore::IR::GPRClass, GeneralRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, StaticRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRClass, GeneralFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, StaticFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, GeneralPairRegisters.size());
RAPass->AddRegisters(FEXCore::IR::ComplexClass, 1);
for (uint32_t i = 0; i < GeneralPairRegisters.size(); ++i) {
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2, FEXCore::IR::GPRPairClass, i);
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2 + 1, FEXCore::IR::GPRPairClass, i);
}
RAPass->PairRegs = PairRegisters;
{
// Set up pointers that the JIT needs to load
// Common
auto &Common = ThreadState->CurrentFrame->Pointers.Common;
auto& Common = ThreadState->CurrentFrame->Pointers.Common;
Common.PrintValue = reinterpret_cast<uint64_t>(PrintValue);
Common.PrintVectorValue = reinterpret_cast<uint64_t>(PrintVectorValue);
@@ -609,7 +567,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
InterpreterOps::FillFallbackIndexPointers(Common.FallbackHandlerPointers);
// Platform Specific
auto &AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
auto& AArch64 = ThreadState->CurrentFrame->Pointers.AArch64;
AArch64.LUDIV = reinterpret_cast<uint64_t>(LUDIV);
AArch64.LDIV = reinterpret_cast<uint64_t>(LDIV);
@@ -624,8 +582,7 @@ Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::In
if (ParanoidTSO()) {
RT_LoadMemTSO = &Arm64JITCore::Op_ParanoidLoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_ParanoidStoreMemTSO;
}
else {
} else {
RT_LoadMemTSO = &Arm64JITCore::Op_LoadMemTSO;
RT_StoreMemTSO = &Arm64JITCore::Op_StoreMemTSO;
}
@@ -645,9 +602,7 @@ void Arm64JITCore::ClearCache() {
EmitDetectionString();
}
Arm64JITCore::~Arm64JITCore() {
}
Arm64JITCore::~Arm64JITCore() {}
bool Arm64JITCore::IsInlineConstant(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const {
auto OpHeader = IR->GetOp<IR::IROp_Header>(WNode);
@@ -704,10 +659,8 @@ bool Arm64JITCore::IsGPRPair(IR::NodeID Node) const {
return Class == IR::GPRPairClass;
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) {
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
@@ -727,9 +680,9 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
CodeData.BlockBegin = GetCursorAddress<uint8_t*>();
// Put the code header at the start of the data block.
ARMEmitter::BackwardLabel JITCodeHeaderLabel{};
ARMEmitter::BackwardLabel JITCodeHeaderLabel {};
Bind(&JITCodeHeaderLabel);
JITCodeHeader *CodeHeader = GetCursorAddress<JITCodeHeader *>();
JITCodeHeader* CodeHeader = GetCursorAddress<JITCodeHeader*>();
CursorIncrement(sizeof(JITCodeHeader));
#ifdef VIXL_DISASSEMBLER
@@ -766,11 +719,11 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
if (CTX->Config.NeedsPendingInterruptFaultCheck) {
// Trigger a fault if there are any pending interrupts
// Used only for suspend on WIN32 at the moment
strb(ARMEmitter::XReg::zr, STATE, offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) -
offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
}
//LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
// LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
SpillSlots = RAData->SpillSlots();
@@ -794,14 +747,13 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == IR::OP_CODEBLOCK, "IR type failed to be a code block");
#endif
auto BlockStartHostCode = GetCursorAddress<uint8_t *>();
auto BlockStartHostCode = GetCursorAddress<uint8_t*>();
{
const auto Node = IR->GetID(BlockNode);
const auto IsTarget = JumpTargets.try_emplace(Node).first;
// if there's a pending branch, and it is not fall-through
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second)
{
if (PendingTargetLabel && PendingTargetLabel != &IsTarget->second) {
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
@@ -812,43 +764,40 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
for (auto [CodeNode, IROp] : IR->GetCode(BlockNode)) {
const auto ID = IR->GetID(CodeNode);
switch (IROp->Op) {
#define REGISTER_OP_RT(op, x) case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, ID); break
#define REGISTER_OP(op, x) case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, ID); break
#define REGISTER_OP_RT(op, x) \
case FEXCore::IR::IROps::OP_##op: std::invoke(RT_##x, this, IROp, ID); break
#define REGISTER_OP(op, x) \
case FEXCore::IR::IROps::OP_##op: Op_##x(IROp, ID); break
#define IROP_DISPATCH_DISPATCH
#include <FEXCore/IR/IRDefines_Dispatch.inc>
#undef REGISTER_OP
default:
Op_Unhandled(IROp, ID);
break;
default: Op_Unhandled(IROp, ID); break;
}
}
if (DebugData) {
DebugData->Subblocks.push_back({
static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t *>() - BlockStartHostCode)
});
DebugData->Subblocks.push_back({static_cast<uint32_t>(BlockStartHostCode - CodeData.BlockEntry),
static_cast<uint32_t>(GetCursorAddress<uint8_t*>() - BlockStartHostCode)});
}
}
// Make sure last branch is generated. It certainly can't be eliminated here.
if (PendingTargetLabel)
{
if (PendingTargetLabel) {
b(PendingTargetLabel);
}
PendingTargetLabel = nullptr;
// CodeSize not including the tail data.
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t *>() - CodeData.BlockBegin;
const uint64_t CodeOnlySize = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
// Add the JitCodeTail
auto JITBlockTailLocation = GetCursorAddress<uint8_t *>();
auto JITBlockTailLocation = GetCursorAddress<uint8_t*>();
auto JITBlockTail = GetCursorAddress<JITCodeTail*>();
CursorIncrement(sizeof(JITCodeTail));
auto JITRIPEntriesLocation = GetCursorAddress<uint8_t *>();
auto JITRIPEntriesLocation = GetCursorAddress<uint8_t*>();
auto JITRIPEntries = GetCursorAddress<JITRIPReconstructEntries*>();
CursorIncrement(sizeof(JITRIPReconstructEntries) * DebugData->GuestOpcodes.size());
@@ -867,8 +816,8 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
uintptr_t CurrentRIPOffset = 0;
uint64_t CurrentPCOffset = 0;
for (size_t i = 0; i < DebugData->GuestOpcodes.size(); i++) {
const auto &GuestOpcode = DebugData->GuestOpcodes[i];
auto &RIPEntry = JITRIPEntries[i];
const auto& GuestOpcode = DebugData->GuestOpcodes[i];
auto& RIPEntry = JITRIPEntries[i];
RIPEntry.HostPCOffset = GuestOpcode.HostEntryOffset - CurrentPCOffset;
RIPEntry.GuestRIPOffset = GuestOpcode.GuestEntryOffset - CurrentRIPOffset;
CurrentPCOffset = GuestOpcode.HostEntryOffset;
@@ -878,7 +827,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry,
CodeHeader->OffsetToBlockTail = JITBlockTailLocation - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t *>() - CodeData.BlockBegin;
CodeData.Size = GetCursorAddress<uint8_t*>() - CodeData.BlockBegin;
JITBlockTail->Size = CodeData.Size;
@@ -933,7 +882,7 @@ void Arm64JITCore::ResetStack() {
}
}
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx, FEXCore::Core::InternalThreadState *Thread) {
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread) {
return fextl::make_unique<Arm64JITCore>(ctx, Thread);
}
@@ -945,4 +894,4 @@ CPUBackendFeatures GetArm64JITBackendFeatures() {
};
}
}
} // namespace FEXCore::CPU
@@ -8,69 +8,83 @@ $end_info$
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/IR.h"
#include "Interface/IR/IntrusiveIRList.h"
#include "Interface/IR/RegisterAllocationData.h"
#include <aarch64/assembler-aarch64.h>
#include <aarch64/disasm-aarch64.h>
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/CPUBackend.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/IR/RegisterAllocationData.h>
#include <FEXCore/fextl/map.h>
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <array>
#include <cstdint>
#include <utility>
#include <variant>
namespace FEXCore::Core {
struct InternalThreadState;
struct InternalThreadState;
}
namespace FEXCore::CPU {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
class Arm64JITCore final : public CPUBackend, public Arm64Emitter {
public:
explicit Arm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
explicit Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
~Arm64JITCore() override;
[[nodiscard]] fextl::string GetName() override { return "JIT"; }
[[nodiscard]]
fextl::string GetName() override {
return "JIT";
}
[[nodiscard]] CPUBackend::CompiledCode CompileCode(uint64_t Entry,
FEXCore::IR::IRListView const *IR,
FEXCore::Core::DebugData *DebugData,
FEXCore::IR::RegisterAllocationData *RAData) override;
[[nodiscard]]
CPUBackend::CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
const FEXCore::IR::RegisterAllocationData* RAData) override;
[[nodiscard]] void *MapRegion(void* HostPtr, uint64_t, uint64_t) override { return HostPtr; }
[[nodiscard]]
void* MapRegion(void* HostPtr, uint64_t, uint64_t) override {
return HostPtr;
}
[[nodiscard]] bool NeedsOpDispatch() override { return true; }
[[nodiscard]]
bool NeedsOpDispatch() override {
return true;
}
void ClearCache() override;
void ClearRelocations() override { Relocations.clear(); }
void ClearRelocations() override {
Relocations.clear();
}
private:
FEX_CONFIG_OPT(ParanoidTSO, PARANOIDTSO);
FEX_CONFIG_OPT(VectorTSOEnabled, VECTORTSOENABLED);
FEX_CONFIG_OPT(MemcpySetTSOEnabled, MEMCPYSETTSOENABLED);
const bool HostSupportsSVE128{};
const bool HostSupportsSVE256{};
const bool HostSupportsRPRES{};
const bool HostSupportsAFP{};
const bool HostSupportsSVE128 {};
const bool HostSupportsSVE256 {};
const bool HostSupportsRPRES {};
const bool HostSupportsAFP {};
ARMEmitter::BiDirectionalLabel *PendingTargetLabel;
FEXCore::Context::ContextImpl *CTX;
FEXCore::IR::IRListView const *IR;
ARMEmitter::BiDirectionalLabel* PendingTargetLabel;
FEXCore::Context::ContextImpl* CTX;
const FEXCore::IR::IRListView* IR;
uint64_t Entry;
CPUBackend::CompiledCode CodeData{};
CPUBackend::CompiledCode CodeData {};
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
[[nodiscard]] FEXCore::ARMEmitter::Register GetReg(IR::NodeID Node) const {
[[nodiscard]]
ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
@@ -84,7 +98,8 @@ private:
FEX_UNREACHABLE;
}
[[nodiscard]] FEXCore::ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
[[nodiscard]]
ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
@@ -98,17 +113,20 @@ private:
FEX_UNREACHABLE;
}
[[nodiscard]] std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
[[nodiscard]]
std::pair<ARMEmitter::Register, ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRPairClass.Val, "Unexpected Class: {}", Reg.Class);
return GeneralPairRegisters[Reg.Reg];
return std::make_pair(GeneralRegisters[Reg.Reg], GeneralRegisters[Reg.Reg + 1]);
}
[[nodiscard]] FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
[[nodiscard]]
FEXCore::IR::RegisterClassType GetRegClass(IR::NodeID Node) const;
[[nodiscard]] IR::PhysicalRegister GetPhys(IR::NodeID Node) const {
[[nodiscard]]
IR::PhysicalRegister GetPhys(IR::NodeID Node) const {
auto PhyReg = RAData->GetNodeRegister(Node);
LOGMAN_THROW_A_FMT(!PhyReg.IsInvalid(), "Couldn't Allocate register for node: ssa{}. Class: {}", Node, PhyReg.Class);
@@ -116,7 +134,8 @@ private:
return PhyReg;
}
[[nodiscard]] FEXCore::ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
[[nodiscard]]
ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
uint64_t Const;
if (IsInlineConstant(Src, &Const)) {
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
@@ -128,36 +147,131 @@ private:
// Converts IR-base shift type to ARMEmitter shift type.
// Will be a no-op, only a type conversion since the two definitions match.
[[nodiscard]] ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) const {
[[nodiscard]]
ARMEmitter::ShiftType ConvertIRShiftType(IR::ShiftType Shift) const {
return Shift == IR::ShiftType::LSL ? ARMEmitter::ShiftType::LSL :
Shift == IR::ShiftType::LSR ? ARMEmitter::ShiftType::LSR :
Shift == IR::ShiftType::ASR ? ARMEmitter::ShiftType::ASR :
ARMEmitter::ShiftType::ROR;
ARMEmitter::ShiftType::ROR;
}
[[nodiscard]] bool IsFPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPR(IR::NodeID Node) const;
[[nodiscard]] bool IsGPRPair(IR::NodeID Node) const;
[[nodiscard]]
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
return Op->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]] FEXCore::ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize,
FEXCore::ARMEmitter::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
[[nodiscard]]
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->Size == 4 || Op->Size == 8, "Invalid size");
return ConvertSize(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8 || ElementSize == 16, "Invalid size");
return ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ARMEmitter::SubRegSize::i128Bit;
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
return ConvertSubRegSize16(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize != 16, "Invalid size");
return ConvertSubRegSize16(ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
return ConvertSubRegSize8(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 8, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
return ARMEmitter::ToVectorSizePair(ConvertSubRegSize16(Op));
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 16, "Invalid size");
return ConvertSubRegSizePair16(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
return ConvertSubRegSizePair8(Op);
}
[[nodiscard]]
ARMEmitter::Condition MapCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
}
}
[[nodiscard]]
bool IsFPR(IR::NodeID Node) const;
[[nodiscard]]
bool IsGPR(IR::NodeID Node) const;
[[nodiscard]]
bool IsGPRPair(IR::NodeID Node) const;
[[nodiscard]]
ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
// NOTE: Will use TMP1 as a way to encode immediates that happen to fall outside
// the limits of the scalar plus immediate variant of SVE load/stores.
//
// TMP1 is safe to use again once this memory operand is used with its
// equivalent loads or stores that this was called for.
[[nodiscard]] FEXCore::ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize,
FEXCore::ARMEmitter::Register Base,
IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType,
uint8_t OffsetScale);
[[nodiscard]]
ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
[[nodiscard]] bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
[[nodiscard]] bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
[[nodiscard]]
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
[[nodiscard]]
bool IsInlineEntrypointOffset(const IR::OrderedNodeWrapper& WNode, uint64_t* Value) const;
struct LiveRange {
uint32_t Begin;
@@ -166,81 +280,83 @@ private:
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass *RAPass;
IR::RegisterAllocationData *RAData;
FEXCore::Core::DebugData *DebugData;
IR::RegisterAllocationPass* RAPass;
const IR::RegisterAllocationData* RAData;
FEXCore::Core::DebugData* DebugData;
void ResetStack();
/**
* @name Relocations
* @{ */
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
uint64_t GetNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
ARMEmitter::ForwardLabel Loc;
uint64_t Lit;
Relocation MoveABI{};
};
/**
* @brief A literal pair relocation object for named symbol literals
*/
struct NamedSymbolLiteralPair {
ARMEmitter::ForwardLabel Loc;
uint64_t Lit;
Relocation MoveABI {};
};
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum &Sum);
/**
* @brief Inserts a thunk relocation
*
* @param Reg - The GPR to move the thunk handler in to
* @param Sum - The hash of the thunk
*/
void InsertNamedThunkRelocation(ARMEmitter::Register Reg, const IR::SHA256Sum& Sum);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant);
/**
* @brief Inserts a guest GPR move relocation
*
* @param Reg - The GPR to move the guest RIP in to
* @param Constant - The guest RIP that will be relocated
*/
void InsertGuestRIPMove(ARMEmitter::Register Reg, uint64_t Constant);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Inserts a named symbol as a literal in memory
*
* Need to use `PlaceNamedSymbolLiteral` with the return value to place the literal in the desired location
*
* @param Op The named symbol to place
*
* @return A temporary `NamedSymbolLiteralPair`
*/
NamedSymbolLiteralPair InsertNamedSymbolLiteral(FEXCore::CPU::RelocNamedSymbolLiteral::NamedSymbol Op);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair &Lit);
/**
* @brief Place the named symbol literal relocation in memory
*
* @param Lit - Which literal to place
*/
void PlaceNamedSymbolLiteral(NamedSymbolLiteralPair& Lit);
fextl::vector<FEXCore::CPU::Relocation> Relocations;
fextl::vector<FEXCore::CPU::Relocation> Relocations;
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
///< Relocation code loading
bool ApplyRelocations(uint64_t GuestEntry, uint64_t CodeEntry, uint64_t CursorEntry, size_t NumRelocations, const char* EntryRelocations);
/** @} */
uint32_t SpillSlots{};
using OpType = void (Arm64JITCore::*)(IR::IROp_Header const *IROp, IR::NodeID Node);
uint32_t SpillSlots {};
using OpType = void (Arm64JITCore::*)(const IR::IROp_Header* IROp, IR::NodeID Node);
using ScalarBinaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2)>;
void VFScalarOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2);
void VFScalarOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit,
ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2);
using ScalarUnaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> SrcVar)>;
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
OpType RT_StoreMemTSO;
#define DEF_OP(x) void Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
// Dynamic Dispatcher supporting operations
DEF_OP(ParanoidLoadMemTSO);
File diff suppressed because it is too large. Load diff
@@ -10,14 +10,13 @@ $end_info$
#endif
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
#include <FEXCore/Core/SignalDelegator.h>
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(GuestOpcode) {
auto Op = IROp->C<IR::IROp_GuestOpcode>();
@@ -28,16 +27,10 @@ DEF_OP(GuestOpcode) {
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val:
dmb(FEXCore::ARMEmitter::BarrierScope::LD);
break;
case IR::Fence_LoadStore.Val:
dmb(FEXCore::ARMEmitter::BarrierScope::SY);
break;
case IR::Fence_Store.Val:
dmb(FEXCore::ARMEmitter::BarrierScope::ST);
break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
case IR::Fence_Load.Val: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::Fence_LoadStore.Val: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::Fence_Store.Val: dmb(ARMEmitter::BarrierScope::ST); break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
}
@@ -55,7 +48,7 @@ DEF_OP(Break) {
.err_code = Op->Reason.ErrorRegister,
};
uint64_t Constant{};
uint64_t Constant {};
memcpy(&Constant, &State, sizeof(State));
LoadConstant(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, Constant);
@@ -136,8 +129,7 @@ DEF_OP(Print) {
if (IsGPR(Op->Value.ID())) {
mov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetReg(Op->Value.ID()));
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintValue));
}
else {
} else {
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r0, GetVReg(Op->Value.ID()), false);
fmov(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::r1, GetVReg(Op->Value.ID()), true);
ldr(ARMEmitter::XReg::x3, STATE, offsetof(FEXCore::Core::CpuStateFrame, Pointers.Common.PrintVectorValue));
@@ -146,12 +138,10 @@ DEF_OP(Print) {
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
if (IsGPR(Op->Value.ID())) {
GenerateIndirectRuntimeCall<void, uint64_t>(ARMEmitter::Reg::r3);
}
else {
} else {
GenerateIndirectRuntimeCall<void, uint64_t, uint64_t>(ARMEmitter::Reg::r3);
}
}
else {
} else {
blr(ARMEmitter::Reg::r3);
}
@@ -231,8 +221,7 @@ DEF_OP(RDRAND) {
if (Op->GetReseeded) {
mrs(Dst.first, ARMEmitter::SystemRegister::RNDRRS);
}
else {
} else {
mrs(Dst.first, ARMEmitter::SystemRegister::RNDR);
}
@@ -245,5 +234,4 @@ DEF_OP(Yield) {
}
#undef DEF_OP
}
} // namespace FEXCore::CPU
@@ -8,15 +8,17 @@ $end_info$
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const *IROp, IR::NodeID Node)
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
LOGMAN_THROW_AA_FMT(Op->Header.Size == 4 || Op->Header.Size == 8, "Invalid size");
const auto EmitSize = Op->Header.Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Src = GetRegPair(Op->Pair.ID());
const std::array<ARMEmitter::Register, 2> Regs = {Src.first, Src.second};
mov(EmitSize, GetReg(Node), Regs[Op->Element]);
const auto Dst = GetReg(Node);
const auto Pair = GetRegPair(Op->Pair.ID());
const auto Src = Op->Element == 0 ? Pair.first : Pair.second;
if (Dst != Src) {
mov(ConvertSize48(IROp), Dst, Src);
}
}
DEF_OP(CreateElementPair) {
@@ -42,6 +44,25 @@ DEF_OP(CreateElementPair) {
}
}
#undef DEF_OP
DEF_OP(Copy) {
auto Op = IROp->C<IR::IROp_Copy>();
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Source.ID()));
}
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A.ID()), B = GetReg(Op->B.ID());
LOGMAN_THROW_AA_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
mov(ARMEmitter::Size::i64Bit, A, B);
mov(ARMEmitter::Size::i64Bit, B, TMP1);
}
DEF_OP(Swap2) {
// Implemented above
}
#undef DEF_OP
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
+3 -3
View File
@@ -1,7 +1,7 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <FEXCore/Core/CPUBackend.h>
#include "Interface/Core/CPUBackend.h"
#include <FEXCore/fextl/memory.h>
namespace FEXCore::Context {
@@ -15,8 +15,8 @@ struct InternalThreadState;
namespace FEXCore::CPU {
class CPUBackend;
[[nodiscard]] fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl *ctx,
FEXCore::Core::InternalThreadState *Thread);
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -13,8 +13,8 @@ $end_info$
#include "Interface/Core/LookupCache.h"
namespace FEXCore {
LookupCache::LookupCache(FEXCore::Context::ContextImpl *CTX)
: BlockLinks_mbr { fextl::pmr::get_default_resource() }
LookupCache::LookupCache(FEXCore::Context::ContextImpl* CTX)
: BlockLinks_mbr {fextl::pmr::get_default_resource()}
, ctx {CTX} {
TotalCacheSize = ctx->Config.VirtualMemSize / 4096 * 8 + CODE_SIZE + L1_SIZE;
@@ -78,5 +78,4 @@ void LookupCache::ClearCache() {
BlockList.clear();
}
}
} // namespace FEXCore
+58 -31
View File
@@ -13,6 +13,9 @@
#include <stddef.h>
#include <utility>
#include <mutex>
#ifdef _M_ARM_64EC
#include <winnt.h>
#endif
namespace FEXCore {
@@ -23,12 +26,12 @@ public:
uintptr_t GuestCode;
};
LookupCache(FEXCore::Context::ContextImpl *CTX);
LookupCache(FEXCore::Context::ContextImpl* CTX);
~LookupCache();
uintptr_t FindBlock(uint64_t Address) {
// Try L1, no lock needed
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
return L1Entry.HostCode;
}
@@ -37,7 +40,7 @@ public:
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Try L2
const auto PageIndex = (Address & (VirtualMemSize -1)) >> 12;
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
const auto PageOffset = Address & (0x0FFF);
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
@@ -48,8 +51,7 @@ public:
// Find there pointer for the address in the blocks
auto BlockPointers = reinterpret_cast<LookupCacheEntry*>(LocalPagePointer);
if (BlockPointers[PageOffset].GuestCode == Address)
{
if (BlockPointers[PageOffset].GuestCode == Address) {
L1Entry.GuestCode = Address;
L1Entry.HostCode = BlockPointers[PageOffset].HostCode;
return L1Entry.HostCode;
@@ -68,6 +70,24 @@ public:
return 0;
}
#ifdef _M_ARM_64EC
bool CheckPageEC(uint64_t Address) {
if (!RtlIsEcCode(Address)) {
return false;
}
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Mark L2 entry for this page as EC by setting the LSB, this can then be
// checked by the dispatcher to see if it needs to perform a call/return to
// EC code.
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
Pointers[PageIndex] |= 1;
return true;
}
#endif
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
// Appends Block {Address} to CodePages [Start, Start + Length)
@@ -77,8 +97,8 @@ public:
bool rv = false;
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length -1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto &CodePage = CodePages[CurrentPage];
for (auto CurrentPage = Start >> 12, EndPage = (Start + Length - 1) >> 12; CurrentPage <= EndPage; CurrentPage++) {
auto& CodePage = CodePages[CurrentPage];
rv |= CodePage.size() == 0;
CodePage.push_back(Address);
}
@@ -87,7 +107,7 @@ public:
}
// Adds to Guest -> Host code mapping
void AddBlockMapping(uint64_t Address, void *HostCode) {
void AddBlockMapping(uint64_t Address, void* HostCode) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
[[maybe_unused]] auto Inserted = BlockList.emplace(Address, (uintptr_t)HostCode).second;
@@ -95,18 +115,18 @@ public:
// There is no need to update L1 or L2, they will get updated on first lookup
// However, adding to L1 here increases performance
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = (uintptr_t)HostCode;
}
void Erase(FEXCore::Core::CpuStateFrame *Frame, uint64_t Address) {
void Erase(FEXCore::Core::CpuStateFrame* Frame, uint64_t Address) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Sever any links to this block
auto lower = BlockLinks->lower_bound({Address, nullptr});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData *>(UINTPTR_MAX)});
auto upper = BlockLinks->upper_bound({Address, reinterpret_cast<FEXCore::Context::ExitFunctionLinkData*>(UINTPTR_MAX)});
for (auto it = lower; it != upper; it = BlockLinks->erase(it)) {
it->second(Frame, it->first.HostLink);
}
@@ -115,7 +135,7 @@ public:
BlockList.erase(Address);
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
if (L1Entry.GuestCode == Address) {
L1Entry.GuestCode = 0;
// Leave L1Entry.HostCode as is, so that concurrent lookups won't read a null pointer
@@ -124,11 +144,11 @@ public:
}
// Do full map
Address = Address & (VirtualMemSize -1);
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t *Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// Page for this code didn't even exist, nothing to do
@@ -141,7 +161,7 @@ public:
BlockPointers[PageOffset].HostCode = 0;
}
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData * HostLink, const FEXCore::Context::BlockDelinkerFunc &delinker) {
void AddBlockLink(uint64_t GuestDestination, FEXCore::Context::ExitFunctionLinkData* HostLink, const FEXCore::Context::BlockDelinkerFunc& delinker) {
std::lock_guard<std::recursive_mutex> lk(WriteLock);
BlockLinks->insert({{GuestDestination, HostLink}, delinker});
@@ -150,14 +170,20 @@ public:
void ClearCache();
void ClearL2Cache();
uintptr_t GetL1Pointer() const { return L1Pointer; }
uintptr_t GetPagePointer() const { return PagePointer; }
uintptr_t GetVirtualMemorySize() const { return VirtualMemSize; }
uintptr_t GetL1Pointer() const {
return L1Pointer;
}
uintptr_t GetPagePointer() const {
return PagePointer;
}
uintptr_t GetVirtualMemorySize() const {
return VirtualMemSize;
}
constexpr static size_t L1_ENTRIES = 1 * 1024 * 1024; // Must be a power of 2
constexpr static size_t L1_ENTRIES_MASK = L1_ENTRIES - 1;
// This needs to be taken before reads or writes to L2, L3, CodePages, Thread::DebugStore,
// This needs to be taken before reads or writes to L2, L3, CodePages,
// and before writes to L1. Concurrent access from a thread that this LookupCache doesn't belong to
// may only happen during cross thread invalidation (::Erase).
// All other operations must be done from the owning thread.
@@ -169,17 +195,17 @@ public:
private:
void CacheBlockMapping(uint64_t Address, uintptr_t HostCode) {
// Do L1
auto &L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
auto& L1Entry = reinterpret_cast<LookupCacheEntry*>(L1Pointer)[Address & L1_ENTRIES_MASK];
L1Entry.GuestCode = Address;
L1Entry.HostCode = HostCode;
// Do ful map
auto FullAddress = Address;
Address = Address & (VirtualMemSize -1);
Address = Address & (VirtualMemSize - 1);
uint64_t PageOffset = Address & (0x0FFF);
Address >>= 12;
uintptr_t *Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uintptr_t* Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
uint64_t LocalPagePointer = Pointers[Address];
if (!LocalPagePointer) {
// We don't have a page pointer for this address
@@ -223,15 +249,16 @@ private:
struct BlockLinkTag {
uint64_t GuestDestination;
FEXCore::Context::ExitFunctionLinkData *HostLink;
FEXCore::Context::ExitFunctionLinkData* HostLink;
bool operator <(const BlockLinkTag& other) const {
if (GuestDestination < other.GuestDestination)
bool operator<(const BlockLinkTag& other) const {
if (GuestDestination < other.GuestDestination) {
return true;
else if (GuestDestination == other.GuestDestination)
} else if (GuestDestination == other.GuestDestination) {
return HostLink < other.HostLink;
else
} else {
return false;
}
}
};
@@ -244,7 +271,7 @@ private:
std::pmr::monotonic_buffer_resource BlockLinks_mbr;
using BlockLinksMapType = std::pmr::map<BlockLinkTag, FEXCore::Context::BlockDelinkerFunc>;
fextl::unique_ptr<std::pmr::polymorphic_allocator<std::byte>> BlockLinks_pma;
BlockLinksMapType *BlockLinks;
BlockLinksMapType* BlockLinks;
fextl::robin_map<uint64_t, uint64_t> BlockList;
@@ -256,7 +283,7 @@ private:
size_t AllocateOffset {};
FEXCore::Context::ContextImpl *ctx;
uint64_t VirtualMemSize{};
FEXCore::Context::ContextImpl* ctx;
uint64_t VirtualMemSize {};
};
}
} // namespace FEXCore
@@ -6,79 +6,81 @@
#include <cstdint>
namespace FEXCore::CodeSerialize {
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct
FEX_PACKED
CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie{};
// If any of the config options mismatch on load then the cache won't be used
// Any of these will result in codegen changes
struct FEX_PACKED CodeObjectSerializationConfig {
// Cookie in the header of the file, isn't part of the config hash
uint64_t Cookie {};
// Instructions per block configuration
int32_t MaxInstPerBlock{};
// Instructions per block configuration
int32_t MaxInstPerBlock {};
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Follows CPUID 4000_0001_EAX[3:0]
unsigned Arch : 4;
// Multiblock enabled
unsigned MultiBlock : 1;
// Multiblock enabled
unsigned MultiBlock : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// Hardware TSO enabled
unsigned HardwareTSOEnabled : 1;
// TSO enabled
unsigned TSOEnabled : 1;
// TSO enabled
unsigned TSOEnabled : 1;
// ABI local flag unsafe optimization
unsigned ABILocalFlags : 1;
// ABI local flag unsafe optimization
unsigned ABILocalFlags : 1;
// Paranoid TSO mode enabled
unsigned ParanoidTSO : 1;
// Paranoid TSO mode enabled
unsigned ParanoidTSO : 1;
// Guest code execution mode (We don't support live mode switch)
unsigned Is64BitMode : 1;
// Guest code execution mode (We don't support live mode switch)
unsigned Is64BitMode : 1;
// SMC checks style
unsigned SMCChecks : 2;
// SMC checks style
unsigned SMCChecks : 2;
// x87 reduced precision
unsigned x87ReducedPrecision : 1;
// x87 reduced precision
unsigned x87ReducedPrecision : 1;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 19;
// Padding to remove uninitialized data warning from asan
// Shows remaining amount of bits available for config
unsigned _Pad : 19;
bool operator==(CodeObjectSerializationConfig const &other) const {
return Cookie == other.Cookie &&
MaxInstPerBlock == other.MaxInstPerBlock &&
Arch == other.Arch &&
MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled &&
TSOEnabled == other.TSOEnabled &&
ABILocalFlags == other.ABILocalFlags &&
ParanoidTSO == other.ParanoidTSO &&
Is64BitMode == other.Is64BitMode &&
SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(CodeObjectSerializationConfig const &other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash{};
Hash <<= 32; Hash |= other.MaxInstPerBlock;
Hash <<= 1; Hash |= other.Arch;
Hash <<= 1; Hash |= other.MultiBlock;
Hash <<= 1; Hash |= other.HardwareTSOEnabled;
Hash <<= 1; Hash |= other.TSOEnabled;
Hash <<= 1; Hash |= other.ABILocalFlags;
Hash <<= 1; Hash |= other.ParanoidTSO;
Hash <<= 1; Hash |= other.Is64BitMode;
Hash <<= 2; Hash |= other.SMCChecks;
Hash <<= 1; Hash |= other.x87ReducedPrecision;
return Hash;
}
};
bool operator==(const CodeObjectSerializationConfig& other) const {
return Cookie == other.Cookie && MaxInstPerBlock == other.MaxInstPerBlock && Arch == other.Arch && MultiBlock == other.MultiBlock &&
HardwareTSOEnabled == other.HardwareTSOEnabled && TSOEnabled == other.TSOEnabled && ABILocalFlags == other.ABILocalFlags &&
ParanoidTSO == other.ParanoidTSO && Is64BitMode == other.Is64BitMode && SMCChecks == other.SMCChecks &&
x87ReducedPrecision == other.x87ReducedPrecision;
}
static uint64_t GetHash(const CodeObjectSerializationConfig& other) {
// For < 64-bits of data just pack directly
// Skip the cookie
uint64_t Hash {};
Hash <<= 32;
Hash |= other.MaxInstPerBlock;
Hash <<= 1;
Hash |= other.Arch;
Hash <<= 1;
Hash |= other.MultiBlock;
Hash <<= 1;
Hash |= other.HardwareTSOEnabled;
Hash <<= 1;
Hash |= other.TSOEnabled;
Hash <<= 1;
Hash |= other.ABILocalFlags;
Hash <<= 1;
Hash |= other.ParanoidTSO;
Hash <<= 1;
Hash |= other.Is64BitMode;
Hash <<= 2;
Hash |= other.SMCChecks;
Hash <<= 1;
Hash |= other.x87ReducedPrecision;
return Hash;
}
};
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is generated!");
}
static_assert(sizeof(CodeObjectSerializationConfig) == 16, "Size changed");
static_assert((sizeof(CodeObjectSerializationConfig) - sizeof(uint64_t)) == 8, "Config size exceeded 64its. Need to change how the hash is "
"generated!");
} // namespace FEXCore::CodeSerialize
@@ -11,120 +11,112 @@
#include <xxhash.h>
namespace FEXCore::CodeSerialize {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename) {
void AsyncJobHandler::AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
#ifndef _WIN32
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
// This function adds a named region *JOB* to our named region handler
// This needs to be as fast as possible to keep out of the way of the JIT
const fextl::string BaseFilename = FHU::Filesystem::GetFilename(filename);
const fextl::string BaseFilename = FHU::Filesystem::GetFilename(filename);
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = fextl::make_unique<CodeRegionEntry>(
Base,
Size,
Offset,
filename,
NamedRegionHandler->DefaultCodeHeader(Base, Offset)
);
if (!BaseFilename.empty()) {
// Create a new entry that once set up will be put in to our section object map
auto Entry = fextl::make_unique<CodeRegionEntry>(Base, Size, Offset, filename, NamedRegionHandler->DefaultCodeHeader(Base, Offset));
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
// Lock the job ref counter so we can block anything attempting to use the entry before it is loaded
Entry->NamedJobRefCountMutex.lock();
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
}
else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
#ifndef _WIN32
// Removing a named region through the job system
// We need to find the entry that we are deleting first
fextl::unique_ptr<CodeRegionEntry> EntryPointer;
CodeRegionMapType::iterator EntryIterator;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto &EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
// Take the pointer from the map
EntryPointer = std::move(it->second);
auto it = EntryMap.emplace(Base, std::move(Entry));
if (!it.second) {
// This happens when an application overwrites a previous region without unmapping what was there
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Lock this entry's Named job reference counter.
// Once this passes then we know that this section has been loaded.
it.first->second->NamedJobRefCountMutex.lock();
// Remove this from the entry map
EntryMap.erase(it);
// Finalize anything the region needs to do first.
CodeObjectCacheService->DoCodeRegionClosure(it.first->second->Base, it.first->second.get());
// munmap the file that was mapped
FEXCore::Allocator::munmap(it.first->second->CodeData, it.first->second->FileSize);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(it.first->second->EntryHeader.OriginalBase);
}
}
else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
// Now overwrite the entry in the map
it = EntryMap.insert_or_assign(Base, std::move(Entry));
EntryIterator = it.first;
} else {
// No overwrite, just insert
EntryIterator = it.first;
}
}
#endif
}
void AsyncJobHandler::AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
// Now that this entry has been added to the map, we can insert a load job using the entry iterator.
// This allows us to quickly unblock the JIT thread when it is loading multiple regions and have the async thread
// do the loading for us.
//
// Create the async work queue job now so it can load
NamedRegionHandler->AsyncAddNamedRegionWorkItem(BaseFilename, filename, true, EntryIterator);
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
#ifndef _WIN32
// Removing a named region through the job system
// We need to find the entry that we are deleting first
fextl::unique_ptr<CodeRegionEntry> EntryPointer;
{
std::unique_lock lk {CodeObjectCacheService->GetEntryMapMutex()};
auto& EntryMap = CodeObjectCacheService->GetEntryMap();
auto it = EntryMap.find(Base);
if (it != EntryMap.end()) {
// Lock the job ref counter since we are erasing it
// Once this passes it will have been loaded
it->second->NamedJobRefCountMutex.lock();
// Take the pointer from the map
EntryPointer = std::move(it->second);
// We can now unmap the file data
FEXCore::Allocator::munmap(EntryPointer->CodeData, EntryPointer->FileSize);
// Remove this from the entry map
EntryMap.erase(it);
// Remove this entry from the unrelocated map as well
{
std::unique_lock lk2 {CodeObjectCacheService->GetUnrelocatedEntryMapMutex()};
CodeObjectCacheService->GetUnrelocatedEntryMap().erase(EntryPointer->EntryHeader.OriginalBase);
}
} else {
// Tried to remove something that wasn't in our code object tracking
return;
}
// Create the async work queue job now so it can finalize what it needs to do
NamedRegionHandler->AsyncRemoveNamedRegionWorkItem(Base, Size, std::move(EntryPointer));
// Tell the async thread that it has work to do
CodeObjectCacheService->NotifyWork();
}
#endif
}
void AsyncJobHandler::AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data) {
// XXX: Actually add serialization job
}
} // namespace FEXCore::CodeSerialize
@@ -7,66 +7,67 @@
#include <FEXCore/fextl/string.h>
namespace FEXCore::CodeSerialize {
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
NamedRegionObjectHandler::NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx) {
DefaultSerializationConfig.Cookie = CODE_COOKIE;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
// Initialize the Arch from CPUID
uint32_t Arch = ctx->CPUID.RunFunction(0x4000'0001, 0).eax & 0xF;
DefaultSerializationConfig.Arch = Arch;
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
DefaultSerializationConfig.MaxInstPerBlock = ctx->Config.MaxInstPerBlock;
DefaultSerializationConfig.MultiBlock = ctx->Config.Multiblock;
DefaultSerializationConfig.TSOEnabled = ctx->Config.TSOEnabled;
DefaultSerializationConfig.ABILocalFlags = ctx->Config.ABILocalFlags;
DefaultSerializationConfig.ParanoidTSO = ctx->Config.ParanoidTSO;
DefaultSerializationConfig.Is64BitMode = ctx->Config.Is64BitMode;
DefaultSerializationConfig.SMCChecks = ctx->Config.SMCChecks;
DefaultSerializationConfig.x87ReducedPrecision = ctx->Config.x87ReducedPrecision;
}
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string &base_filename, const fextl::string &filename, bool Executable) {
// XXX: Add named region objects
void NamedRegionObjectHandler::AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename,
const fextl::string& filename, bool Executable) {
// XXX: Add named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
// XXX: Until entry loading is complete just claim it is loaded
Entry->second->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
void NamedRegionObjectHandler::RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
// XXX: Remove named region objects
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
// XXX: Until entry loading is complete just claim it is loaded
Entry->NamedJobRefCountMutex.unlock();
}
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
void NamedRegionObjectHandler::HandleNamedRegionObjectJobs() {
// Walk through all of our jobs sequentially until the work queue is empty
while (NamedWorkQueueJobs.load()) {
fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem> WorkItem;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
{
// Lock the work queue mutex for a short moment and grab an item from the list
std::unique_lock lk {NamedWorkQueueMutex};
size_t WorkItems = WorkQueue.size();
if (WorkItems != 0) {
WorkItem = std::move(WorkQueue.front());
WorkQueue.pop();
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion *>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
// Atomically update the number of jobs
--NamedWorkQueueJobs;
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion *>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
if (WorkItem) {
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_ADD_NAMED_REGION) {
auto WorkAdd = static_cast<AsyncJobHandler::WorkItemAddNamedRegion*>(WorkItem.get());
AddNamedRegionObject(WorkAdd->Entry, WorkAdd->BaseFilename, WorkAdd->Filename, WorkAdd->Executable);
}
if (WorkItem->GetType() == AsyncJobHandler::NamedRegionJobType::JOB_REMOVE_NAMED_REGION) {
auto WorkRemove = static_cast<AsyncJobHandler::WorkItemRemoveNamedRegion*>(WorkItem.get());
RemoveNamedRegionObject(WorkRemove->Base, WorkRemove->Size, std::move(WorkRemove->Entry));
}
}
}
}
} // namespace FEXCore::CodeSerialize
@@ -6,80 +6,80 @@
#include <FEXCore/Utils/Threads.h>
namespace {
static void* ThreadHandler(void *Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService *This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
static void* ThreadHandler(void* Arg) {
FEXCore::CodeSerialize::CodeObjectSerializeService* This = reinterpret_cast<FEXCore::CodeSerialize::CodeObjectSerializeService*>(Arg);
This->ExecutionThread();
return nullptr;
}
} // namespace
namespace FEXCore::CodeSerialize {
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx)
: CTX {ctx}
, AsyncHandler { &NamedRegionHandler , this }
, NamedRegionHandler { ctx } {
Initialize();
CodeObjectSerializeService::CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx)
: CTX {ctx}
, AsyncHandler {&NamedRegionHandler, this}
, NamedRegionHandler {ctx} {
Initialize();
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
void CodeObjectSerializeService::Shutdown() {
if (CTX->Config.CacheObjectCodeCompilation() == FEXCore::Config::ConfigObjectCodeHandler::CONFIG_NONE) {
return;
}
WorkerThreadShuttingDown = true;
WorkerThreadShuttingDown = true;
// Kick the working thread
WorkAvailable.NotifyAll();
// Kick the working thread
WorkAvailable.NotifyAll();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, fextl::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
CodeObjectFileSection const *CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
FEXCore::Threads::SetThreadName("ObjectCodeSeri\0");
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto &it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
if (WorkerThread->joinable()) {
// Wait for worker thread to close down
WorkerThread->join(nullptr);
}
}
void CodeObjectSerializeService::Initialize() {
// Add a canary so we don't crash on empty map iterator handling
auto it = AddressToEntryMap.insert_or_assign(~0ULL, fextl::make_unique<CodeRegionEntry>());
UnrelocatedAddressToEntryMap.insert_or_assign(~0ULL, it.first->second.get());
uint64_t OldMask = FEXCore::Threads::SetSignalMask(~0ULL);
WorkerThread = FEXCore::Threads::Thread::Create(ThreadHandler, this);
FEXCore::Threads::SetSignalMask(OldMask);
}
void CodeObjectSerializeService::DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it) {
if (Base == ~0ULL) {
// Don't do closure on canary
return;
}
// XXX: Do code region closure
}
const CodeObjectFileSection* CodeObjectSerializeService::FetchCodeObjectFromCache(uint64_t GuestRIP) {
// XXX: Actually fetch code objects from cache
return nullptr;
}
void CodeObjectSerializeService::ExecutionThread() {
// Set our thread name so we can see its relation
FEXCore::Threads::SetThreadName("ObjectCodeSeri\0");
while (WorkerThreadShuttingDown.load() != true) {
// Wait for work
WorkAvailable.Wait();
// Handle named region async jobs first. Highest priority
NamedRegionHandler.HandleNamedRegionObjectJobs();
// XXX: Handle code serialization jobs second.
}
// Do final code region closures on thread shutdown
for (auto& it : AddressToEntryMap) {
DoCodeRegionClosure(it.first, it.second.get());
}
// Safely clear our maps now
AddressToEntryMap.clear();
UnrelocatedAddressToEntryMap.clear();
}
} // namespace FEXCore::CodeSerialize
@@ -17,445 +17,441 @@
#include <shared_mutex>
namespace FEXCore::CodeSerialize {
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {
};
// XXX: Does this need to be signal safe?
using CodeSerializationMutex = std::shared_mutex;
struct CodeSerializationData {};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData *Data;
const char *HostCode;
uint64_t NumRelocations;
const char *Relocations;
};
struct CodeObjectFileSection {
bool Serialized;
bool Invalid;
const CodeSerializationData* Data;
const char* HostCode;
uint64_t NumRelocations;
const char* Relocations;
};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase {};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset {};
// Total amount of code that should be in this file
uint64_t TotalCodeSize {};
// Used to reserve the TSL map
uint64_t NumCodeEntries {};
// The number of relocations that point to this section
uint64_t NumRelocationsTo {};
// Total relocations in this file
uint64_t TotalRelocationsCount {};
};
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base {};
// Size of this code entry
uint64_t Size {};
// The offset inside the file that is mapped to Base
uint64_t Offset {};
// Filename of the object
fextl::string Filename {};
CodeObjectSerializationHeader EntryHeader {};
/** @} */
// The filename of the object cache for this entry
fextl::string ObjectEntrySourceFilename {};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @brief This is the file header that lives at the start of an object cache file
*
* This header is updated from multiple processes!
* Care must be taken to use OS locks when updating the file backing including this header
*/
struct CodeObjectSerializationHeader {
// The configuration that this file has
CodeObjectSerializationConfig Config;
// The original RIP that this object section was mapped at
uint64_t OriginalBase{};
// The original offset in to the file that this object section was loaded from
uint64_t OriginalOffset{};
// Total amount of code that should be in this file
uint64_t TotalCodeSize{};
// Used to reserve the TSL map
uint64_t NumCodeEntries{};
// The number of relocations that point to this section
uint64_t NumRelocationsTo{};
// Total relocations in this file
uint64_t TotalRelocationsCount{};
};
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
struct CodeRegionEntry {
/**
* @name Threaded initialization objects for the initial object creation
* @{ */
// Base address in memory where the code region is at
uint64_t Base{};
// Size of this code entry
uint64_t Size{};
// The offset inside the file that is mapped to Base
uint64_t Offset{};
// Filename of the object
fextl::string Filename{};
CodeObjectSerializationHeader EntryHeader{};
/** @} */
// The filename of the object cache for this entry
fextl::string ObjectEntrySourceFilename{};
// In the case of file corruption that we can detect, we can disable serialization early for an entry
// We should be resiliant to corruption but things happen
bool StillSerializing {true};
// Long lived FD for serialization if we have multiple jobs to serialize
// Bursts of code entries are common and this reduces file lock overhead
//
// Especially useful over network mounts where file locks are very slow
int CurrentSerializedFD {-1};
/**
* @name Objects required to sync objects between threads
* @{ */
// Refcount for the number of outstanding code entries waiting to be written for this object section
CodeSerializationMutex ObjectJobRefCountMutex;
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char *CodeData{};
size_t FileSize{};
fextl::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
fextl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap{};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base,
uint64_t Size,
uint64_t Offset,
fextl::string const &Filename,
CodeObjectSerializationHeader const &DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {
}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = fextl::map<uint64_t, fextl::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = fextl::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
void *HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex *ThreadJobRefCount;
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex *ObjectJobRefCountMutexPtr;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
AsyncJobHandler(NamedRegionObjectHandler *NamedRegionHandler, CodeObjectSerializeService *CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const { return Type; }
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const fextl::string &base, const fextl::string &filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry}
{}
const fextl::string BaseFilename;
const fextl::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, fextl::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
fextl::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler *NamedRegionHandler;
CodeObjectSerializeService *CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::ContextImpl *ctx);
void HandleNamedRegionObjectJobs();
CodeObjectSerializationConfig const &GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const fextl::string &base, const fextl::string &filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemAddNamedRegion> (
base,
filename,
executable,
entry
));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion> (
Base,
Size,
std::move(Entry)
));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs{};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex{};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
fextl::queue<fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue{};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string &base_filename, const fextl::string &filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
// Refcount for outstanding named object region entry loading itself
// Will block JIT code cache look up when this has a unique_lock held
CodeSerializationMutex NamedJobRefCountMutex;
/** @} */
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
* @name Object Entry data management
* @{ */
/**
* @name This is the raw file data that we loaded from the code region entry file
* @{ */
char* CodeData {};
size_t FileSize {};
fextl::vector<CodeObjectFileSection> FileCodeSections;
/** @} */
// This per section map takes the most time to load and needs to be quick
// This is the map of all code segments for this entry
fextl::robin_map<uint64_t, CodeObjectFileSection*> SectionLookupMap {};
/** @} */
// Default initialization
CodeRegionEntry() = default;
// Initializer specifically for threaded loading
CodeRegionEntry(uint64_t Base, uint64_t Size, uint64_t Offset, const fextl::string& Filename, const CodeObjectSerializationHeader& DefaultHeader)
: Base {Base}
, Size {Size}
, Offset {Offset}
, Filename {Filename}
, EntryHeader {DefaultHeader} {}
};
// Map type must use an interator that isn't invalidation on erase/insert
using CodeRegionMapType = fextl::map<uint64_t, fextl::unique_ptr<CodeRegionEntry>>;
using CodeRegionPtrMapType = fextl::map<uint64_t, CodeRegionEntry*>;
class NamedRegionObjectHandler;
class CodeObjectSerializeService;
class AsyncJobHandler final {
public:
/**
* @brief Structure containing all the data required to async serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::ContextImpl *ctx);
struct SerializationJobData {
uint64_t GuestRIP; ///< The RIP for the guest
// XXX: Support multiblock
uint64_t GuestCodeLength; ///< The Guest's code length
uint64_t GuestCodeHash; ///< Hash of the guest code
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
void* HostCodeBegin; ///< Host JIT code starting memory address
size_t HostCodeLength; ///< Host JIT code length
uint64_t HostCodeHash; ///< Host JIT code hash before any backpatching
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
// This is the thread specific ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a thread is shutting down or clearing code cache then the thread will pull a unique lock on this mutex.
// This way it will wait until the async job handler is complete with it.
CodeSerializationMutex* ThreadJobRefCount;
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string &filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
// These are the reolocations for this serialization job
// Relatively small number of entries most of the time
fextl::vector<FEXCore::CPU::Relocation> Relocations;
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @name Objects filled in from the Code Object Serialization service when a job is added
* @{ */
// This is the code region's ref counter for outstanding jobs.
// This shared mutex is incremented when the job is added, then decremented when the job is complete.
// If a named region is being removed then a unique lock will be pulled to wait for all jobs to complete and no new jobs to be added.
CodeSerializationMutex* ObjectJobRefCountMutexPtr;
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(fextl::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex *ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
CodeObjectFileSection const *FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry *it);
CodeSerializationMutex &GetEntryMapMutex() { return EntryMapMutex; }
CodeSerializationMutex &GetUnrelocatedEntryMapMutex() { return EntryMapMutex; }
CodeRegionMapType &GetEntryMap() { return AddressToEntryMap; }
CodeRegionPtrMapType &GetUnrelocatedEntryMap() { return UnrelocatedAddressToEntryMap; }
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() { WorkAvailable.NotifyOne(); }
private:
FEXCore::Context::ContextImpl *CTX;
Event WorkAvailable{};
fextl::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
// This is the code region iterator to reduce the number of map lookups
// This will remain valid while jobs are outstanding for this region
CodeRegionMapType::iterator CodeRegionIterator;
/** @} */
};
}
AsyncJobHandler(NamedRegionObjectHandler* NamedRegionHandler, CodeObjectSerializeService* CodeObjectCacheService)
: NamedRegionHandler {NamedRegionHandler}
, CodeObjectCacheService {CodeObjectCacheService} {}
protected:
friend class CodeObjectSerializeService;
friend class NamedRegionObjectHandler;
/**
* @name Async job submission functions
* @{ */
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename);
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size);
void AsyncAddSerializationJob(fextl::unique_ptr<SerializationJobData> Data);
/** @} */
/**
* @name Async named region handling
* @{ */
/**
* @brief The async named region jobs to handle.
*
* Only two, Code serialization goes in to a different queue.
*/
enum class NamedRegionJobType {
JOB_ADD_NAMED_REGION,
JOB_REMOVE_NAMED_REGION,
};
class NamedRegionWorkItem {
public:
NamedRegionJobType GetType() const {
return Type;
}
protected:
friend class WorkItemAddNamedRegion;
NamedRegionWorkItem(NamedRegionJobType type)
: Type {type} {}
private:
NamedRegionJobType Type;
};
class WorkItemAddNamedRegion : public NamedRegionWorkItem {
public:
WorkItemAddNamedRegion(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_ADD_NAMED_REGION}
, BaseFilename {base}
, Filename {filename}
, Executable {executable}
, Entry {entry} {}
const fextl::string BaseFilename;
const fextl::string Filename;
bool Executable;
CodeRegionMapType::iterator Entry;
};
class WorkItemRemoveNamedRegion : public NamedRegionWorkItem {
public:
WorkItemRemoveNamedRegion(uint64_t base, uint64_t size, fextl::unique_ptr<CodeRegionEntry> entry)
: NamedRegionWorkItem {NamedRegionJobType::JOB_REMOVE_NAMED_REGION}
, Base {base}
, Size {size}
, Entry {std::move(entry)} {}
uint64_t Base;
uint64_t Size;
fextl::unique_ptr<CodeRegionEntry> Entry;
};
/** @} */
private:
NamedRegionObjectHandler* NamedRegionHandler;
CodeObjectSerializeService* CodeObjectCacheService;
};
class NamedRegionObjectHandler final {
public:
NamedRegionObjectHandler(FEXCore::Context::ContextImpl* ctx);
void HandleNamedRegionObjectJobs();
const CodeObjectSerializationConfig& GetDefaultSerializationConfig() const {
return DefaultSerializationConfig;
}
protected:
friend class AsyncJobHandler;
// Return a default code header based off the default serialization config
CodeObjectSerializationHeader DefaultCodeHeader(uint64_t Base, uint64_t Offset) const {
return CodeObjectSerializationHeader {
.Config = DefaultSerializationConfig,
.OriginalBase = Base,
.OriginalOffset = Offset,
.NumCodeEntries = 0,
.NumRelocationsTo = 0,
.TotalRelocationsCount = 0,
};
}
/**
* @brief Adds an asynchronous add named region work item to the object queue
*
* This adds the job that will do the loading of file resources and data tracking.
*/
void AsyncAddNamedRegionWorkItem(const fextl::string& base, const fextl::string& filename, bool executable, CodeRegionMapType::iterator entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemAddNamedRegion>(base, filename, executable, entry));
++NamedWorkQueueJobs;
}
void AsyncRemoveNamedRegionWorkItem(uint64_t Base, uint64_t Size, fextl::unique_ptr<CodeRegionEntry> Entry) {
std::unique_lock lk {NamedWorkQueueMutex};
WorkQueue.emplace(fextl::make_unique<AsyncJobHandler::WorkItemRemoveNamedRegion>(Base, Size, std::move(Entry)));
++NamedWorkQueueJobs;
}
private:
// Code version. If the code emission changes then this needs to increment
constexpr static uint32_t CODE_VERSION = 0x0;
// Default cookie header for the file header
constexpr static uint64_t CODE_COOKIE = FEXCore::IR::COOKIE_VERSION("FEXC", CODE_VERSION);
// Code serialization config for our current process configuration
CodeObjectSerializationConfig DefaultSerializationConfig;
// Atomic counter for number of jobs in the queue without needing to pull the mutex to check
std::atomic<uint64_t> NamedWorkQueueJobs {};
// Mutex for ading new jobs to the work queue
std::mutex NamedWorkQueueMutex {};
// The job queue itself
// Jobs get consumed as a FIFO
// Jobs always get appended to the end
fextl::queue<fextl::unique_ptr<AsyncJobHandler::NamedRegionWorkItem>> WorkQueue {};
/**
* @name Named Region object handling
* @{ */
void AddNamedRegionObject(CodeRegionMapType::iterator Entry, const fextl::string& base_filename, const fextl::string& filename, bool Executable);
void RemoveNamedRegionObject(uintptr_t Base, uintptr_t Size, fextl::unique_ptr<CodeRegionEntry> Entry);
/** @} */
};
/**
* @brief Context specific code object serialization class
*
* Contains everything required for FEXCore to serialize code objects
*/
class CodeObjectSerializeService final {
public:
CodeObjectSerializeService(FEXCore::Context::ContextImpl* ctx);
/**
* @brief Initialize the internal interface
*
* Is a public interface to allow the service to reinitialize after forking
*/
void Initialize();
/**
* @brief Safely shut down the Code Object serialization service.
*
* This service needs to be resiliant to application crashes, but shutting down safely is still preferred.
*/
void Shutdown();
/**
* @name Async interface
* @{ */
/**
* @brief Loads a named region in to the code serialization service. As async as possible.
*
* @param Base - Virtual address that this named region is loaded
* @param Size - The size of the region
* @param Offset - The offset from the file
* @param filename - The filename itself
*/
void AsyncAddNamedRegionJob(uintptr_t Base, uintptr_t Size, uintptr_t Offset, const fextl::string& filename) {
AsyncHandler.AsyncAddNamedRegionJob(Base, Size, Offset, filename);
}
/**
* @brief Unloads a named region from the code serialization service. As async as possible.
*
* @param Base - Virtual address of the named region
* @param Size - The size of the region
*/
void AsyncRemoveNamedRegionJob(uintptr_t Base, uintptr_t Size) {
AsyncHandler.AsyncRemoveNamedRegionJob(Base, Size);
}
/**
* @brief Adds a code object serialization job. As async as possible.
* Code hashing happens prior to async job serialization to catch invalidations due to backpatching.
*
* @param Data - A fully filled out struct containing all the code serialization
*/
void AsyncAddSerializationJob(fextl::unique_ptr<AsyncJobHandler::SerializationJobData> Data) {
AsyncHandler.AsyncAddSerializationJob(std::move(Data));
}
/** @} */
/**
* @name Synchronous interface
* @{ */
/**
* @brief Synchronously waits for this thread's job queue to become empty.
*
* This is necessary for when a thread is shutting down
*
* @param ThreadJobRefCount - The shared mutex to wait on until to be empty
*/
static void WaitForEmptyJobQueue(CodeSerializationMutex* ThreadJobRefCount) {
// Once the shared mutex is empty this unique lock will be gained
std::unique_lock lk {*ThreadJobRefCount};
}
/**
* @brief Fetches object code from the Code Object Cache for JIT.
*
* @param GuestRIP - Which GuestRIP to search the cache for
*
* @return Data required for the JIT to relocate the Object code.
*/
const CodeObjectFileSection* FetchCodeObjectFromCache(uint64_t GuestRIP);
/** @} */
// Public for threading
void ExecutionThread();
protected:
friend class AsyncJobHandler;
/**
* @brief Safely closes out code object regions from the map
*
* @param it - iterator to do a closure on
*/
void DoCodeRegionClosure(uint64_t Base, CodeRegionEntry* it);
CodeSerializationMutex& GetEntryMapMutex() {
return EntryMapMutex;
}
CodeSerializationMutex& GetUnrelocatedEntryMapMutex() {
return EntryMapMutex;
}
CodeRegionMapType& GetEntryMap() {
return AddressToEntryMap;
}
CodeRegionPtrMapType& GetUnrelocatedEntryMap() {
return UnrelocatedAddressToEntryMap;
}
/**
* @brief Notify the async thread that it has work to do
*/
void NotifyWork() {
WorkAvailable.NotifyOne();
}
private:
FEXCore::Context::ContextImpl* CTX;
Event WorkAvailable {};
fextl::unique_ptr<FEXCore::Threads::Thread> WorkerThread;
std::atomic_bool WorkerThreadShuttingDown {false};
AsyncJobHandler AsyncHandler;
NamedRegionObjectHandler NamedRegionHandler;
// Mutex to hold when modifying the entry maps
CodeSerializationMutex EntryMapMutex;
CodeSerializationMutex UnrelocatedEntryMapMutex;
// Entry maps
CodeRegionMapType AddressToEntryMap;
CodeRegionPtrMapType UnrelocatedAddressToEntryMap;
};
} // namespace FEXCore::CodeSerialize
@@ -3,77 +3,77 @@
#include <FEXCore/IR/IR.h>
namespace FEXCore::CPU {
enum class RelocationTypes : uint8_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
enum class RelocationTypes : uint8_t {
// 8 byte literal in memory for symbol
// Aligned to struct RelocNamedSymbolLiteral
RELOC_NAMED_SYMBOL_LITERAL,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// Fixed size named thunk move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocNamedThunkMove
RELOC_NAMED_THUNK_MOVE,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
RELOC_GUEST_RIP_MOVE,
// Fixed size guest RIP move
// 4 instruction constant generation on AArch64
// 64-bit mov on x86-64
// Aligned to struct RelocGuestRIPMove
RELOC_GUEST_RIP_MOVE,
};
struct RelocationTypeHeader final {
RelocationTypes Type;
};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
struct RelocationTypeHeader final {
RelocationTypes Type;
};
RelocationTypeHeader Header {};
struct RelocNamedSymbolLiteral final {
enum class NamedSymbol : uint8_t {
///< Thread specific relocations
// JIT Literal pointers
SYMBOL_LITERAL_EXITFUNCTION_LINKER,
};
NamedSymbol Symbol;
RelocationTypeHeader Header{};
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
NamedSymbol Symbol;
struct RelocNamedThunkMove final {
RelocationTypeHeader Header {};
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
struct RelocNamedThunkMove final {
RelocationTypeHeader Header{};
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// Offset in to the code section to begin the relocation
uint64_t Offset {};
};
// The thunk SHA256 hash
IR::SHA256Sum Symbol;
struct RelocGuestRIPMove final {
RelocationTypeHeader Header {};
// Offset in to the code section to begin the relocation
uint64_t Offset{};
};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
struct RelocGuestRIPMove final {
RelocationTypeHeader Header{};
// Offset in to the code section to begin the relocation
uint64_t Offset {};
// GPR index the constant is being moved to
uint8_t RegisterIndex;
// The unrelocated RIP that is being moved
uint64_t GuestRIP;
};
// Offset in to the code section to begin the relocation
uint64_t Offset{};
union Relocation {
RelocationTypeHeader Header {};
// The unrelocated RIP that is being moved
uint64_t GuestRIP;
};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
union Relocation {
RelocationTypeHeader Header{};
RelocNamedSymbolLiteral NamedSymbolLiteral;
// This makes our union of relocations at least 48 bytes
// It might be more efficient to not use a union
RelocNamedThunkMove NamedThunkMove;
RelocGuestRIPMove GuestRIPMove;
};
}
RelocGuestRIPMove GuestRIPMove;
};
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -22,10 +22,10 @@ class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *RotatedNode{};
Ref RotatedNode {};
if (CTX->HostFeatures.SupportsSHA) {
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
@@ -34,8 +34,7 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
auto Duplicated = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Dest, 3);
auto Sha1HRotated = _VSha1H(Duplicated);
RotatedNode = _VDupElement(OpSize::i128Bit, OpSize::i32Bit, Sha1HRotated, 0);
}
else {
} else {
// SHA1 extension missing, manually rotate.
// Emulate rotate.
auto ShiftLeft = _VShlI(OpSize::i128Bit, OpSize::i32Bit, Dest, 30);
@@ -48,20 +47,20 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
}
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *NewVec = _VExtr(16, 8, Dest, Src, 1);
Ref NewVec = _VExtr(16, 8, Dest, Src, 1);
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
OrderedNode *Result = _VXor(16, 1, Dest, NewVec);
Ref Result = _VXor(16, 1, Dest, NewVec);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
// This instruction mostly matches ARMv8's SHA1SU1 instruction but one of the elements are flipped in an unexpected way.
// Do all the work without it.
@@ -91,45 +90,47 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
}
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(),
"Src1 needs to be literal here to indicate function and constants");
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here to indicate function and constants");
using FnType = OrderedNode* (*)(OpDispatchBuilder&, OrderedNode*, OrderedNode*, OrderedNode*);
using FnType = Ref (*)(OpDispatchBuilder&, Ref, Ref, Ref);
const auto f0 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
return Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._Andn(OpSize::i32Bit, D, B));
};
const auto f1 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
const auto f2 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
return Self.BitwiseAtLeastTwo(B, C, D);
};
const auto f3 = [](OpDispatchBuilder &Self, OrderedNode *B, OrderedNode *C, OrderedNode *D) -> OrderedNode* {
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
};
constexpr std::array<uint32_t, 4> k_array{
constexpr std::array<uint32_t, 4> k_array {
0x5A827999U,
0x6ED9EBA1U,
0x8F1BBCDCU,
0xCA62C1D6U,
};
constexpr std::array<FnType, 4> fn_array{
f0, f1, f2, f3,
constexpr std::array<FnType, 4> fn_array {
f0,
f1,
f2,
f3,
};
const uint64_t Imm8 = Op->Src[1].Data.Literal.Value & 0b11;
const FnType Fn = fn_array[Imm8];
auto K = _Constant(32, k_array[Imm8]);
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
auto W0E = _VExtractToGPR(16, 4, Src, 3);
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
const auto Round0 = [&]() -> RoundResult {
auto A = _VExtractToGPR(16, 4, Dest, 3);
@@ -137,7 +138,8 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
auto C = _VExtractToGPR(16, 4, Dest, 1);
auto D = _VExtractToGPR(16, 4, Dest, 0);
auto A1 = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), W0E), K);
auto A1 =
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), W0E), K);
auto B1 = A;
auto C1 = _Ror(OpSize::i32Bit, B, _Constant(32, 2));
auto D1 = C;
@@ -145,13 +147,13 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
return {A1, B1, C1, D1, E1};
};
const auto Round1To3 = [&](OrderedNode *A, OrderedNode *B, OrderedNode *C,
OrderedNode *D, OrderedNode *E, OrderedNode *Src, unsigned W_idx) -> RoundResult {
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
// Kill W and E at the beginning
auto W = _VExtractToGPR(16, 4, Src, W_idx);
auto Q = _Add(OpSize::i32Bit, W, E);
auto ANext = _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), Q), K);
auto ANext =
_Add(OpSize::i32Bit, _Add(OpSize::i32Bit, _Add(OpSize::i32Bit, Fn(*this, B, C, D), _Ror(OpSize::i32Bit, A, _Constant(32, 27))), Q), K);
auto BNext = A;
auto CNext = _Ror(OpSize::i32Bit, B, _Constant(32, 2));
auto DNext = C;
@@ -163,9 +165,9 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
auto [A1, B1, C1, D1, E1] = Round0();
auto [A2, B2, C2, D2, E2] = Round1To3(A1, B1, C1, D1, E1, Src, 2);
auto [A3, B3, C3, D3, E3] = Round1To3(A2, B2, C2, D2, E2, Src, 1);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Final = Round1To3(A3, B3, C3, D3, E3, Src, 0);
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest3 = _VInsGPR(16, 4, 3, Dest, std::get<0>(Final));
auto Dest2 = _VInsGPR(16, 4, 2, Dest3, std::get<1>(Final));
auto Dest1 = _VInsGPR(16, 4, 1, Dest2, std::get<2>(Final));
auto Dest0 = _VInsGPR(16, 4, 0, Dest1, std::get<3>(Final));
@@ -174,17 +176,17 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
}
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Result{};
Ref Result {};
if (CTX->HostFeatures.SupportsSHA) {
Result = _VSha256U0(Dest, Src);
}
else {
const auto Sigma0 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 7)), _Ror(OpSize::i32Bit, W, _Constant(32, 18))), _Lshr(OpSize::i32Bit, W, _Constant(32, 3)));
} else {
const auto Sigma0 = [this](Ref W) -> Ref {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 7)), _Ror(OpSize::i32Bit, W, _Constant(32, 18))),
_Lshr(OpSize::i32Bit, W, _Constant(32, 3)));
};
auto W4 = _VExtractToGPR(16, 4, Src, 0);
@@ -208,12 +210,13 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
}
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
const auto Sigma1 = [this](OrderedNode* W) -> OrderedNode* {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 17)), _Ror(OpSize::i32Bit, W, _Constant(32, 19))), _Lshr(OpSize::i32Bit, W, _Constant(32, 10)));
const auto Sigma1 = [this](Ref W) -> Ref {
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 17)), _Ror(OpSize::i32Bit, W, _Constant(32, 19))),
_Lshr(OpSize::i32Bit, W, _Constant(32, 10)));
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
auto W14 = _VExtractToGPR(16, 4, Src, 2);
auto W15 = _VExtractToGPR(16, 4, Src, 3);
@@ -230,36 +233,38 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
StoreResult(FPRClass, Op, D0, -1);
}
OrderedNode *OpDispatchBuilder::BitwiseAtLeastTwo(OrderedNode *A, OrderedNode *B, OrderedNode *C) {
// Returns whether at least 2/3 of A/B/C is true.
// Expressed as (A & (B | C)) | (B & C)
//
// Equivalent to expression in SHA calculations: (A & B) ^ (A & C) ^ (B & C)
auto And = _And(OpSize::i32Bit, B, C);
auto Or = _Or(OpSize::i32Bit, B, C);
return _Or(OpSize::i32Bit, _And(OpSize::i32Bit, A, Or), And);
Ref OpDispatchBuilder::BitwiseAtLeastTwo(Ref A, Ref B, Ref C) {
// Returns whether at least 2/3 of A/B/C is true.
// Expressed as (A & (B | C)) | (B & C)
//
// Equivalent to expression in SHA calculations: (A & B) ^ (A & C) ^ (B & C)
auto And = _And(OpSize::i32Bit, B, C);
auto Or = _Or(OpSize::i32Bit, B, C);
return _Or(OpSize::i32Bit, _And(OpSize::i32Bit, A, Or), And);
}
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
const auto Ch = [this](OrderedNode *E, OrderedNode *F, OrderedNode *G) -> OrderedNode* {
const auto Ch = [this](Ref E, Ref F, Ref G) -> Ref {
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
};
const auto Sigma0 = [this](OrderedNode *A) -> OrderedNode* {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), A, ShiftType::ROR, 13), A, ShiftType::ROR, 22);
const auto Sigma0 = [this](Ref A) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), A, ShiftType::ROR, 13), A,
ShiftType::ROR, 22);
};
const auto Sigma1 = [this](OrderedNode *E) -> OrderedNode* {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), E, ShiftType::ROR, 11), E, ShiftType::ROR, 25);
const auto Sigma1 = [this](Ref E) -> Ref {
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), E, ShiftType::ROR, 11), E,
ShiftType::ROR, 25);
};
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
// Hardcoded to XMM0
auto XMM0 = LoadXMMRegister(0);
auto E0 = _VExtractToGPR(16, 4, Src, 1);
auto F0 = _VExtractToGPR(16, 4, Src, 0);
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
OrderedNode *Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
Ref Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
@@ -275,7 +280,7 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
OrderedNode * Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
Ref Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
@@ -299,16 +304,16 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
}
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Result = _VAESImc(Src);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Result = _VAESImc(Src);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESEnc(16, Dest, Src, ZeroRegister);
Ref Result = _VAESEnc(16, Dest, Src, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -319,19 +324,19 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESENC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESEnc(DstSize, State, Key, ZeroRegister);
Ref Result = _VAESEnc(DstSize, State, Key, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESEncLast(16, Dest, Src, ZeroRegister);
Ref Result = _VAESEncLast(16, Dest, Src, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -342,19 +347,19 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESENCLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESEncLast(DstSize, State, Key, ZeroRegister);
Ref Result = _VAESEncLast(DstSize, State, Key, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESDec(16, Dest, Src, ZeroRegister);
Ref Result = _VAESDec(16, Dest, Src, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -365,19 +370,19 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESDEC.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESDec(DstSize, State, Key, ZeroRegister);
Ref Result = _VAESDec(DstSize, State, Key, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESDecLast(16, Dest, Src, ZeroRegister);
Ref Result = _VAESDecLast(16, Dest, Src, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
@@ -388,16 +393,16 @@ void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
// TODO: Handle 256-bit VAESDECLAST.
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
OrderedNode *State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
OrderedNode *Result = _VAESDecLast(DstSize, State, Key, ZeroRegister);
Ref Result = _VAESDecLast(DstSize, State, Key, ZeroRegister);
StoreResult(FPRClass, Op, Result, -1);
}
OrderedNode* OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
const uint64_t RCON = Op->Src[1].Data.Literal.Value;
@@ -407,15 +412,15 @@ OrderedNode* OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
}
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
OrderedNode *Result = AESKeyGenAssistImpl(Op);
Ref Result = AESKeyGenAssistImpl(Op);
StoreResult(FPRClass, Op, Result, -1);
}
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Selector needs to be literal here");
OrderedNode *Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
OrderedNode *Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
auto Res = _PCLMUL(16, Dest, Src, Selector);
@@ -427,12 +432,12 @@ void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
const auto DstSize = GetDstSize(Op);
OrderedNode *Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
OrderedNode *Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
OrderedNode *Res = _PCLMUL(DstSize, Src1, Src2, Selector);
Ref Res = _PCLMUL(DstSize, Src1, Src2, Selector);
StoreResult(FPRClass, Op, Res, -1);
}
}
} // namespace FEXCore::IR
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -22,24 +22,24 @@ class OrderedNode;
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
//Functions in X87.cpp (no change required)
//GetX87Top
//SetX87ValidTag
//GetX87ValidTag
//GetX87Tag (will need changing once special tag handling is implemented)
//SetX87FTW
//GetX87FTW (will need changing once special tag handling is implemented)
//SetX87Top
//X87ModifySTP
//EMMS
//FFREE
//FNSTENV
//FSTCW
//LDSW
//FNSTSW
//FXCH
//FCMOV
//FST(register to register)
// Functions in X87.cpp (no change required)
// GetX87Top
// SetX87ValidTag
// GetX87ValidTag
// GetX87Tag (will need changing once special tag handling is implemented)
// SetX87FTW
// GetX87FTW (will need changing once special tag handling is implemented)
// SetX87Top
// X87ModifySTP
// EMMS
// FFREE
// FNSTENV
// FSTCW
// LDSW
// FNSTSW
// FXCH
// FCMOV
// FST(register to register)
// State loading duplicated from X87.cpp, setting host rounding mode
// See issue
@@ -64,34 +64,31 @@ void OpDispatchBuilder::FNINITF64(OpcodeArgs) {
}
void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
const auto Size = GetSrcSize(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
//ignore the rounding precision, we're always 64-bit in F64.
//extract rounding mode
OrderedNode *roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
ReconstructX87StateFromFSW(NewFSW);
{
// FTW
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
}
}
void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
OrderedNode *NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
//ignore the rounding precision, we're always 64-bit in F64.
//extract rounding mode
OrderedNode *roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
_SetRoundingMode(roundingMode);
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
}
@@ -106,13 +103,13 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs) {
size_t read_width = (width == 80) ? 16 : width / 8;
OrderedNode *data{};
OrderedNode *converted{};
Ref data {};
Ref converted {};
if (!Op->Src[0].IsNone()) {
// Read from memory
data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], read_width, Op->Flags);
// Convert to 64bit float
// Convert to 64bit float
if constexpr (width == 32) {
converted = _Float_FToF(8, 4, data);
} else if constexpr (width == 80) {
@@ -120,8 +117,7 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs) {
} else {
converted = data;
}
}
else {
} else {
// Implicit arg (does this need to change with width?)
auto offset = _Constant(Op->OP & 7);
data = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, orig_top, offset), mask);
@@ -136,12 +132,9 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs) {
_StoreContextIndexed(converted, top, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FLDF64<32>(OpcodeArgs);
template
void OpDispatchBuilder::FLDF64<64>(OpcodeArgs);
template
void OpDispatchBuilder::FLDF64<80>(OpcodeArgs);
template void OpDispatchBuilder::FLDF64<32>(OpcodeArgs);
template void OpDispatchBuilder::FLDF64<64>(OpcodeArgs);
template void OpDispatchBuilder::FLDF64<80>(OpcodeArgs);
void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
// Update TOP
@@ -152,8 +145,8 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
SetX87Top(top);
// Read from memory
OrderedNode *data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
OrderedNode *converted = _F80BCDLoad(data);
Ref data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
Ref converted = _F80BCDLoad(data);
converted = _F80CVT(8, converted);
_StoreContextIndexed(converted, top, 8, MMBaseOffset(), 16, FPRClass);
}
@@ -162,7 +155,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *converted = _F80CVTTo(data, 8);
Ref converted = _F80CVTTo(data, 8);
converted = _F80BCDStore(converted);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
@@ -185,20 +178,13 @@ void OpDispatchBuilder::FLDF64_Const(OpcodeArgs) {
_StoreContextIndexed(data, top, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FLDF64_Const<0x3FF0000000000000>(OpcodeArgs); // 1.0
template
void OpDispatchBuilder::FLDF64_Const<0x400A934F0979A372>(OpcodeArgs); // log2l(10)
template
void OpDispatchBuilder::FLDF64_Const<0x3FF71547652B82FE>(OpcodeArgs); // log2l(e)
template
void OpDispatchBuilder::FLDF64_Const<0x400921FB54442D18>(OpcodeArgs); // pi
template
void OpDispatchBuilder::FLDF64_Const<0x3FD34413509F79FF>(OpcodeArgs); // log10l(2)
template
void OpDispatchBuilder::FLDF64_Const<0x3FE62E42FEFA39EF>(OpcodeArgs); // log(2)
template
void OpDispatchBuilder::FLDF64_Const<0>(OpcodeArgs); // 0.0
template void OpDispatchBuilder::FLDF64_Const<0x3FF0000000000000>(OpcodeArgs); // 1.0
template void OpDispatchBuilder::FLDF64_Const<0x400A934F0979A372>(OpcodeArgs); // log2l(10)
template void OpDispatchBuilder::FLDF64_Const<0x3FF71547652B82FE>(OpcodeArgs); // log2l(e)
template void OpDispatchBuilder::FLDF64_Const<0x400921FB54442D18>(OpcodeArgs); // pi
template void OpDispatchBuilder::FLDF64_Const<0x3FD34413509F79FF>(OpcodeArgs); // log10l(2)
template void OpDispatchBuilder::FLDF64_Const<0x3FE62E42FEFA39EF>(OpcodeArgs); // log(2)
template void OpDispatchBuilder::FLDF64_Const<0>(OpcodeArgs); // 0.0
void OpDispatchBuilder::FILDF64(OpcodeArgs) {
// Update TOP
@@ -210,7 +196,7 @@ void OpDispatchBuilder::FILDF64(OpcodeArgs) {
size_t read_width = GetSrcSize(Op);
// Read from memory
auto data = LoadSource_WithOpSize(GPRClass, Op, Op->Src[0], read_width, Op->Flags);
if(read_width == 2) {
if (read_width == 2) {
data = _Sbfe(OpSize::i64Bit, read_width * 8, 0, data);
}
auto converted = _Float_FromGPR_S(8, read_width == 4 ? 4 : 8, data);
@@ -223,14 +209,14 @@ void OpDispatchBuilder::FSTF64(OpcodeArgs) {
auto orig_top = GetX87Top();
auto data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
if constexpr (width == 64) {
//Store 64-bit float directly
// Store 64-bit float directly
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, data, 8, 1);
} else if constexpr (width == 32) {
//Convert to 32-bit float and store
// Convert to 32-bit float and store
auto result = _Float_FToF(4, 8, data);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, result, 4, 1);
} else if constexpr (width == 80) {
//Convert to 80-bit float
// Convert to 80-bit float
auto result = _F80CVTTo(data, 8);
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, result, 10, 1);
}
@@ -244,19 +230,16 @@ void OpDispatchBuilder::FSTF64(OpcodeArgs) {
}
}
template
void OpDispatchBuilder::FSTF64<32>(OpcodeArgs);
template
void OpDispatchBuilder::FSTF64<64>(OpcodeArgs);
template
void OpDispatchBuilder::FSTF64<80>(OpcodeArgs);
template void OpDispatchBuilder::FSTF64<32>(OpcodeArgs);
template void OpDispatchBuilder::FSTF64<64>(OpcodeArgs);
template void OpDispatchBuilder::FSTF64<80>(OpcodeArgs);
template<bool Truncate>
void OpDispatchBuilder::FISTF64(OpcodeArgs) {
auto Size = GetSrcSize(Op);
auto orig_top = GetX87Top();
OrderedNode *data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
Ref data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
if constexpr (Truncate) {
data = _Float_ToGPR_ZS(Size == 4 ? 4 : 8, 8, data);
} else {
@@ -273,18 +256,16 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs) {
}
}
template
void OpDispatchBuilder::FISTF64<false>(OpcodeArgs);
template
void OpDispatchBuilder::FISTF64<true>(OpcodeArgs);
template void OpDispatchBuilder::FISTF64<false>(OpcodeArgs);
template void OpDispatchBuilder::FISTF64<true>(OpcodeArgs);
template <size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
void OpDispatchBuilder::FADDF64(OpcodeArgs) {
auto top = GetX87Top();
OrderedNode *StackLocation = top;
Ref StackLocation = top;
OrderedNode *arg{};
OrderedNode *b{};
Ref arg {};
Ref b {};
auto mask = _Constant(7);
@@ -292,7 +273,7 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if(width == 16) {
if (width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
@@ -326,26 +307,20 @@ void OpDispatchBuilder::FADDF64(OpcodeArgs) {
_StoreContextIndexed(result, StackLocation, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FADDF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FADDF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<80, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FADDF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FADDF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FADDF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
void OpDispatchBuilder::FMULF64(OpcodeArgs) {
auto top = GetX87Top();
OrderedNode *StackLocation = top;
OrderedNode *arg{};
OrderedNode *b{};
Ref StackLocation = top;
Ref arg {};
Ref b {};
auto mask = _Constant(7);
@@ -353,7 +328,7 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if(width == 16) {
if (width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
@@ -390,34 +365,28 @@ void OpDispatchBuilder::FMULF64(OpcodeArgs) {
_StoreContextIndexed(result, StackLocation, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FMULF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FMULF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<32, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<64, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<80, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FMULF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FMULF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<16, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FMULF64<32, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
auto top = GetX87Top();
OrderedNode *StackLocation = top;
OrderedNode *arg{};
OrderedNode *b{};
Ref StackLocation = top;
Ref arg {};
Ref b {};
auto mask = _Constant(7);
if (!Op->Src[0].IsNone()) {
if (!Op->Src[0].IsNone()) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if(width == 16) {
if (width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
@@ -440,11 +409,10 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
Ref result {};
if constexpr (reverse) {
result = _VFDiv(8, 8, b, a);
}
else {
} else {
result = _VFDiv(8, 8, a, b);
}
@@ -460,50 +428,38 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
_StoreContextIndexed(result, StackLocation, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FDIVF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FDIVF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FDIVF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
auto top = GetX87Top();
OrderedNode *StackLocation = top;
OrderedNode *arg{};
OrderedNode *b{};
Ref StackLocation = top;
Ref arg {};
Ref b {};
auto mask = _Constant(7);
if (!Op->Src[0].IsNone()) {
if (!Op->Src[0].IsNone()) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if(width == 16) {
if (width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
@@ -526,11 +482,10 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *result{};
Ref result {};
if constexpr (reverse) {
result = _VFSub(8, 8, b, a);
}
else {
} else {
result = _VFSub(8, 8, a, b);
}
@@ -547,35 +502,23 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
_StoreContextIndexed(result, StackLocation, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::FSUBF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<32, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<32, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<64, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<64, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<80, false, false, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<80, false, true, OpDispatchBuilder::OpResult::RES_STI>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<16, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<16, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template
void OpDispatchBuilder::FSUBF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<32, true, false, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
template void OpDispatchBuilder::FSUBF64<32, true, true, OpDispatchBuilder::OpResult::RES_ST0>(OpcodeArgs);
void OpDispatchBuilder::FCHSF64(OpcodeArgs) {
auto top = GetX87Top();
@@ -598,7 +541,7 @@ void OpDispatchBuilder::FTSTF64(OpcodeArgs) {
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
auto low = _Constant(0);
OrderedNode *data = _VCastFromGPR(8, 8, low);
Ref data = _VCastFromGPR(8, 8, low);
// We are going to clobber NZCV, make sure it's in a GPR first.
GetNZCV();
@@ -609,7 +552,7 @@ void OpDispatchBuilder::FTSTF64(OpcodeArgs) {
ConvertNZCVToX87();
}
//TODO: This should obey rounding mode
// TODO: This should obey rounding mode
void OpDispatchBuilder::FRNDINTF64(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
@@ -628,11 +571,11 @@ void OpDispatchBuilder::FXTRACTF64(OpcodeArgs) {
auto a = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
auto gpr = _VExtractToGPR(8, 8, a, 0);
OrderedNode* exp = _And(OpSize::i64Bit, gpr, _Constant(0x7ff0000000000000LL));
Ref exp = _And(OpSize::i64Bit, gpr, _Constant(0x7ff0000000000000LL));
exp = _Lshr(OpSize::i64Bit, exp, _Constant(52));
exp = _Sub(OpSize::i64Bit, exp, _Constant(1023));
exp = _Float_FromGPR_S(8, 8, exp);
OrderedNode* sig = _And(OpSize::i64Bit, gpr, _Constant(0x800fffffffffffffLL));
Ref sig = _And(OpSize::i64Bit, gpr, _Constant(0x800fffffffffffffLL));
sig = _Or(OpSize::i64Bit, sig, _Constant(0x3ff0000000000000LL));
sig = _VCastFromGPR(8, 8, sig);
// Write to ST[TOP]
@@ -646,14 +589,14 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
auto top = GetX87Top();
auto mask = _Constant(7);
OrderedNode *arg{};
OrderedNode *b{};
Ref arg {};
Ref b {};
if (!Op->Src[0].IsNone()) {
// Memory arg
if constexpr (Integer) {
arg = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
if(width == 16) {
if (width == 16) {
arg = _Sbfe(OpSize::i64Bit, 16, 0, arg);
}
b = _Float_FromGPR_S(8, width == 64 ? 8 : 4, arg);
@@ -679,8 +622,7 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
_FCmp(8, a, b);
PossiblySetNZCVBits = ~0;
ConvertNZCVToX87();
}
else {
} else {
// Invalidate deferred flags early
// OF, SF, AF, PF all undefined
InvalidateDeferredFlags();
@@ -698,8 +640,7 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
// Set the new top now
top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
SetX87Top(top);
}
else if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
} else if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
// if we are popping then we must first mark this location as empty
SetX87ValidTag(top, false);
// Set the new top now
@@ -708,24 +649,17 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
}
}
template
void OpDispatchBuilder::FCOMIF64<32, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<32, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<64, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<64, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_RFLAGS, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<80, false, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, true>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<16, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<16, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template
void OpDispatchBuilder::FCOMIF64<32, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
template void OpDispatchBuilder::FCOMIF64<32, true, OpDispatchBuilder::FCOMIFlags::FLAGS_X87, false>(OpcodeArgs);
void OpDispatchBuilder::FSQRTF64(OpcodeArgs) {
@@ -746,8 +680,7 @@ void OpDispatchBuilder::X87UnaryOpF64(OpcodeArgs) {
DeriveOp(result, IROp, _F64SIN(a));
if constexpr (IROp == IR::OP_F64SIN ||
IROp == IR::OP_F64COS) {
if constexpr (IROp == IR::OP_F64SIN || IROp == IR::OP_F64COS) {
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
@@ -756,12 +689,9 @@ void OpDispatchBuilder::X87UnaryOpF64(OpcodeArgs) {
_StoreContextIndexed(result, top, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64F2XM1>(OpcodeArgs);
template
void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64SIN>(OpcodeArgs);
template
void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64COS>(OpcodeArgs);
template void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64F2XM1>(OpcodeArgs);
template void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64SIN>(OpcodeArgs);
template void OpDispatchBuilder::X87UnaryOpF64<IR::OP_F64COS>(OpcodeArgs);
template<FEXCore::IR::IROps IROp>
@@ -769,16 +699,15 @@ void OpDispatchBuilder::X87BinaryOpF64(OpcodeArgs) {
auto top = GetX87Top();
auto mask = _Constant(7);
OrderedNode *st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
Ref st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
st1 = _LoadContextIndexed(st1, 8, MMBaseOffset(), 16, FPRClass);
DeriveOp(result, IROp, _F64ATAN(a, st1));
if constexpr (IROp == IR::OP_F64FPREM ||
IROp == IR::OP_F64FPREM1) {
//TODO: Set C0 to Q2, C3 to Q1, C1 to Q0
if constexpr (IROp == IR::OP_F64FPREM || IROp == IR::OP_F64FPREM1) {
// TODO: Set C0 to Q2, C3 to Q1, C1 to Q0
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
}
@@ -786,12 +715,9 @@ void OpDispatchBuilder::X87BinaryOpF64(OpcodeArgs) {
_StoreContextIndexed(result, top, 8, MMBaseOffset(), 16, FPRClass);
}
template
void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM1>(OpcodeArgs);
template
void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM>(OpcodeArgs);
template
void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64SCALE>(OpcodeArgs);
template void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM1>(OpcodeArgs);
template void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64FPREM>(OpcodeArgs);
template void OpDispatchBuilder::X87BinaryOpF64<IR::OP_F64SCALE>(OpcodeArgs);
void OpDispatchBuilder::X87SinCosF64(OpcodeArgs) {
auto orig_top = GetX87Top();
@@ -821,8 +747,8 @@ void OpDispatchBuilder::X87FYL2XF64(OpcodeArgs) {
auto top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, orig_top, _Constant(1)), _Constant(7));
SetX87Top(top);
OrderedNode *st0 = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
Ref st0 = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
Ref st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
if (Plus1) {
auto one = _VCastFromGPR(8, 8, _Constant(0x3FF0000000000000));
@@ -863,7 +789,7 @@ void OpDispatchBuilder::X87ATANF64(OpcodeArgs) {
SetX87Top(top);
auto a = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
Ref st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
auto result = _F64ATAN(st1, a);
@@ -871,7 +797,7 @@ void OpDispatchBuilder::X87ATANF64(OpcodeArgs) {
_StoreContextIndexed(result, top, 8, MMBaseOffset(), 16, FPRClass);
}
//This function converts to F80 on save for compatibility
// This function converts to F80 on save for compatibility
void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
// 14 bytes for 16bit
@@ -893,92 +819,76 @@ void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
// 4 bytes : data pointer offset
// 4 bytes : data pointer selector
auto Size = GetDstSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Dest, Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
OrderedNode *Top = GetX87Top();
const auto Size = GetDstSize(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
Ref Top = GetX87Top();
{
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
_StoreMem(GPRClass, Size, Mem, FCW, Size);
}
{
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
_StoreMem(GPRClass, Size, MemLocation, ReconstructFSW(), Size);
}
{ _StoreMem(GPRClass, Size, ReconstructFSW(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
auto ZeroConst = _Constant(0);
{
// FTW
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
_StoreMem(GPRClass, Size, MemLocation, GetX87FTW(), Size);
_StoreMem(GPRClass, Size, GetX87FTW(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction Offset
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 3));
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
}
{
// Instruction CS selector (+ Opcode)
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 4));
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer offset
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 5));
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
}
{
// Data pointer selector
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 6));
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
}
OrderedNode *ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto TenConst = _Constant(10);
for (int i = 0; i < 7; ++i) {
OrderedNode* data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
data = _F80CVTTo(data, 8);
_StoreMem(FPRClass, 16, ST0Location, data, 1);
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
_StoreMem(FPRClass, 16, data, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
// The final st(7) needs a bit of special handling here
OrderedNode* data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
data = _F80CVTTo(data, 8);
// ST7 broken in to two parts
// Lower 64bits [63:0]
// upper 16 bits [79:64]
_StoreMem(FPRClass, 8, ST0Location, data, 1);
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
_StoreMem(FPRClass, 8, data, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
auto topBytes = _VDupElement(16, 2, data, 4);
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
_StoreMem(FPRClass, 2, topBytes, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
// reset to default
FNINIT(Op);
}
//This function converts from F80 on load for compatibility
// This function converts from F80 on load for compatibility
void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
auto Size = GetSrcSize(Op);
OrderedNode *Mem = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags, {.LoadData = false});
Mem = AppendSegmentOffset(Mem, Op->Flags);
const auto Size = GetSrcSize(Op);
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
//ignore the rounding precision, we're always 64-bit in F64.
//extract rounding mode
OrderedNode *roundingMode = NewFCW;
// ignore the rounding precision, we're always 64-bit in F64.
// extract rounding mode
Ref roundingMode = NewFCW;
auto roundShift = _Constant(10);
auto roundMask = _Constant(3);
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
@@ -987,36 +897,30 @@ void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
auto Top = ReconstructX87StateFromFSW(NewFSW);
{
// FTW
OrderedNode *MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
}
OrderedNode *ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
auto OneConst = _Constant(1);
auto SevenConst = _Constant(7);
auto TenConst = _Constant(10);
auto low = _Constant(~0ULL);
auto high = _Constant(0xFFFF);
OrderedNode *Mask = _VCastFromGPR(16, 8, low);
Ref Mask = _VCastFromGPR(16, 8, low);
Mask = _VInsGPR(16, 8, 1, Mask, high);
for (int i = 0; i < 7; ++i) {
OrderedNode *Reg = _LoadMem(FPRClass, 16, ST0Location, 1);
Ref Reg = _LoadMem(FPRClass, 16, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
// Mask off the top bits
Reg = _VAnd(16, 16, Reg, Mask);
//Convert to double precision
// Convert to double precision
Reg = _F80CVT(8, Reg);
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
}
@@ -1025,20 +929,19 @@ void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
// Lower 64bits [63:0]
// upper 16 bits [79:64]
OrderedNode *Reg = _LoadMem(FPRClass, 8, ST0Location, 1);
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
OrderedNode *RegHigh = _LoadMem(FPRClass, 2, ST0Location, 1);
Ref Reg = _LoadMem(FPRClass, 8, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
Ref RegHigh = _LoadMem(FPRClass, 2, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
Reg = _F80CVT(8, Reg); //Convert to double precision
Reg = _F80CVT(8, Reg); // Convert to double precision
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
}
//FXAM needs change
// FXAM needs change
void OpDispatchBuilder::X87FXAMF64(OpcodeArgs) {
auto top = GetX87Top();
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
OrderedNode *Result = _VExtractToGPR(8, 8, a, 0);
Ref Result = _VExtractToGPR(8, 8, a, 0);
// Extract the sign bit
Result = _Bfe(OpSize::i64Bit, 1, 63, Result);
@@ -1051,9 +954,7 @@ void OpDispatchBuilder::X87FXAMF64(OpcodeArgs) {
auto OneConst = _Constant(1);
// In the case of top being invalid then C3:C2:C0 is 0b101
auto C3 = _Select(FEXCore::IR::COND_EQ,
TopValid, OneConst,
ZeroConst, OneConst);
auto C3 = _Select(FEXCore::IR::COND_EQ, TopValid, OneConst, ZeroConst, OneConst);
auto C2 = TopValid;
auto C0 = C3; // Mirror C3 until something other than zero is supported
@@ -1063,4 +964,4 @@ void OpDispatchBuilder::X87FXAMF64(OpcodeArgs) {
}
}
} // namespace FEXCore::IR
@@ -39,15 +39,15 @@ X86GeneratedCode::X86GeneratedCode() {
// Falling back to this generated code segment still allows a backtrace to work, just might not show
// the symbol as VDSO since there is no ELF to parse.
constexpr std::array<uint8_t, 9> sigreturn_32_code = {
0x58, // pop eax
0x58, // pop eax
0xb8, 0x77, 0x00, 0x00, 0x00, // mov eax, 0x77
0xcd, 0x80, // int 0x80
0x90, // nop
0xcd, 0x80, // int 0x80
0x90, // nop
};
constexpr std::array<uint8_t, 7> rt_sigreturn_32_code = {
0xb8, 0xad, 0x00, 0x00, 0x00, // mov eax, 0xad
0xcd, 0x80, // int 0x80
0xcd, 0x80, // int 0x80
};
CallbackReturn = reinterpret_cast<uint64_t>(CodePtr);
@@ -84,10 +84,9 @@ void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
// We need to have the sigret handler in the lower 32bits of memory space
// Scan top down and try to allocate a location
for (size_t Location = 0xFFFF'E000; Location != 0x0; Location -= 0x1000) {
void *Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
void* Ptr = ::mmap(reinterpret_cast<void*>(Location), Size, PROT_READ | PROT_WRITE, MAP_FIXED_NOREPLACE | MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
if (Ptr != MAP_FAILED &&
reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
if (Ptr != MAP_FAILED && reinterpret_cast<uintptr_t>(Ptr) >= LOCATION_MAX) {
// Failed to map in the lower 32bits
// Try again
// Can happen in the case that host kernel ignores MAP_FIXED_NOREPLACE
@@ -108,5 +107,4 @@ void* X86GeneratedCode::AllocateGuestCodeSpace(size_t Size) {
#endif
}
}
} // namespace FEXCore
+5 -5
View File
@@ -16,12 +16,12 @@ public:
X86GeneratedCode();
~X86GeneratedCode();
uint64_t CallbackReturn{};
uint64_t sigreturn_32{};
uint64_t rt_sigreturn_32{};
uint64_t CallbackReturn {};
uint64_t sigreturn_32 {};
uint64_t rt_sigreturn_32 {};
private:
void *CodePtr{};
void* CodePtr {};
void* AllocateGuestCodeSpace(size_t Size);
};
}
} // namespace FEXCore
+1 -1
View File
@@ -24,4 +24,4 @@ void InitializeInfoTables(Context::OperatingMode Mode) {
InitializeH0F3ATables(Mode);
}
}
} // namespace FEXCore::X86Tables
@@ -102,10 +102,10 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1, nullptr}},
// This should just throw a GP
{0x6C, 1, X86InstInfo{"INSB", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
@@ -183,24 +183,24 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
// Should just throw GP
{0xE4, 2, X86InstInfo{"IN", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_NONE, 1, nullptr}},
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_NONE, 1, nullptr}},
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1, nullptr}},
// Should just throw GP
{0xEC, 2, X86InstInfo{"IN", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{0xFB, 1, X86InstInfo{"STI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
@@ -32,7 +32,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_6, PF_NONE, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -41,7 +41,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_6, PF_F3, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -50,7 +50,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_6, PF_66, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -59,7 +59,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_6, PF_F2, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_6, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
@@ -72,7 +72,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
@@ -81,7 +81,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
@@ -90,7 +90,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
@@ -99,7 +99,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
// GROUP 8
@@ -162,7 +162,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
{OPD(TYPE_GROUP_9, PF_F3, 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_F3, 7), 1, X86InstInfo{"RDPID", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_REG_ONLY, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 0), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
{OPD(TYPE_GROUP_9, PF_66, 1), 1, X86InstInfo{"CMPXCHG8B/16B", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_SF_MOD_MEM_ONLY, 0, nullptr}},
Loaded 100 of 580 files, more files were not shown because too many files have changed in this diff. Show more