Compare commits

...
880 Commits
Author SHA1 Message Date
Ryan Houdek bf66cac272 Docs: Update for release FEX-2309 2023-09-05 22:16:35 -07:00
Ryan Houdek 016c3c0f07 Merge pull request #3057 from Sonicadvance1/support_procfs_interpreter
FEXInterpreter: Supports procfs/interpreter
2023-09-05 21:53:45 -07:00
Ryan Houdek 09a49a3420 FEXInterpreter: Supports procfs/interpreter
This is a new procfs symlink path that changes behaviour of binfmt_misc
when exposed. We need to check both procfs/exe and procfs/interpreter
and see if they exist AND also differ.

Once/if they do then we can disable a bunch of checking of paths once
they do. The fallback when none of this is supported has the same
behaviour has previously where it still does all the regular checking.

During binfmt_misc install cmake will check the kernel version for the
raw binfmt_misc writing. Which will never pass until we have a real
kernel version that it is upstreamed in.

For update-binfmts we add a new optional argument where the tool will
drop the flag if the host kernel version isn't new enough to handle the
option.
2023-09-05 21:30:47 -07:00
Ryan Houdek 16f826c18e Merge pull request #3064 from alyssarosenzweig/ir/rm-phi
IR: Remove phi nodes
2023-09-05 19:43:52 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Ryan Houdek 04228952fa Merge pull request #3063 from alyssarosenzweig/flag/defer-pf-completely
Defer PF calculation completely
2023-09-05 13:18:22 -07:00
Alyssa Rosenzweig 5a3dc8c2ab InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:00:10 -04:00
Alyssa Rosenzweig 8efe2eeef6 OpcodeDispatcher: Defer PF invert
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.

This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:

* Zeroing PF. This now requires writing 1 instead of 0, which may require an
  extra move for the constant. Some of this will go away when we merge PF+AF
  into a single register, which is next up on the list. In that case, the
  two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
  x86 view of PF), that will get absorbed into the or, if we also write AF. If
  we leave AF undefined and need to write a 1, that's a single mov instruction
  and we couldn't do better anyway if not inverted (since we'd still have a mov
  wzr even then). So in view of the future work, this isn't something I'm
  concerned about.

* Float comparisons that put Unordered into PF. These require an extra invert to
  match the new convention. These are already so unnecessarily bloated that I'm
  not convinced I'm making things materially worse here. But we realistically
  need multiple destination support in the IR to fix this particular mess.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 15:58:18 -04:00
Ryan Houdek 8184c55424 Merge pull request #3062 from alyssarosenzweig/flag/no-pf
Remove ABINoPF option
2023-09-05 12:26:45 -07:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Ryan Houdek a8bc6bbb2e Merge pull request #3059 from alyssarosenzweig/flag/defer-af-xor
Defer second XOR for AF
2023-09-05 11:38:47 -07:00
Alyssa Rosenzweig 305bb98cf8 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig 80ac824dd9 Unittests: Fix bogus lahf tests
Logical ops leave AF undefined so we can't expect it to be zero after. Mask the
result of lahf to avoid testing UB. These unit tests would regress from the work
in this MR otherwise.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Ryan Houdek c27f69dd6b Merge pull request #3061 from alyssarosenzweig/flag/cmc
OpcodeDispatcher: Optimize CMC
2023-09-05 11:06:21 -07:00
Alyssa Rosenzweig b6462ee854 OpcodeDispatcher: Optimize CMC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 13:55:05 -04:00
Ryan Houdek 252ca88b3e Merge pull request #3058 from alyssarosenzweig/flag/undef
Stop zeroing undefined flags
2023-09-05 10:49:02 -07:00
Alyssa Rosenzweig 8ba1e91699 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig ff0b514da8 OpcodeDispatcher: Invalidate PF/AF in more cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig 240260576b OpcodeDispatcher: Stop zeroing so many flags
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:10:29 -04:00
Ryan Houdek 486f0ba1e3 Merge pull request #3038 from alyssarosenzweig/flag/af
Defer AF extract
2023-09-05 08:54:47 -07:00
Ryan Houdek 67a26a0e98 Merge pull request #3056 from Sonicadvance1/constpool_heuristic
ConstProp: Adds constpool distance heuristic
2023-09-05 08:46:22 -07:00
Alyssa Rosenzweig a1e3ce3cdf InstCountCI: Update for deferred AF extract
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:39:32 -04:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00
Alyssa Rosenzweig f46ba52e0e OpcodeDispatcher: Use LoadAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Alyssa Rosenzweig c88b022d9f OpcodeDispatcher: Add AF accessors
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Ryan Houdek a24680b0a1 ConstProp: Adds constpool distance heuristic
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.

This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.

FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.

To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.
2023-09-03 18:27:43 -07:00
Mai 2d78b1fbca Merge pull request #3055 from Sonicadvance1/defer_disasm_init
Arm64: Only allocate vixl::Decoder if enabled
2023-09-03 18:58:42 -04:00
Ryan Houdek ca1c33047c Arm64: Only allocate vixl::Decoder if enabled
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.

Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
2023-09-03 12:53:18 -07:00
Ryan Houdek b18592f153 Merge pull request #3050 from Sonicadvance1/fix_flag_reconstruction
OpcodeDispatcher: Fixes NZCV and PF flag compacting
2023-09-03 10:15:44 -07:00
Mai 8017a91e52 Merge pull request #3054 from Sonicadvance1/remove_small_ir_size_assumptions
OpcodeDispatcher: Remove final assumptions about small IR operating sizes
2023-09-03 05:47:38 -04:00
Ryan Houdek a235c3d81a InstCountCI: Update for small operator changes and opt 2023-09-03 02:26:23 -07:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Mai 62fcf6cbd0 Merge pull request #3053 from Sonicadvance1/rflag_handling_32bit
OpcodeDispatcher: Cleans up RFLAGS size handling
2023-09-03 05:13:42 -04:00
Ryan Houdek cb8f183e9f InstCountCI: Update for flags cleanup 2023-09-03 01:41:52 -07:00
Ryan Houdek 44a14e7fd0 OpcodeDispatcher: Cleans up RFLAGS size handling
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.

Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
2023-09-03 01:41:04 -07:00
Mai 2d22176699 Merge pull request #3052 from Sonicadvance1/remove_todo_shlimm
OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
2023-09-03 04:38:52 -04:00
Mai 338cb199a0 Merge pull request #3051 from Sonicadvance1/remove_todo_shiftleft
OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront
2023-09-03 04:38:00 -04:00
Ryan Houdek 16839671e5 InstcountCI: Update for SHLImm changes 2023-09-03 00:28:36 -07:00
Ryan Houdek c425db7284 OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
This one was easy, barely anything changes behaviour, as seen by
InstCountCI changes.
2023-09-03 00:27:28 -07:00
Ryan Houdek ba422c1dc9 InstCountCI: Update for ShiftLeft changes 2023-09-03 00:16:46 -07:00
Ryan Houdek d0595b5f13 OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront 2023-09-03 00:16:26 -07:00
Ryan Houdek 8bae58dcc3 Merge pull request #3049 from bylaws/x87
FEXCore: Rework X87 tag word handling
2023-09-03 00:00:12 -07:00
Ryan Houdek cdfa6939b3 FEXLinuxTests: Adds unit tests to ensure we set EFLAGS correctly
Tests all five of the flags that need specific handling. Without the
prior patch FEX would fail these.
2023-09-02 23:10:57 -07:00
Ryan Houdek cb5d665046 OpcodeDispatcher: Fixes NZCV and PF flag compacting
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.

This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
2023-09-02 23:10:57 -07:00
Billy Laws 393b8c657a InstCountCI: Update for x87 changes 2023-09-02 09:17:33 -07:00
Billy Laws 3792f707dc FEXLoader: Convert between abridged/full tag fmts in signal dispatch
X86 fpstate expects FTW to be saved in the FSAVE format, whereas X64
fpstate expects it to be saved in the abridged format used by FXSAVE.
2023-09-02 09:17:33 -07:00
Billy Laws 13b8f95f85 X87: Switch all stack pointer accesses to 32-bit OpSize 2023-09-02 09:17:33 -07:00
Billy Laws cb49373f47 FEXCore: Rework X87 tag word handling
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
2023-09-02 09:17:33 -07:00
Ryan Houdek ea965810e5 Merge pull request #3048 from Sonicadvance1/construct_eflags
Context: Adds helper to reconstruct and consume packed EFLAGS
2023-09-02 08:14:58 -07:00
Ryan Houdek 435f03c703 Context: Adds helper to reconstruct and consume packed EFLAGS
Currently FEX's internal EFLAGS representation is a perfect 1:1 mapping
between bit offset and byte offset. This is going to change with #3038.
There should be no reason that the frontend needs to understand how to
reconstruct the compacted flags from the internal representation.

Adds context helpers and moves all the logic to FEXCore. The locations
that previously needed to handle this have been converted over to use
this.
2023-09-02 07:05:54 -07:00
Ryan Houdek 81046efcd3 Merge pull request #3047 from alyssarosenzweig/build/no-telem
SignalDelegator: Fix build with telemetry disabled
2023-09-02 05:47:33 -07:00
Alyssa Rosenzweig 69b0747281 SignalDelegator: Fix build with telemetry disabled
/home/alyssa/FEX/Source/Tools/FEXLoader/LinuxSyscalls/SignalDelegator.cpp:1546:7: error: use of undeclared identifier 'CrashMask'
      CrashMask |= (1ULL << Signal);

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-02 08:09:05 -04:00
Mai f0ab9603e1 Merge pull request #3046 from Sonicadvance1/fix_telemetry_exit_group
Syscalls: Fix telemetry with exit_group
2023-09-01 05:15:56 -04:00
Ryan Houdek aea4a88e00 Syscalls: Fix telemetry with exit_group
Picked up on more games exiting with exit_group that I want to ensure
their telemetry data gets saved. Implement support for this.
2023-09-01 01:53:07 -07:00
Ryan Houdek ee8092bdc8 Merge pull request #3035 from alyssarosenzweig/flag/opts
Optimize ADD flag calculation
2023-08-31 16:30:01 -07:00
Alyssa Rosenzweig 31c92a422b InstCountCI: Update for add/sub work
Big Delete The Code energy.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:04 -04:00
Alyssa Rosenzweig 2edce18a59 ConstProp: Optimize XOR with 0
This cleans up the AF flag calculation for `neg`.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:04 -04:00
Alyssa Rosenzweig 659568ec1b ConstProp: Propagate 0 to first argument of SubNZCV
This allows inlining a constant into the comparison for Neg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:01 -04:00
Alyssa Rosenzweig dcb09b085b OpcodeDispatcher: Use AddNZCV/SubNZCV
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:41 -04:00
Alyssa Rosenzweig 45d99a2cce OpcodeDispatcher: Fix source sizes for Sub flags
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:39 -04:00
Alyssa Rosenzweig f688364bf2 IR: Add AddNZCV/SubNZCV op
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.

Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 07:52:12 -04:00
Mai 7c81a0d4fe Merge pull request #3042 from Sonicadvance1/optimize_call
OpcodeDispatcher: Optimize calls with push
2023-08-31 02:23:06 -04:00
Mai 7f997387b0 Merge pull request #3044 from Sonicadvance1/optimize_aes
Arm64: Optimize AES operations by caching a zero register
2023-08-31 00:20:10 -04:00
Ryan Houdek cfe6b929bb InstCountCI: Updates for aes optimization
XMM AES operation is now optimal.
Most other changes are just because of RA changing around the zero
register.
2023-08-30 20:43:11 -07:00
Ryan Houdek 61df7a576a OpcodeDispatcher: Be super defensive when starting a new block
Ensure all cached data is correct.
2023-08-30 20:43:11 -07:00
Ryan Houdek ffaa908475 OpcodeDispatcher: Use the named zero register for each usage
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
2023-08-30 20:43:11 -07:00
Ryan Houdek a2307f28d6 Arm64: Optimize AES operations by caching a zero register
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.

Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.

AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
2023-08-30 19:00:43 -07:00
Ryan Houdek 1446d4fe12 IR: Adds support for named vector zero
This is useful for caching a zero register vector which we use in
various locations. This will be abused soon.
2023-08-30 18:59:38 -07:00
Mai 9fb8c95ef4 Merge pull request #3043 from Sonicadvance1/optimize_blendv
OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
2023-08-30 20:36:14 -04:00
Ryan Houdek 58caee3614 InstCountCI: Update for optimized blendv 2023-08-30 17:13:05 -07:00
Ryan Houdek 24215f7ad0 OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
2023-08-30 17:09:34 -07:00
Ryan Houdek b20c518bf0 OpcodeDispatcher: Optimize calls with push
InstCountCI doesn't cover branch instructions so needs manual
inspection.
2023-08-30 16:25:01 -07:00
Ryan Houdek d738538a34 Merge pull request #3041 from Sonicadvance1/telem_runtime
FEXCore: Allows disabling telemetry at runtime
2023-08-30 13:30:51 -07:00
Ryan Houdek 48df17b836 InstCountCI: Update for changed tests
Since the telemetry array is always in the corestate now some numbers
change.
2023-08-30 13:01:27 -07:00
Ryan Houdek 81a32c3998 FEXCore: Allows disabling telemetry at runtime
This is useful for InstCountCI so you can disable the telemetry
gathering even if enabled so it doesn't affect the CI system.
2023-08-30 12:59:41 -07:00
Mai 02b891c0fe Merge pull request #3040 from Sonicadvance1/optimize_aeskeygen
Arm64: Optimize AESKeyGenAssist
2023-08-30 15:58:29 -04:00
Mai 012750f2bb Merge pull request #3037 from Sonicadvance1/remove_debug_log
OpcodeDispatcher: Removes erroneous debug log
2023-08-30 15:57:31 -04:00
Ryan Houdek b40147a269 InstCountCI: Update for aeskeygenassist 2023-08-30 12:17:07 -07:00
Ryan Houdek d8f131fa3d Arm64: Optimize AESKeyGenAssist
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.

I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
2023-08-30 12:15:09 -07:00
Ryan Houdek 2c2081c977 Merge pull request #3039 from Sonicadvance1/print_OpSize
IR: Adds printer for OpSize
2023-08-30 11:50:13 -07:00
Ryan Houdek 572cc57aa3 IR: Adds printer for OpSize 2023-08-30 11:32:43 -07:00
Ryan Houdek 9bfd4b650f OpcodeDispatcher: Removes erroneous debug log 2023-08-30 11:12:12 -07:00
Ryan Houdek e1c2033fa9 Merge pull request #3036 from Sonicadvance1/OpSize_IRDump_Parse
IR: Adds IR::OpSize to IRDumper
2023-08-30 11:01:06 -07:00
Ryan Houdek d3ee794dcb IR: Adds IR::OpSize to IRDumper
This was missing before.
2023-08-30 10:43:23 -07:00
Mai b40f784da0 Merge pull request #3034 from Sonicadvance1/remove_implicit_add
IR: Removes implicit sized add
2023-08-30 08:50:49 -04:00
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Mai 351c0eee42 Merge pull request #3033 from Sonicadvance1/remove_implicit_bfe
IR: Removes implicit sized bfe
2023-08-29 23:02:02 -04:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Mai 8239f8aa27 Merge pull request #3031 from Sonicadvance1/remove_implicit_and
IR: Removes implicit sized and
2023-08-29 02:32:40 -04:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek 59a4d15907 Merge pull request #3030 from Sonicadvance1/remove_implicit_sub
IR: Removes implicit sized sub
2023-08-28 22:20:30 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Mai 227ba9f9fc Merge pull request #3029 from Sonicadvance1/remove_variable_bfi
IR: Removes bfi from variable size
2023-08-29 00:50:45 -04:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00
Mai 9b55a34e75 Merge pull request #3028 from Sonicadvance1/remove_implicit_xor
IR: Removes implicit sized xor
2023-08-28 23:59:39 -04:00
Ryan Houdek 8bc84c202c IR: Removes implicit sized xor 2023-08-28 19:51:14 -07:00
Ryan Houdek e9a3848602 Merge pull request #3027 from Sonicadvance1/remove_implicit_andn
IR: Removes implicit sized andn
2023-08-28 19:39:49 -07:00
Ryan Houdek 516f27bff5 Merge pull request #3026 from Sonicadvance1/remove_implicit_or
IR: Removes implicit sized or
2023-08-28 19:39:42 -07:00
Ryan Houdek 1699ec9a76 IR: Removes implicit sized andn 2023-08-28 19:16:16 -07:00
Ryan Houdek db6c8852fc IR: Removes implicit sized or 2023-08-28 19:06:05 -07:00
Ryan Houdek 24a9254e37 Merge pull request #3025 from Sonicadvance1/remove_implicit_lshr
IR: Removes implicit sized lshr
2023-08-28 19:05:47 -07:00
Ryan Houdek 65dc6f3e90 IR: Removes implicit sized lshr 2023-08-28 18:16:56 -07:00
Mai 8534d3dfbf Merge pull request #3024 from Sonicadvance1/remove_implicit_lshl
IR: Removes implicit sized lshl
2023-08-28 21:08:05 -04:00
Ryan Houdek 60c4438780 IR: Removes implicit sized lshl 2023-08-28 17:50:41 -07:00
Ryan Houdek 915e52046c Merge pull request #3023 from Sonicadvance1/remove_mul
IR: Removes implicit sized mul ops
2023-08-28 17:49:32 -07:00
Ryan Houdek 898ce1ce8f IR: Removes implicit sized UMulH 2023-08-28 17:20:55 -07:00
Ryan Houdek aa8dfd6af1 IR: Removes implicit sized UMul 2023-08-28 17:20:55 -07:00
Ryan Houdek 6a6d808b0d IR: Removes implicit sized MulH 2023-08-28 17:20:55 -07:00
Ryan Houdek fac5b2ac72 IR: Removes implicit sized Mul 2023-08-28 17:20:55 -07:00
Ryan Houdek a55616e4db Merge pull request #3022 from Sonicadvance1/remove_sext
IR: Removes sext IR helper
2023-08-28 17:20:39 -07:00
Ryan Houdek b9e4a1423f IR: Removes sext IR helper
You hold no power here IR operation.
2023-08-28 17:03:38 -07:00
Ryan Houdek 6f661534a2 Merge pull request #3021 from Sonicadvance1/remove_implicit_op_part_move
IR: Removes implicit sized {Create,Extract}ElementPair
2023-08-28 17:03:12 -07:00
Ryan Houdek c0bb6a053f IR: Removes implicit sized {Create,Extract}ElementPair 2023-08-28 16:50:00 -07:00
Ryan Houdek bc1e89d91d Merge pull request #3020 from Sonicadvance1/remove_implicit_ops_pt_atomic
Remove implicit sized IR ops part atomic
2023-08-28 07:27:47 -07:00
Ryan Houdek c1b4c11e54 Merge pull request #3018 from Sonicadvance1/opcodedispatcher_sizeopt
OpcodeDispatcher: Optimize Get{Src,Dst}Size
2023-08-28 07:13:41 -07:00
Ryan Houdek a36427d01e IR: Removes implicit sized CAS 2023-08-28 07:12:31 -07:00
Ryan Houdek 594baff705 IR: Removes implicit sized CASPair 2023-08-28 07:12:31 -07:00
Ryan Houdek e9f2bc037f IR: Removes non-opsize AtomicAdd/Sub/And/Or
These were unused
2023-08-28 07:12:31 -07:00
Ryan Houdek 6dcfd6eb73 IR: Removes non-opsize AtomicXor 2023-08-28 07:12:31 -07:00
Ryan Houdek 4f8a63459c IR: Removes non-opsize AtomicSwap 2023-08-28 07:12:31 -07:00
Ryan Houdek ab230bf527 IR: Removes non-opsize AtomicFetchAdd 2023-08-28 07:12:31 -07:00
Ryan Houdek 8ba9613972 IR: Removes non-opsize AtomicFetchSub 2023-08-28 07:12:31 -07:00
Ryan Houdek 9370e30af4 IR: Removes non-opsize AtomicFetchAnd 2023-08-28 07:12:31 -07:00
Ryan Houdek 25ce57ef92 IR: Removes non-opsize AtomicFetchOr 2023-08-28 07:12:31 -07:00
Ryan Houdek 436e0f6f86 IR: Removes non-opsize AtomicFetchXor 2023-08-28 07:12:31 -07:00
Ryan Houdek f44c6f394f IR: Removes non-opsize AtomicFetchNeg 2023-08-28 07:12:31 -07:00
Ryan Houdek 5415e4c95f Merge pull request #3017 from Sonicadvance1/remove_implicit_ops_pt1
Remove implicit sized IR ops part 1
2023-08-28 07:11:44 -07:00
Ryan Houdek bb1362b2bf OpcodeDispatcher: Optimize Get{Src,Dst}Size
These functions are called a lot....a lot a lot.
Optimize these in to a couple of ALU operations instead of a whole
table lookup. Confirming with output assembly that this becomes more
optimal.
2023-08-28 05:29:56 -07:00
Ryan Houdek 51baf594c9 OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek 48669b7006 IR: Removes implicit sized orlshl/orlshr 2023-08-28 05:04:32 -07:00
Ryan Houdek 16481c0e55 IR: Removes implicit sized rev 2023-08-28 05:04:32 -07:00
Ryan Houdek b00310674a IR: Removes implicit sized not 2023-08-28 05:04:32 -07:00
Ryan Houdek 5768444ce9 IR: Removes implicit sized abs 2023-08-28 05:04:32 -07:00
Ryan Houdek 1f1473eb74 IR: Removes implicit sized neg 2023-08-28 05:04:32 -07:00
Ryan Houdek cccd7001cb IR: Removes implicit sized ashr 2023-08-28 05:04:32 -07:00
Ryan Houdek ce8392d5ae IR: Removes implicit sized ror 2023-08-28 05:02:01 -07:00
Ryan Houdek 386cf36cfd IR: Removes implicit sized sbfe
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.

We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek 0ddc23a5c9 IR: Removes implicit sized Popcount 2023-08-28 05:02:01 -07:00
Ryan Houdek b95648a4ab IR: Removes implicit sized FindMSB 2023-08-28 05:02:01 -07:00
Ryan Houdek bf18672999 IR: Removes implicit sized FindLSB 2023-08-28 05:02:01 -07:00
Ryan Houdek f405c1be69 IR: Removes inverted EntrypointOffset/InlineEntrypointOffset 2023-08-28 05:02:01 -07:00
Ryan Houdek f3cd115fa7 IR: Removes implicit sized CountLeadingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek a14720e130 IR: Removes implicit sized FindTrailingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek 86ed909de8 IR: Removes implicit sized DIV/REM 2023-08-28 05:02:00 -07:00
Ryan Houdek ac7e75c06b IR: Removes implicit sized UDIV/UREM 2023-08-28 05:02:00 -07:00
Ryan Houdek ec3e7ceeb5 IR: Removes implicit sized EXTR 2023-08-28 05:02:00 -07:00
Ryan Houdek c8c8ddbd4f IR: Removes implicit sized PDEP/PEXT 2023-08-28 05:02:00 -07:00
Ryan Houdek 35013bda37 IR: Adds helper to convert between an integer size and IR::OpSize
This is a nop operation and will get optimized away in release builds.
2023-08-28 05:02:00 -07:00
Ryan Houdek 62a9a075b7 Arm64: Leave a comment that 32-bit division shouldn't leave garbage in upper 64-bits 2023-08-28 05:02:00 -07:00
Ryan Houdek ea6d068cc5 IR: Removes implicit sized LDIV/LREM 2023-08-28 05:02:00 -07:00
Ryan Houdek 5013473ec0 IR: Removes implicit sized LUDIV/LUREM 2023-08-28 05:02:00 -07:00
Alyssa Rosenzweig b801bacb36 Merge pull request #3016 from Sonicadvance1/update_instcountci
InstCountCI: Update for previous changes
2023-08-28 07:42:20 -04:00
Ryan Houdek f02d291da1 InstCountCI: Update for previous changes
Missed thsi in the SRA PR
2023-08-27 23:41:41 -07:00
Ryan Houdek 6f2b3e76ac Merge pull request #3013 from Sonicadvance1/32bit_sra
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek 1d7c280367 Merge pull request #3012 from Sonicadvance1/optimize_movmskps
OpcodeDispatcher: Optimizes SSE movmaskps
2023-08-27 21:29:04 -07:00
Ryan Houdek d6864871ed InstCountCI: Update for movmskps optimization 2023-08-27 21:07:20 -07:00
Ryan Houdek 514a8223d9 OpcodeDispatcher: Optimizes SSE movmaskps
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.

I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek 8d110738ac IR: Add option to disable vector shift range clamping
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.

Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.

When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Mai 590345b125 Merge pull request #3014 from Sonicadvance1/remove_implicit_alu_ops
IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
2023-08-27 20:25:11 -04:00
Mai e2ac97f4db Merge pull request #3015 from Sonicadvance1/instcount_ci_rorx_abovemax
InstCountCI: Test rorx at max mask size
2023-08-27 20:21:45 -04:00
Ryan Houdek a216ea465c InstCountCI: Test rorx at max mask size
Just to ensure we capture the nop nature of passing in the rotate amount
of the operating size.
2023-08-27 01:43:47 -07:00
Ryan Houdek e4bb0df486 IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.

This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.

Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek 7146691360 IR/Passes/RA: Enable SRA for 32-bit GPRs
Noticed that we hadn't ever enabled this, which was a concern when our
GPR operations weren't as strict about leaving garbage in the upper bits
when operating as a 32-bit operation.

Now that our ALU operations are more strict about enforcing upper bit
zeroing we can enable this.

This causes Half-Life: Source FPS to get to > 200FPS finally. Causes
significant performance improvements for 32-bit games because we're no
longer redundantly moving registers before and after every operation.
Causing a bunch of 3-4 instruction sequences to convert to 1.
2023-08-26 18:22:50 -07:00
Ryan Houdek 6cb1afc2ad unittests/asm: Adds ADC/SBB test 2023-08-26 18:22:50 -07:00
Ryan Houdek 572d6cd3e6 OpcodeDispatcher: Fixes ADC and SBB 2023-08-26 18:22:50 -07:00
Ryan Houdek fcc37bf6a8 OpcodeDispatcher: Fixes RCR and ADOX 32-bit
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek 8f7925d06f Arm64: Simple typo fix 2023-08-26 18:22:50 -07:00
Ryan Houdek eace648fa9 OpcodeDispatcher: Fixes bug in UMUL
This was trying to operating on a 32-bit value but BFE the upper
32-bits.

Actually fixes this so it is operating on the 64-bit multiply result.
2023-08-26 18:22:50 -07:00
Ryan Houdek a4ac21a4e4 OpcodeDispatcher: Fixes bug in GetRFLAG with CachedNZCV
This was only operating at byte size but it was attempting to get bit
offsets at greater than operating size.
Change operating size over to 64-bit.
2023-08-26 18:22:50 -07:00
Ryan Houdek a01e69092d Arm64: Ensure Bfe and Sbfe operate at 32-bit or 64-bit op size
For Sbfe at least it ensures the upper bits don't get filled with
garbage.
Bfe it doesn't change behaviour but best to be correct.
2023-08-26 18:22:50 -07:00
Ryan Houdek 2fde2140ef Arm64: Ensure assert is testing correct array 2023-08-26 18:22:50 -07:00
Ryan Houdek 547daf8b6e Merge pull request #2857 from Sonicadvance1/fix_ra_validation
IR: Fixes RAValidation for 32-bit applications
2023-08-26 18:22:18 -07:00
Ryan Houdek e10afefb2b IR: Fixes RAValidation for 32-bit applications
RAValidation was making an assumption that GPR register class would only
have up to 16 registers for either SRA or dynamic registers.

When running a 32-bit application we allow 17 GPRs to be dynamically
allocated, since we can take 8 back from SRA in that case.

Just split the two classes in the RAValidation pass since they will
never overlap their allocation.

Fixes validation in `32Bit_Secondary/15_XX_0.asm` locally that changed
behaviour due to tinkering.
2023-08-26 16:18:03 -07:00
Mai 0195bb6e5a Merge pull request #3010 from Sonicadvance1/exit_group_quit
Linux: Call exit_group when application tries
2023-08-25 21:13:59 -04:00
Ryan Houdek edd27a2723 Linux: Call exit_group when application tries
When the application calls exit_group we no longer need to care about
cleanup because the entire process group is leaving.

Just immediately call exit group and get out. Might revisit this in the
future.

Fixes #2752
2023-08-25 17:41:57 -07:00
Mai 7c6660f634 Merge pull request #3009 from Sonicadvance1/fix_frint64
X8764: Ensure frndint uses host rounding mode
2023-08-25 20:01:01 -04:00
Ryan Houdek 9ba46f429e X8764: Ensure frndint uses host rounding mode
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.

This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.

In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.

[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00
Mai db3dc3edf4 Merge pull request #3003 from Sonicadvance1/optimize_pshuf
OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
2023-08-25 16:36:52 -04:00
Ryan Houdek c26798af83 InstCountCI: Update for shuffles! 2023-08-25 12:59:40 -07:00
Ryan Houdek a76c2c57b0 OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek 7f63d87295 IR: Adds support for new LoadNamedVectorIndexedConstant IR 2023-08-25 12:59:40 -07:00
Mai bf12f08218 Merge pull request #3002 from Sonicadvance1/optimize_movmaskpd
OpcodeDispatcher: Optimize 128-bit movmaskpd
2023-08-25 08:50:28 -04:00
Mai 1f7d138d2a Merge pull request #3008 from Sonicadvance1/optimize_movddup
OpcodeDispatcher: Optimize movddup from register
2023-08-25 08:48:40 -04:00
Mai f36f07055a Merge pull request #3007 from Sonicadvance1/optimize_cvtdq2pd
OpcodeDispatcher: Optimize cvtdq2pd from register source
2023-08-25 08:47:47 -04:00
Mai 30a1a382c4 Merge pull request #3006 from Sonicadvance1/optimize_movq
OpcodeDispatcher: Optimizes movq
2023-08-25 08:46:02 -04:00
Mai 631655dd81 Merge pull request #3005 from Sonicadvance1/nontemporalmoves
OpcodeDispatcher: Optimize nontemporal moves
2023-08-25 08:44:48 -04:00
Mai 0ef439f830 Merge pull request #3004 from Sonicadvance1/optimize_scalar_cvt
OpcodeDispatcher: Generate more optimal code for scalar GPR converts
2023-08-25 08:43:56 -04:00
Ryan Houdek 2aba628a24 InstCountCI: Update for movddup optimization 2023-08-25 03:44:30 -07:00
Ryan Houdek 2fbcf2e4a9 OpcodeDispatcher: Optimize movddup from register
This is now optimal
2023-08-25 03:44:10 -07:00
Ryan Houdek 7367f166b0 InstCountCI: Update for cvtdq2pd optimization 2023-08-25 03:39:45 -07:00
Ryan Houdek 00124205e5 OpcodeDispatcher: Optimize cvtdq2pd from register source
This is now optimal
2023-08-25 03:39:24 -07:00
Ryan Houdek 1a3a59dd38 InstCountCI: Update for movq optimization 2023-08-25 03:30:52 -07:00
Ryan Houdek 81281e2115 OpcodeDispatcher: Optimizes movq
Removes a redundant move between registers and makes it optimal.
Also removes a couple redundant moves on the avx version.
2023-08-25 03:29:58 -07:00
Ryan Houdek a901e5f3df InstCountCI: Update for optimized cvtsi2s{s,d} 2023-08-25 03:19:11 -07:00
Ryan Houdek 1cb2b084b3 OpcodeDispatcher: Generate more optimal code for scalar GPR converts
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
   FPR and converted it in-place.

These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek 189b0da68f JIT/Int: Add support for scalar conversion as well 2023-08-25 03:19:11 -07:00
Ryan Houdek 9fa877e512 InstCountCI: Update for move temporal optimization 2023-08-25 03:13:28 -07:00
Ryan Houdek f3679a99ec OpcodeDispatcher: Optimize nontemporal moves
These are now optimal.
2023-08-25 03:13:10 -07:00
Ryan Houdek 62156f2152 ARM64JIT: Adds support for scalar cvt 2023-08-25 02:34:30 -07:00
Ryan Houdek 3808f97283 InstCountCI: Update for movmaskpd optimization 2023-08-24 17:27:54 -07:00
Ryan Houdek 3a6d25f56a unittests/ASM: Update movmskpd test to include an edge of garbage but no sign bit 2023-08-24 17:27:54 -07:00
Ryan Houdek 9a54898429 OpcodeDispatcher: Optimize 128-bit movmaskpd
I'd consider this optimal now.

Thanks to @dougallj for the optimization idea again!
2023-08-24 17:27:54 -07:00
Ryan Houdek 80d871fb18 Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek 4d58ec1025 Merge pull request #3000 from Sonicadvance1/optimize_cvt
OpcodeDispatcher: Optimize MMX conversion operation
2023-08-24 16:09:04 -07:00
Ryan Houdek 60ad76732a InstCountCI: Update for cvtps2pd optimization 2023-08-24 15:56:02 -07:00
Ryan Houdek 3731e6d88b OpcodeDispatcher: Optimize cvtps2pd
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek e3812f9c3c InstCountCI: Update tests for mmx optimization 2023-08-24 15:46:54 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek 72ce7ddf2d Arm64: Optimize CVT operations for 64-bit variants
Using 128-bit converts for 64-bit versions cuts their throughput in half
on Cortex. Ensure we use the 64-bit version when possible.
2023-08-24 15:45:07 -07:00
Ryan Houdek e025d32531 Merge pull request #2999 from lioncash/catch
Externals: Update Catch2 to v2.13.10
2023-08-24 15:19:23 -07:00
Ryan Houdek 200dbdd0e6 Merge pull request #2998 from Sonicadvance1/optimize_addsub_fcma
OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
2023-08-24 15:19:06 -07:00
Lioncache 989fe22e2d Externals: Update Catch2 to v2.13.10
Updates it to the latest v2 branch tag
2023-08-24 18:04:59 -04:00
Ryan Houdek 4e7adeec85 InstCountCI: Adds new files for FCMA 2023-08-24 15:00:42 -07:00
Ryan Houdek a1210f892a OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.

Needs #2994 merged first.

Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek ba01eac467 IR: Adds support for ARM's FCMA FCADD instruction 2023-08-24 15:00:41 -07:00
Ryan Houdek c5d147322f HostFeatures: Adds support for FCMA 2023-08-24 15:00:41 -07:00
Ryan Houdek df99b7b9b6 Merge pull request #2994 from Sonicadvance1/cache_namedvectorconstants
OpcodeDispatcher: Cache named vector constants in the block
2023-08-24 15:00:05 -07:00
Ryan Houdek 565b30e15e OpcodeDispatcher: Cache named vector constants in the block
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.

Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.

Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Mai ab83ab42dd Merge pull request #2993 from Sonicadvance1/addsub_opt
OpcodeDispatcher: Optimize AddSubP{S,D}
2023-08-24 17:44:41 -04:00
Ryan Houdek b0ec4197ba Merge pull request #2997 from lioncash/fmt
Externals: Update fmt to 10.1.0
2023-08-24 14:19:18 -07:00
Lioncache 9035a29906 Externals: Update fmt to 10.1.0
Updates fmt to the latest version.
2023-08-24 17:01:40 -04:00
Ryan Houdek b547550442 InstCountCI: Update for addsubp 2023-08-23 20:33:55 -07:00
Ryan Houdek f300196d90 OpcodeDispatcher: Optimize AddSubP{S,D}
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.

Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Ryan Houdek c5f358b47a Merge pull request #2992 from lioncash/doc
x86_64/MemoryOps: Fix mislabeled IR op messages
2023-08-23 20:05:39 -07:00
Lioncache 42ccc18606 x86_64/MemoryOps: Fix mislabeled IR op messages 2023-08-23 22:54:36 -04:00
Mai 66c6f96120 Merge pull request #2990 from Sonicadvance1/optimize_pmulh
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek 1aa2c534f7 Merge pull request #2991 from lioncash/pcl
OpcodeDispatcher: Remove redundant moves from PCLMULQDQ and AES operations
2023-08-23 18:47:52 -07:00
Ryan Houdek e9f292462a InstCountCI: Update for pmulh{u,}w optimization 2023-08-23 18:38:05 -07:00
Ryan Houdek 77b6d854b9 OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Ryan Houdek 05b9651279 IR: Implements new vector multiply returning high bits
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 8a622a3c1a OpcodeDispatcher: Remove redundant moves from VAESEnc
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache f4848fd1a7 OpcodeDispatcher: Remove redundant moves from VAESEncLast
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache d37ce08ae9 OpcodeDispatcher: Remove redundant move from VAESDec
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache a6f1a9f8e8 OpcodeDispatcher: Remove redundant moves from VAESDecLast
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache 410e99ba09 OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek 6e4765d48b Merge pull request #2989 from lioncash/ins
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
2023-08-23 18:09:15 -07:00
Ryan Houdek 172c8f3ba6 Merge pull request #2988 from lioncash/half
Arm64/ConversionOps: Add missing half-precision conversions to scalar functions
2023-08-23 17:56:23 -07:00
Lioncache 203a2b1105 Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
If Dst and DestVector alias one another, then we don't need to
move the vector unnecessarily.
2023-08-23 20:50:21 -04:00
Ryan Houdek 4297e13fcf Merge pull request #2986 from lioncash/ext
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
2023-08-23 17:40:50 -07:00
Ryan Houdek 5b8a0f1e0d Merge pull request #2987 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from VSQXTN2/VSQXTUN2/VSQSHL/VSRSHR
2023-08-23 17:40:39 -07:00
Lioncache 5ad56ad52e Arm64/ConversionOps: Add missing half-precision operations to Float_FromGPR_S
Provides parity with vector operations.
2023-08-23 20:36:34 -04:00
Lioncache 24e7baf28f Arm64/ConversionOps: Add missing half-precision conversions to Float_FToF
Provides parity with the vector conversion operations.
2023-08-23 20:31:49 -04:00
Lioncache b248ae4c04 Arm64/VectorOps: Remove redundant moves from SVE SQSHL
We don't need to emit a move if the destination and source alias.
2023-08-23 20:12:09 -04:00
Lioncache 95bea864cf Arm64/VectorOps: Remove redundant moves from SVE SRSHR
We don't need to perform a move is the destination aliases
the source vector to be shifted.
2023-08-23 20:10:51 -04:00
Lioncache 47c4507bb6 Arm64/VectorOps: Remove redundant moves from SVE VSQXTUN2
We don't need to perform a move if the destination aliases the lower vector.
2023-08-23 20:10:29 -04:00
Lioncache 5ea0b6db28 Arm64/VectorOps: Remove redundant moves from SVE VSQXTN2
We don't need to perform a move if the destination aliases the
lower vector.
2023-08-23 20:02:28 -04:00
Mai ee10153d14 Merge pull request #2984 from Sonicadvance1/optimize_pack
OpcodeDispatcher: Use new IR ops for pack instructions
2023-08-23 20:02:16 -04:00
Lioncache d0d94adabe Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
We don't need to do any moves here is the destination aliases the
lower bits.
2023-08-23 19:56:10 -04:00
Ryan Houdek 926b8c2c97 Merge pull request #2985 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from SVE variable/immediate/vector shifts when possible
2023-08-23 16:41:39 -07:00
Lioncache 18ebcdc9de Arm64/VectorOps: Remove redundant moves in VUshrNI2
If the destination and VectorLower alias, then we don't need
to emit a movprfx.
2023-08-23 18:52:32 -04:00
Lioncache f31a9a52e6 Arm64/VectorOps: Remove redundant moves from SVE immediate vector shifts when possible
If the destination and source vector alias one another, then the
operation can largely be done in place.
2023-08-23 18:36:21 -04:00
Lioncache 03504a5f8c Arm64/VectorOps: Remove redundant moves from SVE vector shifts when possible
If the destination and the vector to be shifted alias, then we can
avoid needing to move some data around.
2023-08-23 18:24:58 -04:00
Lioncache d29b4de1ee Arm64/VectorOps: Remove redundant moves from SVE variable vector register shifts when possible
In the event that the destination and the vector to be shifted
alias one another, then we can skip the movprfx, since it's not
necessary.
2023-08-23 18:24:53 -04:00
Ryan Houdek 5e20be756e InstCountCI: Update for pack instruction optimization 2023-08-23 15:16:54 -07:00
Ryan Houdek fc4559d3c4 OpcodeDispatcher: Use new IR ops for pack instructions
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Ryan Houdek c508570da0 IR: Implements VSQXT{U,}NPair operations
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Ryan Houdek ec6548e302 Merge pull request #2983 from lioncash/bsl
Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
2023-08-23 15:11:34 -07:00
Lioncache d5e145c4b0 Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
If the destination and true vector alias one another, then we can
perform the operation in place instead of moving data around.
2023-08-23 17:54:10 -04:00
Ryan Houdek 350bca97c6 Merge pull request #2982 from lioncash/imin
Arm64/VectorOps: Remove redundant moves from SVE V{S,U}Min/V{S,U}Max when possible
2023-08-23 14:53:15 -07:00
Ryan Houdek 226405880f Merge pull request #2981 from lioncash/fmin
Arm64/VectorOps: Remove redundant moves from SVE VFMin/VFMax when possible
2023-08-23 14:46:03 -07:00
Lioncache 37a8cb6821 Arm64/VectorOps: Remove redundant moves from SVE VSMax when possible
When the destination and first source alias one another, then we
can perform the operation in place instead of moving data around.
2023-08-23 17:34:52 -04:00
Lioncache fe2c7dbf97 Arm64/VectorOps: Remove redundant moves from SVE VUMax when possible
When the destination and source alias one another, then we
can perform the operation in place without needing to move
data around.
2023-08-23 17:32:17 -04:00
Lioncache 787b4f37fb Arm64/VectorOps: Remove redundant moves from SVE VSMin when possible
When the destination and first source alias one another, then we can
perform the operation in place without moving any data.
2023-08-23 17:30:08 -04:00
Lioncache c3faa019f5 Arm64/VectorOps: Remove redundant moves from SVE VUMin when possible
If the destination and first source alias, then we can perfom the operation
in place.
2023-08-23 17:27:52 -04:00
Ryan Houdek da098d8204 Merge pull request #2979 from lioncash/div
Arm64/VectorOps: Remove moves from SVE VFDiv if possible
2023-08-23 14:22:46 -07:00
Lioncache 149852b122 Arm64/VectorOps: Remove redundant moves from SVE VFMax if possible
If Dst and Vector1 alias one another, then the operation can be
performed in place instead of moving data around.
2023-08-23 17:18:41 -04:00
Lioncache ecf02846e6 Arm64/VectorOps: Remove redundant moves from SVE VFMin is possible
If Dst and Vector1 alias one another, then we can do the merging move
in place instead of shuffling data around.
2023-08-23 17:16:25 -04:00
Ryan Houdek 2501ebc1cd Merge pull request #2980 from lioncash/avg
Arm64/VectorOps: Remove redundant moves from SVE VURAvg if possible
2023-08-23 14:05:51 -07:00
Lioncache a8f7529847 Arm64/VectorOps: Remove moves from SVE VFDiv if possible
Given the operation is:

Dst = Vector1 / Vector2

If Dst and Vector1 alias one another, then we can just perform
the division as is without any moving of data around.
2023-08-23 16:53:56 -04:00
Lioncache 8431ab43a0 Arm64/VectorOps: Remove redundant moves from SVE VURAvg if possible
If Dst and Vector1 alias one another, then we can perform the operation
without needing to move any data around.
2023-08-23 16:51:17 -04:00
Mai 4b06069c0d Merge pull request #2972 from Sonicadvance1/optimize_scalar_mov
OpcodeDispatcher: Optimizes scalar movd/movq
2023-08-23 16:30:45 -04:00
Ryan Houdek b646f4b781 Merge pull request #2978 from lioncash/misc
OpcodeDispatcher: Remove redundant moves from remaining AVX ops
2023-08-23 13:18:25 -07:00
Ryan Houdek 38853c2a9b InstCountCI: Update for vmovd/vmovq optimization 2023-08-23 12:56:11 -07:00
Ryan Houdek 8836ab8988 OpcodeDispatcher: Optimizes scalar movd/movq
MMX and SSE versions are now optimal.
2023-08-23 12:56:11 -07:00
Ryan Houdek a40526a541 Merge pull request #2977 from lioncash/pack
OpcodeDispatcher: Remove redundant moves from VPACKUSOP/VPACKSSOp
2023-08-23 12:33:05 -07:00
Lioncache ea9747289a OpcodeDispatcher: Remove redundant moves from remaining AVX ops
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 15:31:59 -04:00
Lioncache 735e2060a3 OpcodeDispatcher: Remove redundant moves from VPACKUSOP/VPACKSSOp
Zero-extension will occur automatically if necessary.
2023-08-23 15:09:57 -04:00
Ryan Houdek 86ef6fe48d Merge pull request #2976 from lioncash/mov
OpcodeDispatcher: Remove unnecessary moves from AVX move ops where applicable
2023-08-23 12:02:05 -07:00
Lioncache 8e7e91d61f OpcodeDispatcher: Remove redundant moves in VMOVLPOp
Zero-extension will automatically occur upon storing if necessary.
2023-08-23 14:23:57 -04:00
Lioncache bcba3700c8 OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
Zero-extension will automatically occur if necessary upon storing.

We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache 7d05797e82 OpcodeDispatcher: Remove redundant moves from VMOVHPOp
Zero-extension will automatically occur if necessary upon storing.
2023-08-23 14:14:35 -04:00
Lioncache c409ea78bc OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Ryan Houdek a62ba75ede Merge pull request #2975 from lioncash/scalar
Arm64/ConversionOps: Add scalar support to Vector_FToI
2023-08-23 10:57:38 -07:00
Lioncache 4a7ef3da13 OpcodeDispatcher: Remove unnecessary moves in AVXVectorRound
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 13:41:56 -04:00
Lioncache 990b70dcd6 OpcodeDispatcher: Use scalar rounding for scalar round instructions 2023-08-23 13:34:01 -04:00
Lioncache 393cea2e8b Arm64/ConversionOps: Correct AdvSIMD round-to-nearest Vector_FToI case
This was previously using frinti, which uses the host rounding mode, rather
than round to nearest.
2023-08-23 13:27:44 -04:00
Ryan Houdek 6624f50abf Merge pull request #2974 from lioncash/extend
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
2023-08-23 10:25:56 -07:00
Ryan Houdek cd1f401363 Merge pull request #2973 from lioncash/vfcmp
OpcodeDispatcher: Remove unnecessary moves in AVXVFCMPOp
2023-08-23 10:25:16 -07:00
Lioncache e89321dc60 Arm64/ConversionOps: Add scalar support to Vector_FToI
This can be used for the scalar conversions instead of always using the
vector variants.
2023-08-23 13:23:07 -04:00
Lioncache d99bcbf01b OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
Zero-extension will already occur if necessary upon storing.

Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Mai 0819338dbf Merge pull request #2970 from Sonicadvance1/optimize_pminmax
Arm64: Optimize VFMin/VFMax
2023-08-23 12:39:37 -04:00
Lioncache f516aed4b7 OpcodeDispatcher: Remove unnecessary moves in AVXVFCMPOp
Zero-extension will already occur if necessary upon storing.
2023-08-23 12:37:45 -04:00
Ryan Houdek 76430baf88 Merge pull request #2971 from lioncash/blend
OpcodeDispatcher: Remove redundant moves in AVX blend special cases
2023-08-22 21:42:07 -07:00
Lioncache 3858e4124b OpcodeDispatcher: Remove redundant moves in AVX blend special cases
Zero-extension will happen if necessary upon storing.
2023-08-23 00:08:33 -04:00
Mai 819fe110da Merge pull request #2967 from Sonicadvance1/optimize_storeelement
OpcodeDispatcher: Optimize MOVHP{S,D}
2023-08-23 00:01:25 -04:00
Ryan Houdek 5db5944ad2 InstCountCI: Update for min/max optimization 2023-08-22 21:00:27 -07:00
Ryan Houdek f0b1030e54 Arm64: Optimize VFMin/VFMax
We can be more optimal on the selects. This makes 3DNow! and SSE packed
min/max operations optimal.
2023-08-22 20:59:11 -07:00
Ryan Houdek adfd6787c0 Merge pull request #2969 from lioncash/insert
OpcodeDispatcher: Remove unnecessary moves from AVX inserts
2023-08-22 20:51:55 -07:00
Ryan Houdek b35ad8d8ed InstCountCI: Update for movhp{s,d} optimization 2023-08-22 20:42:24 -07:00
Ryan Houdek 0ee2579a5e OpcodeDispatcher: Optimize MOVHP{S,D}
Loads can turn in to element Loads.
Stores can turn in to element stores.

These four instruction variants are now optimal.
2023-08-22 20:42:24 -07:00
Ryan Houdek 6aa2cab41c IR: Implement support for vector store element
Matches ARM64 ST1 semantics
2023-08-22 20:42:24 -07:00
Mai bb2f7107cd Merge pull request #2963 from Sonicadvance1/optimize_loadelement
OpcodeDispatcher: Optimize MOVLP{S,D} loads
2023-08-22 23:41:59 -04:00
Lioncache c33f3ff8df OpcodeDispatcher: Remove unnecessary moves from AVX inserts
We already zero-extend on stores when necessary.
2023-08-22 23:29:40 -04:00
Ryan Houdek 5d44a445dd Merge pull request #2968 from lioncash/shift2
OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
2023-08-22 20:23:12 -07:00
Ryan Houdek 519c670374 InstCountCI: Update for load element optimization
Adds movhlps special case which was missed before.
2023-08-22 20:15:16 -07:00
Ryan Houdek de239cde67 OpcodeDispatcher: Optimize MOVLP{S,D} loads
This now uses the new load element IR operation and makes these
instructions optimal.

LRPCPC3 will introduce instructions in the future for TSO emulation to
help these operations, but that doesn't exist today.
2023-08-22 20:15:16 -07:00
Ryan Houdek 5e57ec94cf IR: Implement support for vector load element
Matches Arm64 LD1 semantics.
2023-08-22 20:15:16 -07:00
Lioncache 2f5fae7677 OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
Zero-extension will occur automatically when necessary upon storing.
2023-08-22 23:06:53 -04:00
Ryan Houdek fb65fb29c7 Merge pull request #2966 from lioncash/shift
OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
2023-08-22 20:02:14 -07:00
Lioncache 8f8062eb4e OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
These zero-extensions will occur automatically when applicable.
2023-08-22 22:50:10 -04:00
Ryan Houdek 36a54183f5 Merge pull request #2965 from lioncash/mov
OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
2023-08-22 19:38:46 -07:00
Lioncache e5f5629ffc OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Ryan Houdek 4443c667ec Merge pull request #2964 from Sonicadvance1/missed_optimal
InstCountCI: Fix some mislabeled instructions
2023-08-22 19:05:14 -07:00
Ryan Houdek ead141fd90 Merge pull request #2962 from lioncash/variable
OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
2023-08-22 18:52:47 -07:00
Ryan Houdek 2c64523317 InstCountCI: Fix some mislabeled instructions
These were all optimal. Fixed now.
2023-08-22 18:47:17 -07:00
Lioncache 2b071e282e OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
We already zero-extend on a store if necessary.
2023-08-22 21:20:34 -04:00
Ryan Houdek 14144523f7 Merge pull request #2961 from lioncash/minpos
OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
2023-08-22 18:19:46 -07:00
Ryan Houdek 1f2c5fc6c6 Merge pull request #2960 from lioncash/index
Arm64: Optimize SVE VInsElement
2023-08-22 18:19:13 -07:00
Mai 42200bf7b6 Merge pull request #2959 from Sonicadvance1/movlpd_store
X86Tables: Optimize MOVLPD stores
2023-08-22 21:02:23 -04:00
Lioncache fa17d9fae9 OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
We already do a zero-extend if necessary in StoreResult.

This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Lioncache 398a70312e Arm64: Optimize SVE VInsElement
This can be done without storing any data to memory and also
compressing the amount of instructions being used.

Thanks to @dougallj for the optimization suggestions.
2023-08-22 20:36:29 -04:00
Ryan Houdek 76afc653e8 InstCountCI: Update for movlpd store optimization 2023-08-22 17:34:15 -07:00
Ryan Houdek d3ed9766e8 X86Tables: Optimize MOVLPD stores
Just use the full register size and store the lower bits.
2023-08-22 17:33:34 -07:00
Ryan Houdek ed7f1b017d Merge pull request #2958 from Sonicadvance1/optimize_phminpos
OpcodeDispatcher: Optimize phminposuw
2023-08-22 16:58:27 -07:00
Ryan Houdek fb60f9e406 InstCountCI: Update for phminposuw optimization 2023-08-22 16:29:06 -07:00
Ryan Houdek c795d42d21 OpcodeDispatcher: Optimize phminposuw
I would now consider the XMM version of this to be optimal.

Thanks to @rygorous for giving the idea for how to optimize this!
2023-08-22 16:29:06 -07:00
Ryan Houdek bbf9cb9d52 IR: Implements new VRev32 and LoadNamedVectorConstant ops
VRev32 matches Arm64 semantics directly.
LoadNamedVectorConstant allows FEX to quickly load "named constants".
This will allow us to have specific hardcoded vector constant values
that we can load with a ldr(State)+ldr(Value) and will be more abused in
the future.
This also allows us to do a very simple optimization in the future where
we can optimize away redundant loads of these loads if they are used
multiple times in the same block. (Not implemented here).
2023-08-22 16:29:06 -07:00
Mai 6c7933e7b1 Merge pull request #2957 from Sonicadvance1/optimize_pfnacc
OpcodeDispatcher: Optimize PFNACC
2023-08-22 10:10:11 -04:00
Ryan Houdek 364f084604 Merge pull request #2956 from Sonicadvance1/optimize_hsubp
OpcodeDispatcher: Optimize hsubp
2023-08-21 20:47:50 -07:00
Ryan Houdek ffa8f1e3dc Merge pull request #2955 from lioncash/sign
OpcodeDispatcher: Remove redundant move from VPSIGN
2023-08-21 20:47:41 -07:00
Ryan Houdek be5d5b06f8 InstCountCI: Update for pfnacc 2023-08-21 20:38:25 -07:00
Ryan Houdek ad6738939b OpcodeDispatcher: Optimize PFNACC
Turns out this can be even more optimal.
2023-08-21 20:38:18 -07:00
Mai 9df94d8a93 Merge pull request #2954 from Sonicadvance1/optimize_pmuludq
OpcodeDispatcher: Optimize pmuludq
2023-08-21 23:33:45 -04:00
Ryan Houdek 9ae85b2251 InstCountCI: Update for hsubp 2023-08-21 20:24:37 -07:00
Ryan Houdek dcb3e4ee86 OpcodeDispatcher: Optimize hsubp
This makes the SSE version optimal.
This dramatically improves the AVX version as well.
2023-08-21 20:22:14 -07:00
Lioncache dbbe6288de OpcodeDispatcher: Remove redundant move from VPSIGN
StoreResult will already zero-extend if the vector is 128-bit.
2023-08-21 23:11:07 -04:00
Ryan Houdek de1f75f7a5 InstCountCI: Update for pmuludq 2023-08-21 20:08:19 -07:00
Ryan Houdek 1563398d2c OpcodeDispatcher: Optimize pmuludq
MMX version was already optimal, SSE version is now also.
AVX version is significantly improved.
2023-08-21 20:07:35 -07:00
Ryan Houdek 71984fc0ea Merge pull request #2953 from lioncash/scalar
OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
2023-08-21 19:51:23 -07:00
Lioncache 920a0fb132 OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
Our store will already zero-extend if the vector is 128-bit.
2023-08-21 22:39:07 -04:00
Ryan Houdek 3c88671cca Merge pull request #2952 from lioncash/alu
OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
2023-08-21 19:26:22 -07:00
Lioncache ce8169794f OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
We already zero-extend on a store if we have 256-bit vectors and the stored
vector is 128-bit.
2023-08-21 22:13:29 -04:00
Mai 185e3bfcb6 Merge pull request #2950 from Sonicadvance1/optimize_pmaddwd
OpcodeDispatcher: Optimize pmaddwd
2023-08-21 21:39:30 -04:00
Mai 3c49b3238a Merge pull request #2949 from Sonicadvance1/optimize_phsub
OpcodeDispatcher: Optimize phsub
2023-08-21 21:39:02 -04:00
Ryan Houdek 2d7a3a578e Merge pull request #2931 from Sonicadvance1/optimize_psign
Optimize PSIGN and VBSL
2023-08-21 18:30:55 -07:00
Ryan Houdek 3a2a576c35 Merge pull request #2951 from lioncash/shift
OpcodeDispatcher: Handle zero immediate shifts better
2023-08-21 18:20:18 -07:00
Ryan Houdek fe4de26250 InstCountCI: Updates tests for pmaddwd optimization
MMX and SSE implementations are now optimal.
2023-08-21 17:55:20 -07:00
Ryan Houdek 869136b907 OpcodeDispatcher: Optimize pmaddwd
This is actually fairly trivial looking at it.
2023-08-21 17:53:32 -07:00
Lioncache af8b6766d8 OpcodeDispatcher: Handle zero immediate shifts better
In the SSE and lower cases, we don't need to do anything,
since the value is already in the destination.
2023-08-21 20:46:41 -04:00
Ryan Houdek d283d1ba11 InstCountCI: Update for phsub optimization
MMX and SSE now optimal
2023-08-21 17:38:24 -07:00
Ryan Houdek 30e9beba51 JIT64: Fixes i32v2 unzips
This has just been broken since it was implemented. Turns out we had
never used these with XMM operations before today.
2023-08-21 17:38:24 -07:00
Ryan Houdek 8ee6262e5e Merge pull request #2946 from Sonicadvance1/optimize_mpsad
OpcodeDispatcher: Optimizes mpsadbw
2023-08-21 17:33:03 -07:00
Ryan Houdek 637a5d3b18 InstCountCI: Updates results from mpsadbw optimization
Pretty sure the SSE versions are optimal implementations.
The AVX versions still have spurious moves that can probably get
removed.
2023-08-21 16:57:13 -07:00
Ryan Houdek a29076244f OpcodeDispatcher: Optimizes mpsadbw
Two optimizations here:
1) The final VInsElement was generating three instructions
   - This itself could have been change to vzip, which would have
     removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
   instruction off the calculation.
   - addp odd elements and even elemnts first
   - Then transpose those elements
   - Then use one final addp to generate the result in the correct
     order.

The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
2023-08-21 16:56:41 -07:00
Ryan Houdek 445c43792b OpcodeDispatcher: Optimize phsub
This was...surprisingly bad. I blame myself entirely.
2023-08-21 16:54:35 -07:00
Mai f4f9b20f32 Merge pull request #2948 from Sonicadvance1/fix_newline
InstCountCI: Add newline to end of file
2023-08-21 19:50:53 -04:00
Mai 34a7feffe8 Merge pull request #2947 from Sonicadvance1/optimize_pmaddubsw
OpcodeDispatcher: Optimize SSE/AVX pmaddubsw
2023-08-21 19:50:10 -04:00
Ryan Houdek b559982515 InstCountCI: Update for newlines 2023-08-21 16:26:46 -07:00
Ryan Houdek ca96a25a7a InstCountCI: Add newline to end of file
This way these endlines don't constantly keep getting toggled.
2023-08-21 16:26:20 -07:00
Ryan Houdek 3f82c8cfe5 InstCountCI: Update for pmaddubsw
Not quite optimal but a heck of a lot better.
2023-08-21 16:22:00 -07:00
Ryan Houdek b35f4df798 OpcodeDispatcher: Optimize SSE pmaddubsw
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
2023-08-21 16:21:01 -07:00
Ryan Houdek e765bd8986 Merge pull request #2945 from lioncash/interpdata
Interpreter: Tie SSA data elements to supported vector width
2023-08-21 11:52:01 -07:00
Lioncache f1d020ce95 Interpreter: Tie SSA data elements to supported vector width
Now, if we ever increase our vector sizes, the allocated data elements
will follow suit without needing to remember to handle this as well.
2023-08-21 14:41:00 -04:00
Ryan Houdek 8b9ee997b2 Merge pull request #2944 from lioncash/interparray
Interpreter: Use alias for temporary vector data
2023-08-21 11:28:24 -07:00
Lioncache a9a7cbce21 Interpreter: Use alias for temporary vector data
Lets us extract the size into one location for easy size
changes in the future if necessary.
2023-08-21 14:12:55 -04:00
Ryan Houdek 084d102c9c Merge pull request #2943 from Sonicadvance1/wide_shifts
IR: Implements support for wide scalar shifts
2023-08-21 09:14:38 -07:00
Ryan Houdek 214dad25b5 InstCountCI: Updates instruction tables for wide shifts
Even on platforms without SVE these have improved slightly but the real
improvement comes from SVE.

Adds some new InstCountCI files for SVE128 enabled testing.

Also enables SVE128 in the VEX maps. Host features should probably
enable SVE128 when SVE256/AVX is enabled, but that isn't the case today.
2023-08-20 19:19:15 -07:00
Ryan Houdek a3bf952f2b IR: Implements support for wide scalar shifts
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.

With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).

This is a significant improvement even on platforms that only support
128-bit SVE.
2023-08-20 19:16:40 -07:00
Mai 5cc30bdf10 Merge pull request #2942 from Sonicadvance1/fix_nontelemetry_compilation
FEXInterpeter: Fixes compilation when telemetry is disabled
2023-08-20 17:16:13 -04:00
Ryan Houdek c3b4bfbc8d FEXInterpeter: Fixes compilation when telemetry is disabled
Oops, this has been broken for a while now.
2023-08-20 13:50:33 -07:00
Ryan Houdek b39e8ed5f0 InstCountCI: Update for improvements to psign/vbsl
Quite a few improvements here.
2023-08-20 13:46:01 -07:00
Ryan Houdek ac77986c44 IR: Implements support for saturating/rounding vector shifts 2023-08-20 13:38:23 -07:00
Ryan Houdek c5f5a03c68 OpcodeDispatcher: Optimize PSIGN
This dramatically improves the performance of the PSIGN instructions.
2023-08-20 13:38:23 -07:00
Ryan Houdek 1ed9ec63be Arm64: Optimize BSL when possible.
With ASIMD FEX would never optimize BSL out of fear if some registers
overlapped it would break things. So it had previously always moved to a
temporary first and then moved the result back out when done.

Now instead check upfront if any of the source registers overlap the
destination. If the destination register overlaps any of the three
sources we can bsl, bit, or bif depending on which register gets
overlapped.

Worst case the destination doesn't overlap any of the source registers
and still needs these moves.
2023-08-20 13:38:23 -07:00
Ryan Houdek a523858f66 Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek 9e4888c6a1 Merge pull request #2930 from Sonicadvance1/support_push
OpcodeDispatcher: Implement support for push IR operation
2023-08-20 10:25:26 -07:00
Ryan Houdek 277345d2a5 Merge pull request #2941 from lioncash/sha1
OpcodeDispatcher: Improve SHA1MSG1 output
2023-08-20 10:18:10 -07:00
Lioncache 8167626a07 x86_64/VectorOps: Properly handle VExtr element sizes other than bytes
We need to convert the index into a byte index.
2023-08-20 13:03:27 -04:00
Lioncache c3778a9729 OpcodeDispatcher: Improve SHA1MSG1 output
We can simplify these inserts down to a single EXT
2023-08-20 12:50:07 -04:00
Ryan Houdek 34722348e8 Merge pull request #2912 from Sonicadvance1/optimize_flag_clearing
OpcodeDispatcher: Minor optimization around clearing flags
2023-08-19 23:03:40 -07:00
Ryan Houdek 4286d44e92 Merge pull request #2940 from lioncash/crypto
Arm64/EncryptionOps: Use MOVI reg, #0 to zero vectors
2023-08-19 20:58:35 -07:00
Lioncache 1fbf193739 Arm64/EncryptionOps: Use MOVI reg, #0 to zero vectors
This is a little more optimal than XORing the vector by itself.
2023-08-19 23:20:36 -04:00
Ryan Houdek 8302ef8c22 InstCountCI: Update changed instructions due to flag clearing improvements
Some instructions reordered, but a bunch of operations had their number of instructions reduced as well
2023-08-19 20:19:39 -07:00
Ryan Houdek 92c3014aaa OpcodeDispatcher: Minor optimization around clearing flags
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.

Back to back bfi is actually less optimal due to dependency tracking.

With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
2023-08-19 20:14:37 -07:00
Mai 6960fca256 Merge pull request #2929 from Sonicadvance1/signaldelegator_getconfig
SignalDelegator: Allow getting the internal configuration
2023-08-19 23:12:06 -04:00
Mai affbcd2241 Merge pull request #2928 from Sonicadvance1/remove_x18_saving
Arm64Emitter: Stop saving and restoring platform register
2023-08-19 23:11:44 -04:00
Mai a2e5c231ae Merge pull request #2908 from Sonicadvance1/optimize_stc_clc
IR/ConstProp: Ensure that BFI with constant bitfields can optimize to Andn or Or
2023-08-19 23:10:48 -04:00
Ryan Houdek b973c193be Merge pull request #2939 from lioncash/round
OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
2023-08-19 19:29:01 -07:00
Ryan Houdek 4c409ea47d Merge pull request #2938 from lioncash/fcmp
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
2023-08-19 19:27:24 -07:00
Ryan Houdek ac53913c37 Merge pull request #2937 from lioncash/scalarunary
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
2023-08-19 19:22:10 -07:00
Ryan Houdek 2224c23c79 Merge pull request #2934 from lioncash/scalarfp
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
2023-08-19 19:20:20 -07:00
Ryan Houdek 6ce380d7a9 Merge pull request #2936 from lioncash/scalaralu
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorScalarALUOp
2023-08-19 19:09:06 -07:00
Lioncache 5152854b98 OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
2023-08-19 22:07:25 -04:00
Ryan Houdek 8dade7eea1 Merge pull request #2935 from lioncash/scalarfp2
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
2023-08-19 19:05:52 -07:00
Ryan Houdek add5baeba5 Merge pull request #2933 from lioncash/movss
OpcodeDispatcher: Remove some extraneous MOVs from VMOVSD/VMOVSS
2023-08-19 19:03:24 -07:00
Lioncache 6ba42e5cf1 OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
When dealing with scalar source registers, we can opt to not zero-extend
the vector and just perform the scalar operation and then insert the result.
2023-08-19 21:27:59 -04:00
Lioncache 343b00818d OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
We dealing with scalar vector sources, we don't need to zero-extend
the vector, and we can just use it as is.
2023-08-19 21:20:03 -04:00
Lioncache 09addb217a OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
2023-08-19 20:09:19 -04:00
Lioncache 4d1f002dea OpcodeDispatcher: Remove unnecessary moves in AVXVectorScalarALUOp
Same thing as the SSE variant, but for AVX.
2023-08-19 19:22:42 -04:00
Lioncache 1158ad7b2a OpcodeDispatcher: Remove unnecessary moves in VectorScalarALUOp
We can explicitly specify the vector width when working with a
vector source, so that we don't do any unnecessary zero-extending
on the element.
2023-08-19 19:15:13 -04:00
Lioncache 6907fdca6b OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
2023-08-19 18:46:08 -04:00
Lioncache c31329609f OpcodeDispatcher: Unify handling code for MOVSD and MOVSS
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
2023-08-19 17:44:45 -04:00
Lioncache 1fe8470933 OpcodeDispatcher: Remove extraneous moves from VMOVSS/VMOVSD xmm to mem case
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
2023-08-19 17:35:52 -04:00
Lioncache 99b5aaa426 OpcodeDispatcher: Remove extraneous moves in VMOVSS/VMOVSD register case
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
2023-08-19 17:35:12 -04:00
Ryan Houdek db60a2fd4b Merge pull request #2932 from lioncash/perm
OpcodeDispatcher: Handle broadcasting cases in VPERMQ/VPERMPD
2023-08-18 22:48:30 -07:00
Lioncache 84f228a75a x86_64/VectorOps: Simplify index handling in VInsElement
While we're in the area, we can simplify these cases down a little.
2023-08-19 01:28:45 -04:00
Lioncache 83a330b039 x86_64/VectorOps: Fix insertion bugs in VDupElement for 256-bit
Previously we weren't hitting this because we were never broadcasting
from the upper lane with VDupElement.
2023-08-19 01:21:12 -04:00
Lioncache bbed4d73ed OpcodeDispatcher: Improve VPERMQ/VPERMPD broadcast cases
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
2023-08-19 01:08:30 -04:00
Ryan Houdek a7e81ca731 InstCountCI: Update for push optimization
Some of these instructions improved a decent amount. Some even ending up
as being optimal.
2023-08-18 14:28:03 -07:00
Ryan Houdek 8b051b5e63 OpcodeDispatcher: Implement support for push IR operation
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.

While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
2023-08-18 14:19:16 -07:00
Ryan Houdek 1fdc4d2c62 IR: Implement support for a push operation
This is a bit of tricky operation where due to our our usage of SSA, the
incoming source isn't guaranteed to end its live-range at this
instruction.

This gives us a behaviour where to be optimal we need to take different
paths depending on if the incoming address register is the same as the
destination node.
Once we have form of RA constraints or non-SSA IR form that can
guarantee this restriction then this will go away.
2023-08-18 14:14:38 -07:00
Ryan Houdek ea5c67da80 SignalDelegator: Allow getting the internal configuration
Not used by FEX today but will be used by the WINE integration.
2023-08-18 11:56:52 -07:00
Ryan Houdek 0373826f46 Arm64Emitter: Stop saving and restoring platform register
FEX doesn't use the platform register on wine platforms so there is no
reason to save  and restore it.

On Linux we can still use it at some point but for now it isn't part of
our RA.
2023-08-18 11:49:44 -07:00
Ryan Houdek f912715df3 InstCountCI: STC and CLC is now optimal
Interestingly the segment register move instructions were previously
considered optimal on accident. They are actually optimal now which is
funny.
2023-08-18 11:41:21 -07:00
Ryan Houdek 7db2e487c3 IR/ConstProp: Remove some UBSAN behaviour
Changes the idiom used for constant mask generation to a ternary.
This pattern is definitely used elsewhere in code but we can get rid of
all instances here.
2023-08-18 11:41:11 -07:00
Ryan Houdek 6e5111b876 IR/ConstProp: Ensure that BFI with constant bitfields can optimize to Andn or Or
This optimizes the clc and stc instructions for flag setting and
clearing.
2023-08-18 11:33:11 -07:00
Ryan Houdek d502ad63f4 IR/ConstProp: Ensure ANDN is optimized 2023-08-18 11:27:06 -07:00
Ryan Houdek fc84f6b345 Merge pull request #2927 from bylaws/interrupt
FEXCore: Allow for interrupting the JIT on block entry
2023-08-18 06:14:24 -07:00
Billy Laws de63fd05d0 FEXCore: Allow for interrupting the JIT on block entry
This takes a similar approach to deferred signal handling and allows any given
thread to be interrupted while running JIT code by protecting the appropriate
page as RO. When the thread then enters a new block, it will try to acccess
that page and segfault. This is safer than just sending a signal to the thread
as that could stop in a place where JIT context couldn't be recovered correctly.
2023-08-18 05:58:51 -07:00
Ryan Houdek d3f0c7e969 Merge pull request #2925 from bylaws/winfile
Support for Config.json loading on WIN32
2023-08-18 05:00:22 -07:00
Ryan Houdek f1aa6208eb Merge pull request #2926 from bylaws/mingw
CMake: Add mingw toolchain file
2023-08-18 05:00:09 -07:00
Ryan Houdek b4d172649e Merge pull request #2924 from bylaws/logs
LogMan: Commonise log level to string conversion
2023-08-18 04:57:15 -07:00
Billy Laws 00556023c2 Remove unnecessary WIN32 file handling TODOs
With WOW, all allocations from 64-bit code use the full address space
and limiting is handled on the syscall thunk side so theres need to
worry about STL allocations stealing AS.
2023-08-18 04:37:40 -07:00
Billy Laws 5de0714766 FileLoading: Fix handling of non-existent files on WIN32 2023-08-18 04:37:40 -07:00
Billy Laws b862203491 Config: Add windows config loading support
This relies on wine's behaviour passing through linux paths and env vars,
so that the config in the user's home directory can be accessed outside
of the wine prefix.
2023-08-18 04:37:40 -07:00
Billy Laws bbfd15f801 LogMan: Commonise log level to string conversion 2023-08-18 04:36:31 -07:00
Billy Laws 0954c7eb9f AllocatorHooks: Add C++17 aligned new/delete functions 2023-08-18 04:32:16 -07:00
Billy Laws 8b2be809ff CI: Update to use mingw toolchain file 2023-08-18 04:31:41 -07:00
Billy Laws 193157812f CMake: Add mingw toolchain file 2023-08-18 04:31:41 -07:00
Ryan Houdek f09d9af3db Merge pull request #2922 from lioncash/psrld
OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
2023-08-17 17:05:35 -07:00
Ryan Houdek d19e2507e5 FEXCore: Adds telemetry around legacy segment register setting
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.

Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.

It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.

Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.

A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.

[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
   - 3.6 - Restricted Subset of Segmentation
      - `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
        registers; the base for CS, DS, ES, and SS is ignored for 32-bit
        mode, same as 64-bit mode (treated as zero).`
   - 4.2.17 - MOV to Segment Register
      - Will fault if SS is written (Breaking anything that writes to
        SS).
      - Will not fault if CS, DS, ES are written (Thus it sets the
        segment but gets ignored due to 3.6).
2023-08-17 17:00:41 -07:00
Lioncache 9e54ec2724 OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
2023-08-17 19:33:09 -04:00
Ryan Houdek 461ca6fe7c Merge pull request #2921 from lioncash/shift
OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
2023-08-17 16:03:16 -07:00
Lioncache 5a1f32c339 OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
PSLLIImpl already checks for and handles a shift value of zero.
2023-08-17 18:47:50 -04:00
Ryan Houdek a4a68b47ce Merge pull request #2920 from lioncash/ddup
OpcodeDispatcher: Improve VMOVDDUP output
2023-08-17 15:47:06 -07:00
Lioncache 01515cea2c OpcodeDispatcher: Improve VMOVDDUP output
We can make use of TRN1 here to collapse a bunch of these moves.
2023-08-17 18:30:44 -04:00
Ryan Houdek f70b6f37a2 Merge pull request #2919 from lioncash/sldup
OpcodeDispatcher: Improve output of {V}MOVSLDUP/{V}MOVSHDUP
2023-08-17 14:49:09 -07:00
Lioncache 764c844225 OpcodeDispatcher: Improve output of {V}MOVSHDUP 2023-08-17 17:01:50 -04:00
Lioncache 31719aac6a OpcodeDispatcher: Improve output of {V}MOVSLDUP 2023-08-17 17:01:46 -04:00
Ryan Houdek c9856daaee Merge pull request #2891 from alyssarosenzweig/move-fexcore
Move External/FEXCore/ to FEXCore/
2023-08-17 13:56:07 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00
Ryan Houdek 425b0347a9 Merge pull request #2918 from lioncash/quad
IR: Allow 128-bit broadcasts in VBroadcastFromMem
2023-08-17 11:32:04 -07:00
Lioncache e14e71aaff IR: Allow 128-bit broadcasts in VBroadcastFromMem
Now all vbroadcast implementations go down the more optimal path.

For non-SVE 128-bit cases where we only have 128-bit wide registers,
we behave like ld1rqb and just act as a normal 128-bit load for
interface convenience.
2023-08-17 14:20:12 -04:00
Ryan Houdek a9dea29f03 Merge pull request #2917 from lioncash/quad
ARMEmitter: Handle SVE load and broadcast quadword groups
2023-08-17 10:59:27 -07:00
Lioncache f97df2a40f ARMEmitter: Handle SVE load and broadcast quadword (scalar plus scalar) group 2023-08-17 13:32:20 -04:00
Lioncache 6f9bc1e2fe ARMEmitter: Handle SVE load and broadcast quadword (scalar plus imm) category 2023-08-17 13:32:17 -04:00
Ryan Houdek 49b8b7cd2c Merge pull request #2916 from Sonicadvance1/128bit_predicate
Arm64Emitter: Ensure that 128-bit predicate is generated with SVE
2023-08-17 09:56:56 -07:00
Ryan Houdek cfc6368064 Merge pull request #2914 from lioncash/broad
IR: Add VBroadcastFromMem opcode
2023-08-17 09:45:11 -07:00
Ryan Houdek 059c022255 Arm64Emitter: Ensure that 128-bit predicate is generated with SVE
In the case of running on a 128-bit SVE system this predicate wasn't
setup. Since we never had any predicate usage before this wasn't an
issue. Now that #2914 is using the 128-bit predicate we need to make
sure that we are generating it.
2023-08-17 09:37:55 -07:00
Lioncache 25708be807 IR: Add TSO handling to VBroadcastFromMem 2023-08-17 12:30:57 -04:00
Lioncache c8e3ca481f OpcodeDispatcher: Remove explicit zero-extending in VBROADCASTOp
Since the implementations zero the upper lanes when appropriate, we can
remove the unnecessary explicit move.
2023-08-17 12:17:24 -04:00
Lioncache 879bc5176e IR: Add VBroadcastFromMem opcode
Allows the implementations of the vbroadcast instructions to perform the
load and broadcast in one operation as opposed to doing the load and then
broadcast separately.

Notably, the broadcasting loads can also be used on systems that have SVE 128-bit
support as well, not only 256-bit.

On non-SVE systems, we use the equivalent AdvSIMD instructions.
2023-08-17 12:17:24 -04:00
Ryan Houdek 6d562f8b3b Merge pull request #2911 from Sonicadvance1/stop_abusing_orr
Arm64: Stop abusing orr in LoadConstant
2023-08-17 09:13:36 -07:00
Ryan Houdek 1029bb1fae Merge pull request #2910 from Sonicadvance1/minor_bfi_opt
Arm64: Optimize non-optimal BFI move case
2023-08-17 09:12:37 -07:00
Ryan Houdek 1343c14db0 Merge pull request #2909 from Sonicadvance1/optimize_clzero_clear
Arm64: Optimize CacheLine{Clear,Clean}
2023-08-17 09:10:45 -07:00
Ryan Houdek 6580796169 Merge pull request #2915 from lioncash/unused
OpcodeDispatcher: Remove unused variable in AVXVectorUnaryOpImpl
2023-08-17 08:44:18 -07:00
Lioncache 77c64285cb OpcodeDispatcher: Remove unused variable in AVXVectorUnaryOpImpl
Forgot to remove this when getting rid of the unnecessary
explicit zero-extending behavior
2023-08-17 11:26:41 -04:00
Ryan Houdek 9f3730f251 Merge pull request #2913 from Sonicadvance1/string_view_self
Filemanagement: Optimize GetSelf using a string_view
2023-08-17 06:32:51 -07:00
Ryan Houdek 2be8e22e14 Filemanagement: Optimize GetSelf using a string_view
Instead of allocating a temporary copy of the string, return a view of
it instead. Should improve the performance of system calls that take
file paths. Since it was allocating a string for every single syscall
that uses them in this case.
2023-08-16 21:39:05 -07:00
Ryan Houdek fe37c89109 FEXCore/Config: Stop making temporary string copies
For config values that were string objects we were unnecessary creating
copies each time the string was accessed.

Convert the () operator over to returning a reference.
2023-08-16 21:35:13 -07:00
Ryan Houdek d11563f48f InstCountCI: Update instructions for data movement
Instruction counts don't change at all here, just the instructions being
used.
2023-08-16 19:36:12 -07:00
Ryan Houdek 23fd79a3b3 Arm64: Stop abusing orr in LoadConstant
The current implementation uses orr excessively. This has FEX missing
hardware optimization opportunities where some CPU cores will zero-cycle
move constants that fit in to the 16-bits of movz/movk.

First evaluate up front if the number of 16-bit segments is > 1, in
those cases we should check if it is a bitfield that can be moved in one
instruction with orr.

After that point we will use movz for 16-bit constant moves.

Additionally this optimizes the case where a constant of zero is loaded
to be a `mov <reg>, zr` which gets renamed in most hardware.
2023-08-16 19:35:15 -07:00
Ryan Houdek 13a4009bd3 InstCountCI: Update changed operations due to bfi operation
No instruction count changes here, just moving from lsr to mov.
2023-08-16 14:37:36 -07:00
Ryan Houdek a3b40c37c2 Arm64: Optimize non-optimal BFI move case
Commonly we are doing a BFI into a 32-bit register, which is hitting the
ubfx (lsr alias) path.

In the case of 32-bit destination we can also do a regular move, which
will take advantage of CPU's rename functionality and give a minor speed
boost.
2023-08-16 14:35:41 -07:00
Ryan Houdek 6cb0f52e94 Merge pull request #2907 from Sonicadvance1/fix_dumpir_bug
FEXCore/IR: Fixes bug in IRDumper without specification
2023-08-16 14:24:33 -07:00
Ryan Houdek f283ba4b05 InstCountCI: clwb/clfush/clflushopt is now optimal
One or two instructions depending.
2023-08-16 14:20:22 -07:00
Ryan Houdek 4522a766e0 Arm64: Optimize CacheLine{Clear,Clean}
When the cacheline size matches the expected x86 cacheline size then we
can remove the spurious move + add.
2023-08-16 14:20:22 -07:00
Ryan Houdek dc0cf98a81 Merge pull request #2898 from Sonicadvance1/instcount_asm
InstCountCI: Support encoding expected Arm64 ASM in JSON
2023-08-16 13:51:43 -07:00
Ryan Houdek fc12958095 FEXCore/IR: Fixes bug in IRDumper without specification
Didn't notice this in the previous PR, When DUMPIR=stderr without and
selection of where to place it in PASSMANAGERDUMPIR it was supposed to
put the dumper at the end of the passes.

We need to make sure that it it placed at the end of the passes rather
than current `it`.
2023-08-16 13:51:03 -07:00
Ryan Houdek fd40e058e8 InstCountCI: Update tests with inline asm
This will result in a decent amount of data churn but it's all
automated so it is a non-issue.
2023-08-16 13:38:45 -07:00
Ryan Houdek 8706f895d0 InstCountCI: Sanitize out vixl address calculations 2023-08-16 13:38:45 -07:00
Ryan Houdek 9753ebdebc InstCountCI: Support encoding expected Arm64 ASM in JSON
This will allow investigating the Arm64 directly next to the test, plus
publicly linking directly to badly behaving tests.

Perfect for nerdsniping implementations.
2023-08-16 01:50:06 -07:00
Mai f95ef7092c Merge pull request #2905 from Sonicadvance1/instcountci_updatetests
InstCountCI: Update tests from actual ARM64 device
2023-08-16 04:13:05 -04:00
Ryan Houdek d4d5bc9635 InstCountCI: Update tests from actual ARM64 device
Previous numbers were from the simulator which were slightly different.
2023-08-15 14:39:37 -07:00
Mai df3d4efc80 Merge pull request #2904 from Sonicadvance1/instcountci_only_arm
GIthub: Only enable InstCountCI on an ARM platform
2023-08-15 17:33:10 -04:00
Mai dac220a6ff Merge pull request #2903 from Sonicadvance1/instcountci_rng
InstCountCI: Adds RNG support
2023-08-15 17:21:03 -04:00
Ryan Houdek 9ba7f2fd0e InstCountCI: Fixes instcountci_tests depends 2023-08-15 14:14:24 -07:00
Ryan Houdek 1441cb76b9 HostFeatures: Adds support for overriding ARMv8.1 LSE atomics
Always enable it on the InstCountCI.
2023-08-15 14:12:27 -07:00
Ryan Houdek 2d20513e34 GIthub: Only enable InstCountCI on an ARM platform
simulator generates some instruction count differences that we don't
care about. Just run this on ARM platforms only instead.
2023-08-15 14:12:27 -07:00
Mai da334fe7d5 Merge pull request #2901 from Sonicadvance1/instcountci_only_disassembler
InstCountCI: Disables tests with unsupported configurations
2023-08-15 16:01:47 -04:00
Ryan Houdek 04f1f073c8 InstCountCI: Adds RNG support
Some instructions in the SecondaryGroup need RNG support for testing.
2023-08-15 12:58:53 -07:00
Ryan Houdek 0e52158cef Merge pull request #2902 from lioncash/unary
OpcodeDispatcher: Eliminate unnecessary moves in AVXVectorUnaryOpImpl
2023-08-15 12:54:47 -07:00
Lioncache 17956eac5f OpcodeDispatcher: Eliminate unnecessary moves in AVXVectorUnaryOpImpl
We no longer need to do any manual zero-extending here, since this
will occur automatically on hardware with SVE when 128-bit AdvSIMD
is used.
2023-08-15 15:43:34 -04:00
Ryan Houdek 6ad053a1e6 Merge pull request #2900 from lioncash/sqrt
Arm64/VectorOps: Remove redundant move in VFRSqrt SVE path
2023-08-15 12:39:02 -07:00
Ryan Houdek b18e5e2f63 InstCountCI: Disables tests with unsupported configurations
Need to have the vixl disassembler enabled for instcountci.

Also need to make sure the host platform is using the ARM64 JIT.
2023-08-15 12:27:32 -07:00
Lioncache 2708374d95 Arm64/VectorOps: Remove redundant move in VFRSqrt SVE path
We can perform the SQRT first and then broadcast 1.0 into the destination
since all the intermediary work is done, meaning we don't have to worry
about Dst and Vector aliasing one another.
2023-08-15 15:22:21 -04:00
Mai 3c99fb84e2 Merge pull request #2899 from Sonicadvance1/instcountci_fix_asm_name
InstCountCI: Ensure output nasm name doesn't conflict
2023-08-15 14:44:44 -04:00
Ryan Houdek 710a3928ff Merge pull request #2897 from lioncash/broad
ARMEmitter: Handle SVE load and broadcast element group
2023-08-15 11:36:20 -07:00
Ryan Houdek ac1d2ec1d9 InstCountCI: Ensure output nasm name doesn't conflict
Pretty sure this is why CI is unhappy. If a test in a different file has
the same name then it is highly likely to conflict when nasm is
generating files and will overwrite and erase, causing CI to break.

Include the incoming json filename as part of the asm keys so it can't
conflict here.
2023-08-15 11:29:39 -07:00
Lioncache 6acce60855 ARMEmitter: Handle SVE load and broadcast element group
These can be used to improve vbroadcast implementations from
doing a mem load+dup in the non-GPR case into just directly
loading into the destination.
2023-08-15 13:47:12 -04:00
Mai 63f28eae4f Merge pull request #2896 from Sonicadvance1/primary_table
InstCountCI: Adds primary tables
2023-08-15 13:07:35 -04:00
Ryan Houdek eda67eb0a5 Merge pull request #2895 from lioncash/scalar
ARMEmitter: Handle load/store multiple structures (scalar plus scalar) groups
2023-08-15 09:12:40 -07:00
Ryan Houdek 21cac6ef0b InstCountCI: Adds primary tables
Surprisingly few instructions are optimal here.
This is all the instruction tables completed now!
2023-08-15 09:11:39 -07:00
Ryan Houdek 1193b55150 InstCountCI: Script auto change line 2023-08-15 09:11:29 -07:00
Ryan Houdek 27a53280bd InstCountCI: Fixes bitness in script 2023-08-15 09:11:07 -07:00
Mai 135b9ac425 Merge pull request #2894 from Sonicadvance1/instcount_secondary_prefixes
InstCountCI: Adds secondary prefix tables
2023-08-15 10:22:28 -04:00
Lioncache 81115f64f6 ARMEmitter: Handle SVE Store Multiple Structures (scalar plus scalar) 2023-08-15 10:18:37 -04:00
Lioncache 0176efa3bb ARMEmitter: Handle SVE Load Multiple Structures (scalar plus scalar) group 2023-08-15 10:01:15 -04:00
Ryan Houdek 2508274ddb InstCountCI: Adds secondary prefix tables
Most of these 128-bit vector ops are looking pretty good. Scalar and
edge case ops aren't always optimal though.
2023-08-15 06:57:47 -07:00
Mai 24d01cd8d2 Merge pull request #2893 from Sonicadvance1/instcount_secondary
InstCountCI: Adds secondary tables
2023-08-15 04:24:40 -04:00
Ryan Houdek 2fb72f822d InstCountCI: Adds secondary tables
A decent number of instructions that are optimal but still quite a lot
that aren't.
2023-08-14 16:04:31 -07:00
Ryan Houdek a2c2b042b2 CodeSizeValidation: Fixes nullptr dereference 2023-08-14 16:04:20 -07:00
Ryan Houdek 398e76be89 X86Tables: Fixes typo 2023-08-14 16:04:05 -07:00
Ryan Houdek 3885bc42f6 Merge pull request #2892 from Sonicadvance1/minor_config_changes
Config: Minor changes
2023-08-14 15:10:42 -07:00
Ryan Houdek c5439b294c TestHarnessRunner: InitializeConfig paths
Just removes an erroneous message in the TestHarnessRuner on each
execution.
2023-08-14 12:37:35 -07:00
Ryan Houdek f248e7f3e7 Config: If DumpIR is enabled, default enable a passmanager option
If DumpIR is enabled but the PassManagerDumpIR option isn't enabled then
this currently does nothing.

As a convenience, enable dumping the final optimized IR if an option
hasn't been specified.
2023-08-14 12:29:56 -07:00
Ryan Houdek e51606c669 Config: Fixes mixup in PassManagerDumpIR
The opt and pass options were inverted in PassManager.
Renames the enum to make this more clear.
2023-08-14 12:28:37 -07:00
Ryan Houdek 648d8aeb65 Config: Adds missing server option to DumpIR description
This was accepted but I failed to describe it when added.
2023-08-14 12:22:35 -07:00
Ryan Houdek 112c463655 Config: Ensure OutputLog to server doesn't try to expand path
"server" isn't a path, this was missed when it was added.
2023-08-14 12:20:58 -07:00
Ryan Houdek 2f0c690d54 Merge pull request #2890 from alyssarosenzweig/constprop/set-but-not
ConstProp: Fix set-but-not-used mask variable
2023-08-14 10:17:42 -07:00
Alyssa Rosenzweig 7ecbbd6c04 ConstProp: Fix set-but-not-used mask variable
I think this was the intended logic?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-14 11:59:59 -04:00
Mai 103e6044fa Merge pull request #2889 from Sonicadvance1/instcount_primarygroup
InstCountCI: Adds Primary group tables
2023-08-13 23:07:25 -04:00
Mai e334278f80 Merge pull request #2887 from Sonicadvance1/instcount_vex_map3
InstCountCI: Adds VEX map3 tables
2023-08-13 23:06:55 -04:00
Mai 7fe2d3b305 Merge pull request #2888 from Sonicadvance1/instcount_vex_map_group
InstCountCI: Adds VEX map group tables
2023-08-13 23:06:06 -04:00
Ryan Houdek 51d4bd5f05 InstCountCI: Adds Primary group tables
A lot of things fairly close to optimal but not quite there.
2023-08-13 17:26:09 -07:00
Ryan Houdek 2a871957f8 InstCountCI: Adds VEX map3 tables
Pretty much no instructions are optimal here. 64-bit RORX is the
outlier.
2023-08-13 13:00:00 -07:00
Ryan Houdek d9266167e1 InstCountCI: Adds VEX map group tables
No instructions are optimal in this table.
2023-08-13 12:59:11 -07:00
Ryan Houdek cbfea75037 Increase maximum instruction name to 128-bytes
Some long name instructions are going beyond 32-bytes. Obviously not
hurting to encode a few additional bytes.
2023-08-13 12:32:14 -07:00
Ryan Houdek ef3887ca4f X86Tables: Fixes typo in VEX table 2023-08-13 12:32:14 -07:00
Mai 871434dab9 Merge pull request #2886 from Sonicadvance1/instcount_vex_map2
InstCountCI: Adds FEX map2 tables
2023-08-13 07:42:01 -04:00
Ryan Houdek 33b84337ea InstCountCI: Adds FEX map2 tables
A handful of instructions are optimal, a few random ones are close, but
a lot aren't close.
2023-08-13 04:29:19 -07:00
Ryan Houdek 1e2b5ecd5d InstCountParser: Allow the ability to skip instructions 2023-08-13 04:29:07 -07:00
Mai a4647982cd Merge pull request #2885 from Sonicadvance1/vex_map1
InstCountCI: Adds VEX map1 tables
2023-08-13 07:05:52 -04:00
Ryan Houdek 047bbca5bb InstCountCI: Adds VEX map1 tables
None of the instructions here are optimal, a couple are close.

Will be splitting each of the map tables in to their own json files
since each one can get fairly large.
2023-08-13 02:15:53 -07:00
Mai 1146c42044 Merge pull request #2884 from Sonicadvance1/instcount_x87
InstCountCI: Adds x87 table
2023-08-12 11:12:01 -04:00
Ryan Houdek fc85431a0e InstCountCI: Adds x87 table
Most of these instructions aren't optimal.
Even outside of the ones that jump out of the JIT they have a issues in
most cases.
2023-08-12 06:53:30 -07:00
Mai a15934bd55 Merge pull request #2883 from Sonicadvance1/instcount_h0f38
InstCountCI: Adds H0F38 table
2023-08-12 09:05:27 -04:00
Ryan Houdek e1ef84f933 InstCountCI: Adds H0F38 table
Some of these instructions haven't been fully audited, which is why they
are classified as Unknown.
2023-08-12 05:20:44 -07:00
Ryan Houdek 139dd4ccbc Merge pull request #2882 from lioncash/adr
ARMEmitter: Handle SVE ADR
2023-08-11 20:52:42 -07:00
Lioncache 8fd810c3c0 ARMEmitter: Migrate adr off SVEMemOperand
We need to move the modifier enum out of the SVEMemOperand class
since it's also used with adr. Plus, this can also be convenient
not being tied down to the class itself.

This also makes accessing modifiers less noisy, since the class
2023-08-11 22:17:46 -04:00
Lioncache 068db933bf ARMEmitter: Handle SVE ADR 2023-08-11 19:50:21 -04:00
Ryan Houdek 72357e50a2 Merge pull request #2881 from lioncash/imm8
ARMEmitter: Handle SVE CPY (immediate)
2023-08-11 15:50:39 -07:00
Lioncache 0bf74a1f3e ARMEmitter: Use signed imm8 handler with dup_imm
Lets us deduplicate the behavior used for dup and cpy.
2023-08-11 18:30:45 -04:00
Lioncache 78f06c7fcb ARMEmitter: Handle SVE CPY (immediate)
Also adds the relevant aliases.
2023-08-11 18:24:26 -04:00
Ryan Houdek 6d1fcfce09 Merge pull request #2877 from Sonicadvance1/classification_adds
InstructionCountCI: Adds three more instruction tables
2023-08-11 15:09:29 -07:00
Ryan Houdek 5a0a6dd0ca Merge pull request #2880 from lioncash/vfp
ARMEmitter: Migrate off vixl float utils
2023-08-11 14:35:23 -07:00
Mai da17e24996 Merge pull request #2879 from Sonicadvance1/fix_adcx
FEXCore: Fixes bug with 32-bit adcx
2023-08-11 17:27:00 -04:00
Lioncache aef0795dc8 ARMEmitter: Make FloatToEquivalentUInt a little more robust
Rather than compare sizes, we should be comparing the types directly,
prevents any shenanigans from happening if interface changes occur to
Float16.
2023-08-11 17:16:14 -04:00
Lioncache f63498a558 ARMEmitter: Migrate off vixl float utils 2023-08-11 17:12:34 -04:00
Ryan Houdek 9aa3fde174 FEXCore: Fixes bug with 32-bit adcx
When a 32-bit adcx instruction was encountered, it was getting treated
as a 16-bit adcx instruction instead. This is because of the 0x66 prefix
required to handle this instruction.

Adds a unit test to ensure it doesn't break again.
2023-08-11 14:09:31 -07:00
Ryan Houdek 8fce13386a Merge pull request #2878 from lioncash/fcpy
ARMEmitter: Handle SVE FCPY (predicated)
2023-08-11 13:58:31 -07:00
Lioncache 247c7ce784 ARMEmitter: Handle SVE FCPY (predicated)
While we're at it, we can reduce our dependence on vixl's utils by
implementing our own based off the pseudocode of VFPExpandImm.
2023-08-11 16:18:20 -04:00
Ryan Houdek 7e12b056dc InstructionCountCI: Adds three more instruction tables
3DNow! table, H0F3A table, and SecondaryModRM table.

H0F3A table skips the SSE4.2 string instructions because they're
nightmares for now.
2023-08-11 11:13:42 -07:00
Ryan Houdek e8e52af2e8 CodeSizeValidation: Adds support for overriding CLZero support
Vixl simulator by default doesn't support this.
2023-08-11 11:12:46 -07:00
Ryan Houdek 6c7371af60 Merge pull request #2872 from Sonicadvance1/instruction_count_ci
FEX: Adds instruction count CI
2023-08-11 11:12:00 -07:00
Ryan Houdek acc7f2fa8f FEX: Adds instruction count CI
Implements CI for tracking instruction counts for generate blocks of
code when transforming from x86 to ARM64 assembly.

This will end up encompassing every instruction in our instruction
tables similarly to how our assembly tests try to test everything in our
instruction tables.

Incidentally, the data for this CI is generated using our assembly
tests. By enabling disassembly and instruction stats when executing a
suite of instructions, this gives the stats that can be added to a json
file.

The current implementation only implements the SecondGroup table of
instructions because it is a relatively small table and has known
inefficiencies in the instruction implementations. As this gets merged I
will be adding more tables of instructions to additional json files for
testing.

These JSON files will support adjusting CPU features regardless of the
host features so it can test implementations depending on different CPU
features. This will let us test things like one instruction having
different "optimal" implementations depending on if it supports SVE128,
SVE256, SVEI8MM, etc.

This initial instruction auditing is what found the bug in our vector
shift instructions by size of zero. If inspecting the result of the CI
run, you can tell that these instructions still aren't "optimal" because
they are doing loads and stores that can be eliminated.

The "Optimal" in the JSON is purely for human readable and grepping
ability to see what is optimal versus not. Same with the "Comment"
section.

According to my auditing spreadsheet, the total number of instructions
that will end up in these json files will be about 1000, but we will
likely end up with more since there will be edge cases that can be more
optimal depending on arguments.
2023-08-11 09:10:36 -07:00
Ryan Houdek b8b4dd8008 FEXCore/Utils: Add the ability to write a fextl::string 2023-08-11 09:08:54 -07:00
Ryan Houdek ec8855f8fb Arm64: Consolidate simulator and diassembler code in to Arm64Emitter
This was confusingly split between Arm64Emitter, Arm64Dispatcher, and
Arm64JIT.

- Arm64JIT objects were unnecessary and free to be deleted.
- Arm64Dispatcher simulator and decoder moved to Arm64Emitter
- Arm64Emitter disassembler and decoder renamed
  - Dropped usage of the PrintDisassembler since it is hardcoded to go
    through a FILE* type
  - We instead want its output to go through LogMan, which means using a
    split Decoder+Disassembler object pair.
  - Can't reuse the object from the vixl simulator since the simulator
    registers the decoder as a visitor, causing the simulator to execute
    while disassembling instructions if reused.
- Disassembly output for blocks and dispatcher now output through Logman
  - Blocks wrapped in Begin/End text for tracking purposes for CI.
2023-08-11 08:05:10 -07:00
Ryan Houdek 969ad9b3b0 Merge pull request #2869 from Sonicadvance1/sve_128bit_ci
Github: Adds a CI runner for 128-bit SVE testing
2023-08-11 07:37:47 -07:00
Ryan Houdek 5de7eeea20 Merge pull request #2876 from lioncash/comment
ARMEmitter: Remove resolved TODO comment
2023-08-11 07:22:19 -07:00
Ryan Houdek 0109e88082 Merge pull request #2875 from lioncash/ff
ARMEmitter: Handle contiguous first fault load (scalar plus scalar) group
2023-08-11 07:09:32 -07:00
Lioncache 73288f377f ARMEmitter: Remove resolved TODO comment
I forgot to remove this after implementing the normal gather instruction handling.
2023-08-11 10:03:10 -04:00
Lioncache 0aaa9503c9 ARMEmitter: Add missing ld1w (scalar plus scalar) tests
ld1sw was mistakenly tested twice.

Also groups the tests by data sizes.
2023-08-11 09:51:04 -04:00
Lioncache 14cc23b6c3 ARMEmitter: Handle contiguous first fault load (scalar plus scalar) group
Adds the only missing implementation category for the first-faulting loads,
making the interface more consistent.
2023-08-11 09:46:40 -04:00
Ryan Houdek 5404dba360 Github: Adds a CI runner for 128-bit SVE testing
We don't currently have a device in CI that can run SVE with 128-bit
width registers. Until we have a device with this, make sure the vixl
simulator is also running the ASM tests in this width.
2023-08-10 22:27:59 -07:00
Ryan Houdek 833c07e9e2 CoreState: Zero initialize some important members
This was causing test failure locally where some values were set to
uninitialized data. Ensure that gregs, YMM, and MMX registers are all
zero initialized.
2023-08-10 22:27:59 -07:00
Ryan Houdek 186ec201aa Config: Stop passing a temporary std::string_view outside of scope
Was causing strenum variables to be parsed, then leaving scope would
break the string.
2023-08-10 22:27:59 -07:00
Ryan Houdek 887c47c451 Config: Adds an option to override SVE width for CI 2023-08-10 21:25:57 -07:00
Ryan Houdek 0f3460e025 Config: Fixes typo in HostFeatures disable{sve,avx} 2023-08-10 21:25:57 -07:00
Ryan Houdek fadba9a3e1 External: Update vixl
Fixes simulator bug
2023-08-10 21:25:57 -07:00
Ryan Houdek 9d26af95ab Merge pull request #2873 from neobrain/refactor_warning_fixes
Various warning fixes
2023-08-10 16:58:14 -07:00
Tony Wasserka aed4dda3e4 Arm64: Remove unused function 2023-08-10 18:45:14 +02:00
Tony Wasserka e0d21e61cc Syscalls: Fix warnings about unused variables in Release builds 2023-08-10 18:45:14 +02:00
Tony Wasserka 45d0f0d349 ARMEmitter: Fix warnings about unused variables in Release builds 2023-08-10 18:45:14 +02:00
Tony Wasserka f1cc76614b Include VIXL as a system library
This suppresses warnings from VIXL headers.
2023-08-10 18:45:14 +02:00
Ryan Houdek 099f29f1ed Merge pull request #2871 from Sonicadvance1/fix_stats_missing_member
FEXCore: Fixes Arm64 stats disassembly
2023-08-10 06:23:02 -07:00
Ryan Houdek b1a3f82923 FEXCore: Fixes Arm64 stats disassembly
Requires the IR headerop to house the number of host instructions this
code is translating for the stats.

Fixes compiling with disassembly enabled, will be used with the
instruction count CI.
2023-08-10 03:23:25 -07:00
Ryan Houdek 6f4a23dd15 Merge pull request #2870 from lioncash/indexed
ARMEmitter: Handle SVE FP multiply-add long groups
2023-08-09 21:35:33 -07:00
Ryan Houdek f3182036bc Merge pull request #2867 from Sonicadvance1/dummy_thin_handlers
FEX: Create a CommonTools static library
2023-08-09 21:34:46 -07:00
Lioncache 444961ad79 ARMEmitter: Handle SVE FP multiply-add long group 2023-08-09 15:20:04 -04:00
Lioncache 48a3271fbc ARMEmitter: Handle SVE FP multiply-add long (indexed) group 2023-08-09 15:19:49 -04:00
Mai ea8fbc61c2 Merge pull request #2868 from Sonicadvance1/irdumper_passmanager
IR: Adds Option to run the IRDumper with more configurations
2023-08-09 10:28:25 -04:00
Ryan Houdek 35e97ec9bc IR: Adds Option to run the IRDumper with more configurations
This is incredibly useful and I find myself hacking this feature in
every time I am optimizing IR. Adds a new configuration option which
allows dumping IR at various times.

Before any optimization passes has happened
After all optimizations passes have happened
Before and After each IRPass to see what is breaking something.

Needs #2864 merged first
2023-08-09 05:58:20 -07:00
Ryan Houdek 53ac8abce9 Merge pull request #2863 from Sonicadvance1/stats
Arm64: Adds stats to the disassembly
2023-08-09 04:06:22 -07:00
Ryan Houdek fe351353f6 Merge pull request #2865 from Sonicadvance1/first_sve_opt
Arm64: Implement first SVE-128bit optimization
2023-08-09 04:06:05 -07:00
Ryan Houdek a23cb0447b Arm64: Implement first SVE-128bit optimization
This is a /very/ simple optimization purely because of a choice that ARM
made with SVE in latest Cortex.

Cortex-A715:
   - sxtl/sxtl2/uxtl/uxtl2 can execute 1 instruction per cycle.
   - sunpklo/sunpkhi/uunpklo/uunpkhi can execute 2 instructions per cycle.

Cortex-X3:
   - sxtl/sxtl2/uxtl/uxtl2 can execute 2 instruction per cycle.
   - sunpklo/sunpkhi/uunpklo/uunpkhi can execute 4 instructions per cycle.

This is fairly quirky since this optimization only works on SVE systems
with 128-bit Vector length. Which since it is all of the current
consumer platforms, it will work.
2023-08-09 03:51:57 -07:00
Ryan Houdek f2aa2ce4bb Arm64: Rename HostSupportsSVE
We need to know the difference between the host supporting SVE with
128-bit registers versus 256-bit registers. Ensure we know the
difference.

No functional change here.
2023-08-09 03:51:56 -07:00
Ryan Houdek cf93652708 Config: Adds support for overriding host features
This allows use to both enable and disable regardless of what the host
supports. This replaces the old `EnableAVX` option.

Unlike the old EnableAVX option which was a binary option which could
only disable, each of these options are technically trinary states.
Not setting an option gives you the default detection, while explicitly
enabling or disabling will toggle the option regardless of what the host
supports.

This will be used by the instruction count CI in the future.
2023-08-09 03:51:37 -07:00
Ryan Houdek eaed5c4704 Merge pull request #2862 from Sonicadvance1/optimize_vector_zero
ARM64: Optimize vector zeroing
2023-08-09 03:51:04 -07:00
Mai c77ed78f5a Merge pull request #2861 from Sonicadvance1/fix_vector_shift_by_zero
FEXCore: Fixes vector shifts by zero
2023-08-09 05:52:10 -04:00
Ryan Houdek 348844a95b FEX: Create a CommonTools static library
Moves the dummy handlers over to this library. This will end up getting
used for more than the mingw test harness runner once the instruction
count CI is operational.
2023-08-09 02:27:13 -07:00
Ryan Houdek e8fb322025 unittests: Adds tests for vector shifts with zero immediate
To ensure FEX doesn't encounter the encoding bug again.
2023-08-09 02:16:17 -07:00
Ryan Houdek 5f0efda8fe ARM64: Fixes shift by immediate zero
These would emit invalid instructions in most cases. Turn in to a move
or a no-op if the shift is zero.
2023-08-09 02:16:17 -07:00
Ryan Houdek d198d701aa OpcodeDispatcher: Fixes vector shifts by immediate zero
pslldq logic was wrong in the case of zero shift.
The rest should just return their source in the case of zero shift.
2023-08-09 02:16:17 -07:00
Mai c4c7620ed5 Merge pull request #2866 from Sonicadvance1/remove_unnecessary_loadconstant
Arm64: Remove erroneous LoadConstant
2023-08-09 05:10:11 -04:00
Ryan Houdek bf5719770e Arm64: Remove erroneous LoadConstant
This was a debug LoadConstant that would load the entry in to a temprary
register to make it easier to see what RIP a block was in.

This was implemented when FEX stopped storing the RIP in the CPU state
for every block. This is now no longer necessary since FEX stores the
in the tail data of the block.

This was affecting instructioncountci when in a debug build.
2023-08-08 22:56:36 -07:00
Ryan Houdek 0f6a268243 Arm64: Adds stats to the disassembly
I use this locally when looking for optimization opportunities in the
JIT.
The instruction count CI in the future will use this as well.
Just get it upstreamed right away.
2023-08-08 22:28:52 -07:00
Ryan Houdek e0461497a0 ARM64: Optimize vector zeroing
`eor <reg>, <reg>, <reg>` is not the optimal way to zero a vector
register on ARM CPUs. Instead we should move by constant or zero
register to take advantage of zero-latency moves.
2023-08-08 22:24:11 -07:00
Ryan Houdek e9e5b6fb0b Docs: Update for release FEX-2308 2023-08-06 02:34:55 -07:00
Ryan Houdek 68cb6e61d1 Merge pull request #2860 from lioncash/alias
ARMEmitter: Add missing atomic aliases
2023-08-06 02:05:30 -07:00
Mai 5d0b2060e2 Merge pull request #2858 from Sonicadvance1/fix_clzero
X86Tables: Fixes CLZero destination address
2023-08-06 05:05:07 -04:00
Lioncache b7d05a65c7 ARMEmitter: Add missing atomic aliases 2023-08-04 21:49:59 -04:00
Lioncache 93fe2fe06c ARMEmitter: Detemplatize LoadStoreAtomicLSE
Lets us lessen some template instantiations.
2023-08-04 19:47:03 -04:00
Ryan Houdek 21eb6e03c7 Merge pull request #2859 from bylaws/ooo
Fix 16-bit popa insertion behaviour
2023-08-04 16:37:09 -07:00
Lioncache 1b53337925 ARMEmitter: Simplify LoadStoreAtomicLSE variant
We can move the base opcode into the implementation function.
2023-08-04 19:35:54 -04:00
Billy Laws 52e5b8ccd9 OpcodeDispatcher: Fix 16-bit popa insertion behaviour
The 16-bit writes shouldn't overwrite the upper half of the 32-bit register for
POPA.
2023-08-04 17:24:55 +01:00
Billy Laws 8c8a8c84df unittests: Test for 16-bit popa insertion behaviour 2023-08-04 17:24:52 +01:00
Ryan Houdek 0d6837f1a1 Merge pull request #2856 from Sonicadvance1/allow_override_linker
CMake: Allow overriding linker
2023-08-04 03:37:17 -07:00
Ryan Houdek 7ef3cb88f9 CMake: Allow overriding linker
While the ENABLE_LLD and ENABLE_MOLD options are nice, they don't handle
the case when the linker of `lld` or `mold` doesn't match the compiler.

This particularly crops up when overriding the C compiler to a new
version of clang but the globally installed `ld.lld` is still the old
clang version.
This then causes clang to fail with unusual errors when upstream breaks
compatibility with itself.

Easy enough to use by passing the linker to cmake:
`-DUSE_LINKER=/usr/bin/ld.lld-15`

This also removes the ENABLE_LLD and ENABLE_MOLD options to use
USE_LINKER directly.
- ldd: `-DUSE_LINKER=lld`
- mold: `-DUSE_LINKER=mold`

Example of compiler failure when built with clang-15 but attempting to
link with ld.lld 14:
```bash
ld.lld-14: error: unittests/APITests/CMakeFiles/Filesystem.dir/Filesystem.cpp.o: Opaque pointers are only supported in -opaque-pointers mode (Producer: 'LLVM15.0.7' Reader: 'LLVM 14.0.6')
```
2023-08-04 02:34:15 -07:00
Ryan Houdek 617977357a X86Tables: Fixes CLZero destination address
This needs to default to 64-bit addresses, this was previously
defaulting to 32-bit which was meaning the destination address was
getting truncated. In a 32-bit process the address is still 32-bit.

I'm actually surprised this hasn't caused spurious SIGSEGV before this
point.

Adds a 32-bit test to ensure that side is tested as well.
2023-08-04 02:31:30 -07:00
Ryan Houdek 5a53c9231b Merge pull request #2855 from alyssarosenzweig/tst-instead-of-cmn
JIT: Use TST instead of CMN
2023-08-02 15:07:28 -07:00
Alyssa Rosenzweig a996e5300e JIT: Use TST instead of CMN
This is more obvious. llvm-mca says TST is half the cycle count of CMN
for whatever it's defaulting to. dougallj's reference shows both as the
same performance.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 17:51:33 -04:00
Ryan Houdek 25b2af14fd Merge pull request #2854 from alyssarosenzweig/flags/rotate-harder
OpcodeDispatcher: Optimize rotates
2023-08-02 14:09:40 -07:00
Alyssa Rosenzweig 76059949ea Merge pull request #2852 from Sonicadvance1/optimize_phsubsw
OpcodeDispatcher: Optimize phsubsw/phaddsw
2023-08-02 17:02:13 -04:00
Ryan Houdek 6e15c9c213 Merge pull request #2851 from Sonicadvance1/optimize_cas128_select
OpcodeDispatcher: Optimize CMPXCHG{8B,16B} final comparison
2023-08-02 14:01:58 -07:00
Alyssa Rosenzweig 01fcca884b OpcodeDispatcher: Optimize rotates
In the non-immediate cases, we can amortize some work between the two
flags to come out 1 instruction ahead.

In the immediate case, costs us an extra 2 instructions compared to
before we packed NZCV flags, but this mitigates a bigger instr count
regression that this PR would otherwise have. Coming out ahead will
require FlagM and smarter RA, but is doable.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 16:54:56 -04:00
Ryan Houdek 91bd3aa62a Merge pull request #2832 from alyssarosenzweig/flags/pack-nzcv
Pack NZCV flags
2023-08-02 13:42:56 -07:00
Alyssa Rosenzweig 7a0119b092 OpcodeDispatcher: Optimize right shifts
Same technique as the left shifts. Gets rid of all our COND_FLAG_SET
use, which is good because it's a performance footgun.

Overall saves 17 instructions (!!!!) from the flag calculation code for
`sar eax, cl`.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 14:45:44 -04:00
Alyssa Rosenzweig fa42c1616e OpcodeDispatcher: Preserve AF for non-immediate shift
The selection logic is expensive. Saves 5 instructions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 14:45:10 -04:00
Alyssa Rosenzweig 18783948f7 OpcodeDispatcher: Optimize non-immediate shift
Similar to the immediate case, but now we select between the entire old
and new NZCV registers. This is faster than selecting each bit
independently. Saves 11 instructions for calculating flags for "shl eax,
cl".
2023-08-02 14:45:10 -04:00
Alyssa Rosenzweig 969d2e4b6a OpcodeDispatcher: Zero OF for shift > 1
It is undefined in this case. We prefer to zero (rather than preserve
the existing value) as it avoids a costly RMW of the NZCV register.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 14:14:08 -04:00
Alyssa Rosenzweig 7ceaf56407 OpcodeDispatcher: Use SetNZ_ZeroCV for immediate shifts
We need to be careful to preserve V if needed. For `shl 1` and `shr 1`,
saves 2 instruction overall compared to before the PR. For `sar 1`,
saves 3 overall.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 14:12:00 -04:00
Alyssa Rosenzweig 7c52375267 OpcodeDispatcher: Optimize (U)MUL flags calculation
csel the value we want directly. Saves 2 instructions. Could do better
still but hey, progress is progress. Currently looks like:

2995: 0x0000ffff6c600040  320407f5		orr w21, wzr, #0x30000000
2995: 0x0000ffff6c600044  f10000df		cmp x6, #0x0 (0)
2995: 0x0000ffff6c600048  9a950294		csel x20, x20, x21, eq
2995: 0x0000ffff6c60004c  b902db94		str w20, [x28, #728]

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:40:44 -04:00
Alyssa Rosenzweig a4ea792d03 OpcodeDispatcher: Optimize GetRFLAG of definitely-0 flag
We can just return zero, no need to do a pointless Bfe. Saves yet
another instruction for GetPackedRLAG in a test I'm looking at.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig cec98637c9 OpcodeDispatcher: Avoid extra OR in GetPackedRFLAG
We know that bit 0 is CF, so we can do CF first and then avoid setting
Original to 2 (for reserved) with a silly `or xzr, #2` instruction.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig d5026f5815 OpcodeDispatcher: Handle SF/ZF together for GetPackedRFLAG
They're together on both x86 and arm64, so this is faster if we're
getting both.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig 1e4456ec40 OpcodeDispatcher: Use orlshl in GetPackedRFLAG
Saves some moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig e285c7c9a0 OpcodeDispatcher: Use TEST when possible
Faster sign/negate testing for 32-bit/64-bit inputs. This could maybe be
extended to 8/16-bit if we have FlagM but that's for later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig 8244c7f2a6 OpcodeDispatcher: Use orlshl when possible
If we can prove that a flag bit could not possibly be set, we can use
orlshl rather than bfi, which can be more efficient.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig 747c5e17f8 OpcodeDispatcher: Add ZeroNZCV helper
In some cases we just want to insert in one bit at a time, add a helper
to zero the 4 flags together so we can avoid the extra RMW cycle at the
beginning.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 13:18:39 -04:00
Alyssa Rosenzweig 52ce027e9b OpcodeDispatcher: Set N flag more efficiently
We can set N more efficiently with some bit math, and zero ZCV at the
same time. In the future we'll be able to use TST for this to make it
even faster.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 0fb3903889 OpcodeDispatcher: Pack NZCV flags together
Later, this will let us take advantage of the arm64 flags. For
now, this just turns some strb's into bfi's for dubious benefit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 2832afd0f7 OpcodeDispatcher: Calculate deferred flags more
If we read or write NZCV flags we need to call CalculateDeferredFlags on
block boundaries, if only to flush out the cached copy.

Also, when leaving a block we call it to flush out. This is annoyingly
invasive but I don't know of a better way to do this that doesn't
involve rearchitecting the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 4af42477a7 OpcodeDispatcher: Use GetRFLAG more
We'll add an extra caching layer in a moment so can't call _LoadFlag
directly and expect correct results.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 4786aa479c OpcodeDispatcher: Unify SetRFLAG impls
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 8a049aa0c3 Context: Make BackendFeatures public
So that the OpcodeDispatcher can check the supported features and emit
code accordingly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 64bac687b5 IR: Add TEST opcode
Maps to arm64 tst, except properly SSA. This will need some RA support
to avoid redundant mrs/msr sequences.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 6e05494ad0 IR: Add Orlshr
Similar to Orlshl. This will let us save an instruction in
GetPackedRFLAG.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 4d13f5d97d IR: Add Orlshl op
On arm64, orr (with a shifted register) is maybe fewer cycles and
definitely easier on the RA than bfi (=> fewer moves generated). So,
it's preferred when we know the corresponding bit is 0 in the
destination.

It's not useful on other targets, so it's gated behind a backend feature
bit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Alyssa Rosenzweig 1ad928dc2c IR: Add special NZCV "flag"
Reserve 4 bytes of "flags" to model the 32-bit arm64 NZCV register, so
we can start porting FEX's flag handling code over to using NZCV without
needing the whole compiler to be aware of instructions that might
clobber host flags.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-02 12:38:00 -04:00
Ryan Houdek c235ab883d OpcodeDispatcher: Optimize phsubsw/phaddsw
Goes from 41 instructions down to 20(!) instructions.
The primary optimization here is removing a bunch of dull inserts and
instead using zip logic to get the elements where we want them.

The previous implementation was trying to retain semantics around how
the original instruction is implemented, which makes no sense at all.
Unzip the even and odd elements and just do the saturating operations
directly!

Huge shoutout to @dougallj again, showing that I was thinking about this implementation far too much like an x86 developer.

Before:
```asm
0x0000ffff8f300c50  10ffffe0    adr x0, #-0x4 (addr 0xffff8f300c4c)
0x0000ffff8f300c54  f9005f80    str x0, [x28, #184]
0x0000ffff8f300c58  3dc010c4    ldr q4, [x6, #64]
0x0000ffff8f300c5c  6e60bae5    neg v5.8h, v23.8h
0x0000ffff8f300c60  6e60b886    neg v6.8h, v4.8h
0x0000ffff8f300c64  4eb71ee0    mov v0.16b, v23.16b
0x0000ffff8f300c68  6e0614a0    mov v0.h[1], v5.h[1]
0x0000ffff8f300c6c  4ea01c07    mov v7.16b, v0.16b
0x0000ffff8f300c70  6e0614c4    mov v4.h[1], v6.h[1]
0x0000ffff8f300c74  6e0e34a7    mov v7.h[3], v5.h[3]
0x0000ffff8f300c78  6e0e34c4    mov v4.h[3], v6.h[3]
0x0000ffff8f300c7c  6e1654a7    mov v7.h[5], v5.h[5]
0x0000ffff8f300c80  6e1654c4    mov v4.h[5], v6.h[5]
0x0000ffff8f300c84  4ea71ce0    mov v0.16b, v7.16b
0x0000ffff8f300c88  6e1e74a0    mov v0.h[7], v5.h[7]
0x0000ffff8f300c8c  4ea01c05    mov v5.16b, v0.16b
0x0000ffff8f300c90  6e1e74c4    mov v4.h[7], v6.h[7]
0x0000ffff8f300c94  0f10a4a6    sxtl v6.4s, v5.4h
0x0000ffff8f300c98  4f10a4a5    sxtl2 v5.4s, v5.8h
0x0000ffff8f300c9c  0f10a487    sxtl v7.4s, v4.4h
0x0000ffff8f300ca0  4f10a484    sxtl2 v4.4s, v4.8h
0x0000ffff8f300ca4  4ea5bcc5    addp v5.4s, v6.4s, v5.4s
0x0000ffff8f300ca8  4ea4bce4    addp v4.4s, v7.4s, v4.4s
0x0000ffff8f300cac  0e6148a5    sqxtn v5.4h, v5.4s
0x0000ffff8f300cb0  4ea51ca0    mov v0.16b, v5.16b
0x0000ffff8f300cb4  4e614880    sqxtn2 v0.8h, v4.4s
0x0000ffff8f300cb8  4ea01c17    mov v23.16b, v0.16b
0x0000ffff8f300cbc  58000040    ldr x0, pc+8 (addr 0xffff8f300cc4)
0x0000ffff8f300cc0  d63f0000    blr x0
0x0000ffff8f300cc4  a4e97128    ldff1h {z8.d}, p4/z, [x9, x9, lsl #1]
0x0000ffff8f300cc8  0000ffff    udf #0xffff
0x0000ffff8f300ccc  000100dd    unallocated (Unallocated)
0x0000ffff8f300cd0  00000000    udf #0x0
[DEBUG] RIP: 0x100d7
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 41
[DEBUG] Blow-up Amt: 41x
```

After:
```asm
0x0000ffffe2500a04  10ffffe0            adr x0, #-0x4 (addr 0xffffe2500a00)
0x0000ffffe2500a08  f9005f80            str x0, [x28, #184]
0x0000ffffe2500a0c  3dc010c4            ldr q4, [x6, #64]
0x0000ffffe2500a10  4e441ae5            uzp1 v5.8h, v23.8h, v4.8h
0x0000ffffe2500a14  4e445ae4            uzp2 v4.8h, v23.8h, v4.8h
0x0000ffffe2500a18  4e642cb7            sqsub v23.8h, v5.8h, v4.8h
0x0000ffffe2500a1c  58000040            ldr x0, pc+8 (addr 0xffffe2500a24)
0x0000ffffe2500a20  d63f0000            blr x0
0x0000ffffe2500a24  f7fec128            unallocated (Unallocated)
0x0000ffffe2500a28  0000ffff            udf #0xffff
0x0000ffffe2500a2c  000100dd            unallocated (Unallocated)
0x0000ffffe2500a30  00000000            udf #0x0
[DEBUG] RIP: 0x100d7
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 20
[DEBUG] Blow-up Amt: 20x
```
2023-08-02 00:10:32 -07:00
Ryan Houdek 4bf3a0888b OpcodeDispatcher: Optimize CMPXCHG{8B,16B} final comparison
Optimizes the instruction blow-up from 36x to 34x.
The issue with this instruction is that AArch64 doesn't do something
like x86 where it sets a flag if the CAS was successful. This means we
need to do additional comparisons after the fact to see if it was
actually successful.

Previously this was implemented as eor+eor+orr+cmp+cset, Now it is
cmp+ccmp+cset, Saving two instructions.
Simple optimization but easy to do. This instruction is still mostly
killed by the overhead of moving registers all over.

Before:
```asm
0x0000ffffe25002ec  10ffffe0            adr x0, #-0x4 (addr 0xffffe25002e8)
0x0000ffffe25002f0  f9005f80            str x0, [x28, #184]
0x0000ffffe25002f4  aa0403f4            mov x20, x4
0x0000ffffe25002f8  aa0603f5            mov x21, x6
0x0000ffffe25002fc  aa0703f6            mov x22, x7
0x0000ffffe2500300  aa0503f7            mov x23, x5
0x0000ffffe2500304  aa1403e2            mov x2, x20
0x0000ffffe2500308  aa1503e3            mov x3, x21
0x0000ffffe250030c  4862ffb6            caspal x2, x3, x22, x23, [x29]
0x0000ffffe2500310  aa0203f4            mov x20, x2
0x0000ffffe2500314  aa0303f5            mov x21, x3
0x0000ffffe2500318  aa1403f6            mov x22, x20
0x0000ffffe250031c  aa1503f4            mov x20, x21
; ZF Flag + branch
0x0000ffffe2500320  ca0402d5            eor x21, x22, x4
0x0000ffffe2500324  ca060297            eor x23, x20, x6
0x0000ffffe2500328  aa1702b5            orr x21, x21, x23
0x0000ffffe250032c  f10002bf            cmp x21, #0x0 (0)
0x0000ffffe2500330  9a9f17f7            cset x23, eq
0x0000ffffe2500334  390b1b97            strb w23, [x28, #710]
0x0000ffffe2500338  b4000075            cbz x21, #+0xc (addr 0xffffe2500344)

0x0000ffffe250033c  aa1603e4            mov x4, x22
0x0000ffffe2500340  aa1403e6            mov x6, x20
0x0000ffffe2500344  58000040            ldr x0, pc+8 (addr 0xffffe250034c)
0x0000ffffe2500348  d63f0000            blr x0
0x0000ffffe250034c  f7fec128            unallocated (Unallocated)
0x0000ffffe2500350  0000ffff            udf #0xffff
0x0000ffffe2500354  00010053            unallocated (Unallocated)
0x0000ffffe2500358  00000000            udf #0x0
[DEBUG] RIP: 0x1004f
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 36
[DEBUG] Blow-up Amt: 36x
```

After:
```asm
0x0000ffffe25002ec  10ffffe0            adr x0, #-0x4 (addr 0xffffe25002e8)
0x0000ffffe25002f0  f9005f80            str x0, [x28, #184]
0x0000ffffe25002f4  aa0403f4            mov x20, x4
0x0000ffffe25002f8  aa0603f5            mov x21, x6
0x0000ffffe25002fc  aa0703f6            mov x22, x7
0x0000ffffe2500300  aa0503f7            mov x23, x5
0x0000ffffe2500304  aa1403e2            mov x2, x20
0x0000ffffe2500308  aa1503e3            mov x3, x21
0x0000ffffe250030c  4862ffb6            caspal x2, x3, x22, x23, [x29]
0x0000ffffe2500310  aa0203f6            mov x22, x2
0x0000ffffe2500314  aa0303f7            mov x23, x3
0x0000ffffe2500318  aa1603f8            mov x24, x22
0x0000ffffe250031c  aa1703f9            mov x25, x23
; ZF Flag + branch
0x0000ffffe2500320  eb1402df            cmp x22, x20
0x0000ffffe2500324  fa5502e0            ccmp x23, x21, #nzcv, eq
0x0000ffffe2500328  9a9f17f4            cset x20, eq
0x0000ffffe250032c  390b1b94            strb w20, [x28, #710]
0x0000ffffe2500330  b5000074            cbnz x20, #+0xc (addr 0xffffe250033c)

0x0000ffffe2500334  aa1803e4            mov x4, x24
0x0000ffffe2500338  aa1903e6            mov x6, x25
0x0000ffffe250033c  58000040            ldr x0, pc+8 (addr 0xffffe2500344)
0x0000ffffe2500340  d63f0000            blr x0
0x0000ffffe2500344  f7fec128            unallocated (Unallocated)
0x0000ffffe2500348  0000ffff            udf #0xffff
0x0000ffffe250034c  00010053            unallocated (Unallocated)
0x0000ffffe2500350  00000000            udf #0x0
[DEBUG] RIP: 0x1004f
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 34
[DEBUG] Blow-up Amt: 34x
```
2023-08-01 20:01:02 -07:00
Ryan Houdek 6834fe32e4 IR: Adds support for GPRPair to Select IR op
This IR operation was limited to GPR only previously for the values
getting compared.

This adds support for GPRPair (and technically FPR) so that it can be
used directly with GPR pairs. I say technically FPR because the
IREmitter disallowed FPRs for the comparison, but this was already
supported in all of the backends, we just didn't ever use it.

Some minor changes to the constant prop pass to ensure that we don't try
to propagate a select in to a CondJump, otherwise pair comparisons would
be duplicated. This code is expecting to be able to merge a simple
comparison in to a `cbnz`, which doesn't happen with GPR pairs.
2023-08-01 19:50:15 -07:00
Ryan Houdek d196709162 Merge pull request #2850 from Sonicadvance1/put_the_man_in_its_place
FEXCore: Ensure that the man page follows DESTDIR
2023-08-01 14:28:06 -07:00
Ryan Houdek 475a6a38fa FEXCore: Ensure that the man page follows DESTDIR
When cmake's `install` function is invoked with a relative path, then it
is interpreted as being relative to the `CMAKE_INSTALL_PREFIX` variable.

This variable follows both `DESTDIR` and `CMAKE_INSTALL_PREFIX` so it is
best to use relative addresses in the install path.

Thanks for the report Mike!
Fixes #2849
2023-08-01 13:28:21 -07:00
Mai 0f748bd724 Merge pull request #2829 from Sonicadvance1/mingw_runner
Github: Adds mingw build test workflow
2023-07-31 21:45:24 -04:00
Ryan Houdek 47bd331239 Github: Adds mingw build test workflow
Currently only does a build, doing a CI run means figuring out why
TestHarnessRunner doesn't find libraries correctly.

I want to ensure we don't break building at least while sorting out the
rest of this.
2023-07-31 17:55:24 -07:00
Ryan Houdek 173b70d191 Merge pull request #2817 from Sonicadvance1/psad_you_know
OpcodeDispatcher: Optimize PSAD* to use vuabdl{2,}
2023-07-31 14:54:57 -07:00
Ryan Houdek 96aa0a844e Merge pull request #2845 from Sonicadvance1/thunks_guestlibs
Thunks: Set bitness flags for 64-bit guests
2023-07-31 12:43:11 -07:00
Ryan Houdek 48121cd585 Thunks: Set bitness flags for 64-bit guests
This will help non-multiarch aware distros
2023-07-31 12:22:33 -07:00
Ryan Houdek 52d7efda10 Merge pull request #2839 from Sonicadvance1/shrdi_of
OpcodeDispatcher: Fixes SHRD by immediate OF flag calculation
2023-07-31 10:56:40 -07:00
Ryan Houdek 30ab4d3f58 Merge pull request #2848 from lioncash/aliases
ARMEmitter: Add cinc/cinv/csetm aliases
2023-07-30 20:18:29 -07:00
Lioncache f9fa0f25e2 ARMEmitter: Add cinc/cinv/csetm aliases
Also corrects a cneg test using CC_AL. The ARM ARM says this
shouldn't be using both AL and NV
2023-07-30 22:41:03 -04:00
Ryan Houdek eebcbfda96 Merge pull request #2847 from lioncash/aliases
ARMEmitter: Add ngc/ngcs aliases
2023-07-30 19:07:32 -07:00
Lioncache 1483ddb538 ARMEmitter: Add ngc/ngcs aliases 2023-07-30 21:55:54 -04:00
Ryan Houdek d2bca9b997 Merge pull request #2846 from lioncash/aliases
ARMEmitter: Add bfc/bfxil aliases
2023-07-30 18:49:01 -07:00
Lioncache 8f48021a63 ARMEmitter: Add bfxil alias 2023-07-30 21:18:20 -04:00
Lioncache 1da5d7f2e6 ARMEmitter: Add bfc alias 2023-07-30 21:18:17 -04:00
Ryan Houdek 0a5db6c404 Merge pull request #2843 from Sonicadvance1/implement_socketcall_recvsendmmsg
Linux: Implement {recv,send}mmsg inside of socketcall
2023-07-30 18:04:44 -07:00
Ryan Houdek 4681061011 Merge pull request #2844 from lioncash/aliases
ARMEmitter: Add sbfiz/tst/ubfiz aliases
2023-07-30 18:01:34 -07:00
Lioncache 4e9eeb1a16 ARMEmitter: Add SBFIZ/UBFIZ aliases 2023-07-30 20:19:37 -04:00
Lioncache 95d728b634 ARMEmitter: Add TST aliases 2023-07-30 19:36:56 -04:00
Ryan Houdek 187e551a0c Linux: Implement {recv,send}mmsg inside of socketcall
These are the last two missing socketcall operations. Implementing them
is mostly using the other implementation just in a helper function.
2023-07-30 15:45:01 -07:00
Ryan Houdek 7c4e4c4409 Merge pull request #2842 from Sonicadvance1/syscall_accept4
Linux: Implement accept4 inside of socketcall
2023-07-30 15:44:18 -07:00
Ryan Houdek bdd8df2100 Linux: Implement accept4 inside of socketcall
Steam started using this recently. So implement it.
Still missing recvmmsg and sendmmsg here.
2023-07-30 15:11:45 -07:00
Ryan Houdek 1d1bdfb96d OpcodeDispatcher: Fixes SHRD by immediate OF flag calculation
We were calculating this like the regular SHR instruction which isn't
correct.

With this resolved, Denuvo games get slightly farther.
2023-07-28 18:00:50 -07:00
Ryan Houdek 2fc6542d15 Merge pull request #2838 from lioncash/imm
ARMEmitter: Finish off remaining SVE Integer Wide Immediate - Unpredicated categories
2023-07-28 17:57:23 -07:00
Lioncache 5e88be3c99 ARMEmitter: Handle SVE MUL immediate (unpredicated) 2023-07-28 19:55:06 -04:00
Lioncache be52844ce0 ARMEmitter: Handle SVE MIN/MAX immediate (unpredicated) 2023-07-28 19:47:32 -04:00
Lioncache d629884147 ARMEmitter: Handle SVE ADD/SUB immediate (unpredicated) 2023-07-28 19:23:06 -04:00
Ryan Houdek 64c72430bf Merge pull request #2837 from lioncash/predicate
Arm64/Emitter: Add remaining missing SVE predicate range assertions
2023-07-28 14:44:43 -07:00
Lioncache 1b13e8db7d Arm64/Emitter: Simplify SVE constructive prefix
Moves the assert and op definition into the implementing function.
Also enforces valid use of predicate registers.
2023-07-28 17:17:51 -04:00
Ryan Houdek 1ce0ea8f3b Merge pull request #2835 from Sonicadvance1/cmn_alias
ARMEmitter: Implement cmn alias
2023-07-28 14:15:01 -07:00
Lioncache 8557656259 Arm64/Emitter: Simplify SVE FP unary operations predicated group
Moves the asserts into the implementation function and also enforces
valid use of predicates.
2023-07-28 17:07:15 -04:00
Ryan Houdek c87c04db59 ARMEmitter: Implement cmn alias
Inspired by #2832, since it was using adds directly.
cmn is an alias of adds with the destination register being zr.
2023-07-28 13:58:23 -07:00
Ryan Houdek f5262446a0 Merge pull request #2833 from Sonicadvance1/remove_erroneous_asserts
FEXCore: Remove erroneous asserts in the project
2023-07-28 13:53:42 -07:00
Ryan Houdek 0a1820d444 Merge pull request #2834 from lioncash/sveops
Arm64/Emitter: Deduplicate some more SVE implementations pt. 2
2023-07-28 13:50:12 -07:00
Ryan Houdek a255813e99 FEXCore: Remove erroneous asserts in the project
Inspired by #2832 by going through the source and removing uses of
`assert`.

assert doesn't work for us in debug builds because some games will
capture SIGABRT and continue running. So we need to use FEX's built in
assert handlers which call `FEX_TRAP_EXECUTION` which will take down the
FEX process as expected.

There wasn't too much usage of this in the source, so this is relatively
straightforward.
2023-07-28 13:36:05 -07:00
Lioncache 4c6336b560 Arm64/Emitter: Simplify load/store contiguous (scalar plus scalar) group
Lets us move the base opcode into the implementing function.
While we're at it, we can also add an assert to enforce predicate ranges.
2023-07-28 16:30:34 -04:00
Lioncache 145c7799a5 Arm64/Emitter: Simplify load/store contiguous (scalar plus immediate) group
We can move asserts and the base opcode into the implementing function.
While we're at it, we can add an assert to ensure predicate registers are in range.
2023-07-28 16:30:34 -04:00
Lioncache 91df503e55 Arm64/Emitter: Simplify load/store multiple contiguous (scalar plus immediate) group
We can centralize the base opcode and some of the asserts in the implementing function.
We can also add an assert that validates the predicate register range
and also make the offset assertions much more informative.
2023-07-28 16:30:30 -04:00
Lioncache 3c417cee84 Arm64/Emitter: Reduce qualifying in SVE FP/Int Convert To FP/Int groups
Since everything we need is already in the ARMEmitter namespace, we
don't need to qualify all type usages, which reduces the verbosity
a little during reading.
2023-07-28 14:57:43 -04:00
Lioncache 857780b46e Arm64/Emitter: Fix sorting of SVE integer to float group
This is actually already implemented but was grouped with the
SVE floating-point convert to integer group. We can extract this
out to be organized a little nicer.
2023-07-28 14:52:57 -04:00
Lioncache cb34aa3dee Arm64/Emitter: Simplify SVE FP/Int Convert To FP/Int groups
We can move the base opcode and asserts into the implementing function.
We can also add an assert to ensure the predicate is in a valid range.
2023-07-28 14:44:54 -04:00
Lioncache 0d74777954 Arm64/Emitter: Simplify SVE FP Unary Operations Predicated group
We can move the op into the implementing function and get rid of
an unnecessary function.

We can also add an assert to ensure the predicates are valid as well.
2023-07-28 14:24:28 -04:00
Ryan Houdek 092a023900 OpcodeDispatcher: Use VZip to make merge more optimal in PSAD
Thanks to @dougallj for noticing this.
Fixes #2818

New ASM
```asm
0x00007fffe2080808  10ffffe0            adr x0, #-0x4 (addr 0x7fffe2080804)
0x00007fffe208080c  f9005f80            str x0, [x28, #184]
0x00007fffe2080810  3dc03cc4            ldr q4, [x6, #240]
0x00007fffe2080814  2e247305            uabdl v5.8h, v24.8b, v4.8b
0x00007fffe2080818  6e247304            uabdl2 v4.8h, v24.16b, v4.16b
0x00007fffe208081c  4e71b8a5            addv h5, v5.8h
0x00007fffe2080820  4e71b884            addv h4, v4.8h
0x00007fffe2080824  4ec438b8            zip1 v24.2d, v5.2d, v4.2d
0x00007fffe2080828  58000040            ldr x0, pc+8 (addr 0x7fffe2080830)
0x00007fffe208082c  d63f0000            blr x0
```
2023-07-28 10:56:59 -07:00
Ryan Houdek b57dd5c3aa OpcodeDispatcher: Optimize PSAD* to use vuabdl{2,}
Optimizes this instruction from 18 down to 12 instructions.

Still has some extraneous moves at the tail of the operation that don't need to exist.

Before:
```asm
0x00007fffe2080908  10ffffe0            adr x0, #-0x4 (addr 0x7fffe2080904)
0x00007fffe208090c  f9005f80            str x0, [x28, #184]
0x00007fffe2080910  3dc03cc4            ldr q4, [x6, #240]
0x00007fffe2080914  2f08a705            uxtl v5.8h, v24.8b
0x00007fffe2080918  6f08a706            uxtl2 v6.8h, v24.16b
0x00007fffe208091c  2f08a487            uxtl v7.8h, v4.8b
0x00007fffe2080920  6f08a484            uxtl2 v4.8h, v4.16b
0x00007fffe2080924  6e6784a5            sub v5.8h, v5.8h, v7.8h
0x00007fffe2080928  6e6484c4            sub v4.8h, v6.8h, v4.8h
0x00007fffe208092c  4e60b8a5            abs v5.8h, v5.8h
0x00007fffe2080930  4e60b884            abs v4.8h, v4.8h
0x00007fffe2080934  4e71b8a5            addv h5, v5.8h
0x00007fffe2080938  4e71b884            addv h4, v4.8h
0x00007fffe208093c  4ea51ca0            mov v0.16b, v5.16b
0x00007fffe2080940  6e180480            mov v0.d[1], v4.d[0]
0x00007fffe2080944  4ea01c18            mov v24.16b, v0.16b
0x00007fffe2080948  58000040            ldr x0, pc+8 (addr 0x7fffe2080950)
0x00007fffe208094c  d63f0000            blr x0
```

After:
```asm
0x00007fffe2080848  10ffffe0            adr x0, #-0x4 (addr 0x7fffe2080844)
0x00007fffe208084c  f9005f80            str x0, [x28, #184]
0x00007fffe2080850  3dc03cc4            ldr q4, [x6, #240]
0x00007fffe2080854  2e247305            uabdl v5.8h, v24.8b, v4.8b
0x00007fffe2080858  6e247304            uabdl2 v4.8h, v24.16b, v4.16b
0x00007fffe208085c  4e71b8a5            addv h5, v5.8h
0x00007fffe2080860  4e71b884            addv h4, v4.8h
0x00007fffe2080864  4ea51ca0            mov v0.16b, v5.16b
0x00007fffe2080868  6e180480            mov v0.d[1], v4.d[0]
0x00007fffe208086c  4ea01c18            mov v24.16b, v0.16b
0x00007fffe2080870  58000040            ldr x0, pc+8 (addr 0x7fffe2080878)
0x00007fffe2080874  d63f0000            blr x0
```
2023-07-28 10:56:59 -07:00
Ryan Houdek 0bea508935 IR: Adds support for VUABDL2 operation
FEX already supported VUABDL, but we had missed VUABDL2.
2023-07-28 10:56:59 -07:00
Ryan Houdek 9c175da0e2 Merge pull request #2831 from lioncash/moreop
Arm64/Emitter: Deduplicate some more SVE implementations
2023-07-28 10:47:50 -07:00
Lioncache cc359f09ad Arm64/Emitter: Simplify SVE FP convert precision odd elements
We can move the base opcode into the implementation function.
We can also add an assert to ensure that the predicate is in a valid range.
2023-07-28 13:34:04 -04:00
Lioncache dedeba0575 Arm64/Emitter: Simplify SVE2 Character Match group
We can move the base opcode into the implementation function.
2023-07-28 13:34:04 -04:00
Lioncache dc3ccbcc43 Arm64/Emitter: Simplify SVE SEL instruction
We can move the base opcode and asserts into the implementing function.
We can also implement the mov alias in terms of sel()
2023-07-28 13:34:04 -04:00
Lioncache ebd9b49dbe Arm64/Emitter: Simplify SVE2 bitwise ternary operations category
We can move the asserts into the implementing function.
2023-07-28 13:34:04 -04:00
Lioncache 43e21ad749 Arm64/Emitter: Simplify SVE permute vector elements category
We can move the asserts and the base opcode into the implementing function.
2023-07-28 13:34:04 -04:00
Lioncache 1b3beec085 Arm64/Emitter: Add missing SVE TBL encoding
Adds the missing SVE2 double table encoding.

Crosses off a TODO.
2023-07-28 13:34:04 -04:00
Lioncache c88f2e0730 Arm64/Emitter: Simplify SVE table lookup category
We can move the base opcode into the implementing function.
2023-07-28 13:34:04 -04:00
Lioncache 1645f6e582 Arm64/Emitter: Simplify SVE integer add/sub vectors unpredicated category
We can move the base op encoding into the implementing function.
2023-07-28 13:34:04 -04:00
Mai af35e18979 Merge pull request #2830 from Sonicadvance1/fix_smc_race
FEXLinuxTests: Fixes race in smc-mt-2
2023-07-28 13:23:08 -04:00
Ryan Houdek 012c0ae062 FEXLinuxTests: Fixes race in smc-mt-2
This test was written to test SMC where one thread is doing execution
while the other thread is modifying.

According to the printf documentation it is supposed to "wait for code
to be modified" but actually it was testing a race between a printf on
one thread and the primary thread modifying the code.

Fix this test so it is actually waiting for code modification to happen
rather than testing a race condition. This is likely what the original
author intended.

CI is hitting this flake more frequently now because it is even faster
it seems, so fixing this test is necessary to resolve these flakes.
2023-07-27 19:24:57 -07:00
Ryan Houdek dc8f063a4b Merge pull request #2823 from Sonicadvance1/fix_mingw_build
Mingw: Fixes compiling again
2023-07-27 18:14:07 -07:00
Ryan Houdek 759747b3c5 Merge pull request #2828 from lioncash/unused
LongDivideRemovalPass: Remove unused variable
2023-07-27 16:27:56 -07:00
Lioncache 4c6d26f167 LongDivideRemovalPass: Remove unused variable 2023-07-27 19:13:29 -04:00
Ryan Houdek 7562ff2308 Merge pull request #2827 from lioncash/alias
Context: Pull out some long std::function declarations into aliases
2023-07-27 16:11:20 -07:00
Lioncache e2baa1106b Context: Make alias for code invalidation functor
Makes it so any changes wouldn't need to go across everything implementing the Context class.
2023-07-27 18:59:15 -04:00
Lioncache 68880fd66b Context: Make alias for custom IR handlers
Puts the type declaration in one place.
2023-07-27 18:43:19 -04:00
Lioncache 62eef6c548 AOTIR: std::move function instances where applicable
Ensures no unnecessary allocations occur.
2023-07-27 18:35:30 -04:00
Lioncache cfdc1bda47 AOTIR: Add aliases for long functors
Puts the really long declarations in one spot, which makes them
nicer to change over time.
2023-07-27 18:30:07 -04:00
Ryan Houdek f9a9645ef3 Merge pull request #2826 from lioncash/unused
Frontend: Remove unused ModRMDecoded instance
2023-07-27 14:30:19 -07:00
Lioncache 8fc903cc51 Frontend: Remove unused ModRMDecoded instance
This is assigned to but never used.
2023-07-27 17:18:16 -04:00
Ryan Houdek c81b1c3432 Merge pull request #2825 from lioncash/lookup
Frontend: Remove redundant lookups in BranchTargetInMultiblockRange()
2023-07-27 14:00:40 -07:00
Lioncache db9ba16a53 Frontend: Remove redundant lookups in BranchTargetInMultiblockRange()
insert() will only ever perform an insertion if the relevant element
doesn't exist within the set, so we were doing an unnecessary lookup
in two spots.

We don't use emplace() here because it will need to construct the
key for the element inside the allocated node (since the key may
be non-copyable/non-movable), causing an allocation even if an insert
doesn't actually occur.

Conversely, insert() will not need to allocate and construct a node
ahead of time if an element already exists in the set.
2023-07-27 16:28:37 -04:00
Ryan Houdek 0be68a54d0 Merge pull request #2824 from lioncash/opsimp
Arm64/Emitter: Reorganize some base opcode and assert locations
2023-07-27 13:18:23 -07:00
Lioncache 6e140e9ff2 Arm64/Emitter: Simplify SVE Integer Compare With Unsigned Immediate category
We can move the asserts into the implementation function.
2023-07-27 14:12:16 -04:00
Lioncache d793f7295a Arm64/Emitter: Add missing cmpne signed immediate op
Noticed while simplifying the signed immediate implementation.
2023-07-27 14:08:26 -04:00
Lioncache 8f6018a1b4 Arm64/Emitter: Simplify SVE Integer Compare With Signed Immediate category
We can move all asserts and handling into the implementation function.
2023-07-27 14:03:34 -04:00
Lioncache 0a97d4456d Arm64/Emitter: Simplify SVE Predicate Logical Operations category
We can move the base opcode into the implementation function.
2023-07-27 13:51:48 -04:00
Lioncache c2dbe8309b Arm64/Emitter: Simplify SVE Bitwise Logical Operations Predicated category
We can move the asserts and base opcode into the implementation function.
We can also add an assert to ensure a valid predicate range.
2023-07-27 13:41:49 -04:00
Lioncache 06d66a0188 Arm64/Emitter: Simplify SVE Integer Mul/Div Vectors Predicated category
We can move the base opcode into the implementation function.
We can also move the division assert into it as well.
2023-07-27 13:36:22 -04:00
Lioncache ef254da7ca Arm64/Emitter: Simplify SVE Integer Min/Max/Difference Predicated category
We can move the asserts into the implementation function and also add
another assert to validate the predicate range.
2023-07-27 13:29:15 -04:00
Lioncache 162dcf3cd0 Arm64/Emitter: Simplify SVE Integer Add/Sub Vectors Predicated category
We can move the opcode into the implementation function.
2023-07-27 13:19:56 -04:00
Lioncache dcc47074d1 Arm64/Emitter: Simplify SVE Integer Multiply-Add Predicated category
We can move the op definition into the implementation function.
2023-07-27 13:16:11 -04:00
Lioncache 4dfee31012 Arm64/Emitter: Simplify SVE FP Recursive Reduction category
We can move the op constant into the implementation function.
2023-07-27 13:10:12 -04:00
Lioncache 3749fa6569 Arm64/Emitter: Simplify SVE Float Arithmetic Unpredicated category
We can move the common assert and op value into the implementation function.
2023-07-27 13:07:46 -04:00
Mai 94273fbf4f Merge pull request #2760 from Sonicadvance1/switch_to_half_barriers
Arm64: Switch to using half barriers
2023-07-27 09:07:14 -04:00
Ryan Houdek 162bbf2937 Mingw: Fixes compiling again
Getting the CI machines setup to handle this so we can stop breaking it.
2023-07-26 16:52:19 -07:00
Mai 296adf1830 Merge pull request #2822 from Sonicadvance1/fix_struct_pack_verifier_cursor
StructPackVerifier: Fixes missing cursorkind again
2023-07-26 14:47:43 -04:00
Ryan Houdek a86109280a StructPackVerifier: Fixes missing cursorkind again
python-clang is getting pretty bad with missing cursor kinds these days.
This was hit from clang-15, will be required for new CI runners
2023-07-26 11:21:16 -07:00
Mai d83960d4dd Merge pull request #2821 from Sonicadvance1/rename_valuenode
RCLSE: Rename `Node` to `ValueNode`
2023-07-25 16:06:12 -04:00
Ryan Houdek bba97823a8 RCLSE: Rename Node to ValueNode
Every time I look at this I can never remember which node is the value
node and which one is the store node.

Rename the opaque `Node` to `ValueNode` and have documentation comments
to explain which one it is. This way I won't forget again.
2023-07-25 12:19:15 -07:00
Ryan Houdek abf5e8c6a5 Merge pull request #2820 from alyssarosenzweig/pf/shift
Preserve PF across zero shift
2023-07-25 08:43:16 -07:00
Alyssa Rosenzweig 8d288300c4 unittests: Add test for preserving PF across zero shift
Tests for the regression from 7e6bb04db ("OpcodeDispatcher: Extract
CalculatePF"). This fails on main but passes with this PR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-25 11:21:52 -04:00
Alyssa Rosenzweig 0caf263c77 OpcodeDispatcher: Preserve PF across zero shifts
Reported by CATFELLA on Discord. uwu

Fixes: 7e6bb04db ("OpcodeDispatcher: Extract CalculatePF")
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-25 11:21:52 -04:00
Ryan Houdek 29dc77c44b Merge pull request #2819 from alyssarosenzweig/const/bfi
ConstProp: Handle constant Bfi
2023-07-25 07:34:57 -07:00
Alyssa Rosenzweig 4c1f53c1ff ConstProp: Handle constant Bfi
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-25 10:19:56 -04:00
Ryan Houdek 5821175ddb Merge pull request #2813 from Sonicadvance1/fix_ra_lr
Arm64: Fixes LR corruption in 128-bit divides
2023-07-24 15:08:45 -07:00
Ryan Houdek 003c88e537 Merge pull request #2816 from alyssarosenzweig/ir/bump-16
IR: Expand to 16-bit opcodes
2023-07-24 15:08:36 -07:00
Alyssa Rosenzweig 27a1ebc2f5 IR: Expand to 16-bit opcodes
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-24 17:21:11 -04:00
Ryan Houdek e689c6fcfa Merge pull request #2814 from lioncash/shift
Arm64/Emitter: Simplify SVE immediate shift helper
2023-07-24 12:28:53 -07:00
Ryan Houdek 61e905a339 Arm64: Fixes LR corruption in 128-bit divides
Need to save and restore LR before branching out to the helpers.
Confirmed that the rest of the JIT handles this correctly.
2023-07-24 12:04:16 -07:00
Lioncache 10fdcaa109 Arm64/Emitter: Simplify SVE immediate shift helper
We can collapse the entire if statement to be much simpler.
2023-07-24 15:04:14 -04:00
Ryan Houdek ba672f868c Arm64: Switch to using half barriers
Inspired from: https://github.com/dotnet/runtime/issues/8072

Currently FEX is /very/ heavy handed with our backpatching where we wrap
every backpatched loadstore with `dmb ish`.

This can be relaxed slightly according to the linked issue.

For TSO load instructions the instruction sequence changes to:
  ldr <args>;
  dmb ld; <-- Slightly less strict dmb

For TSO store instructions the instruction sequence changes to:
  dmb ish; <-- Still the all encompassing dmb
  str <args>;

For backpatching loadstores this does the same thing where only one side
needs the nop and it uses the same instruction sequence when
backpatched.

The minor change is that on load backpatching, we are no longer backing
up a single instruction, instead just re-executing the instruction we
patched directly.

Took a long time to come back to this (Last looked in August 2020).
Previously when I was implementing this idea it didn't work, but that
was because our CompareExchange operation was broken back then. With the
CAS now, it should just work.
2023-07-24 11:06:20 -07:00
Ryan Houdek d6697fce32 Merge pull request #2812 from lioncash/dup
Arm64/Emitter: Collapse encoding cases for indexed dup
2023-07-24 11:06:04 -07:00
Lioncache 9c3a843df7 Arm64/Emitter: Move indexed dup handling into SVEDup
SVEDup is only used by dup(), so we can move all the implementation
details into it instead of keeping it all in the public function.
2023-07-24 13:53:57 -04:00
Lioncache 8defa2b55f Arm64/Emitter: Collapse encoding cases for indexed dup
Lets us hoist out the asserts and also collapse all the
branching into one series of operations.
2023-07-24 13:53:42 -04:00
Ryan Houdek 77c88ffe53 Merge pull request #2804 from alyssarosenzweig/eor-zero
Optimize `xor %eax, %eax`
2023-07-24 09:01:53 -07:00
Ryan Houdek 5b261b0d2e Merge pull request #2809 from lioncash/assert
Arm64/VectorOps: Hoist asserts out of VInsElement cases
2023-07-24 09:00:41 -07:00
Alyssa Rosenzweig 2ce15ddc89 OpcodeDispatcher: Partially defer PF calculation
We expect that PF is written more often than it's read, so we want to
get the expensive popcount out of the hot path. (Thank you to Dougall
for suggesting that.)

There are two cases:

1. PF is written by an integer instruction. In this case, we calculate
   with the formula `popcount(x ^ 1) & 1`.
2. PF is written by a float instruction, copying a host flag.

What we really want is to defer the relatively expensive popcount. So,
to unify these cases, we have integer instructions write `x ^ 1` and
(unchanged) float instructions write the host flag. Then, when reading
PF, we do `popcount(value) & 1` on the byte read in.

If PF is written but not read, this saves the expensive popcount and
leaves only the cheap xor.

If PF is written by an integer op and read, this maybe shuffles some
code but does not materially change anything.

If PF is written by a float op and read, this is worse because now we're
doing an extra pointless popcount. This is a tradeoff... However, this
is only relevant to unordered float comparisons, which I expect to be
obscure for games. So this should be worth it over all (for games, if
not weird numerical computing workloads).

How does this connect to my register zeroing quest? The constant folding
code doesn't currently deal with FPRs and I'm not in a mood to change
this. So before, a block ending with `xor eax, eax` would still do a
popcount for PF. Now it just writes a constant 1 since the xor constant
folds and the popcount never happens at all.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-24 11:17:31 -04:00
Alyssa Rosenzweig 3d1b55383e ConstProp: Handle Select::EQ
For flag calculation after moving a constant. This cleans up the code
generated for zeroing at the end of a block.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-24 11:17:31 -04:00
Alyssa Rosenzweig 69deaa0976 LongDivideRemovalPass: Don't detect xor zero
It is now optimized out (canonicalized) in the frontend.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-24 11:17:31 -04:00
Alyssa Rosenzweig fc72fa9e5f OpcodeDispatcher: Optimize xor zeroing
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-24 11:17:31 -04:00
Mai 6baee3b7c1 Merge pull request #2808 from Sonicadvance1/timeout_github_upload
github-actions: Adds timeout to upload results
2023-07-24 11:14:06 -04:00
Lioncache 362a5a019a Arm64/VectorOps: Move VInsElement size assert checks to be first
Ensures the asserts will be hit first before emitting code.
2023-07-24 11:08:35 -04:00
Lioncache d529a58893 Arm64/VectorOps: Hoist asserts out of VInsElement cases
Lets us deduplicate the asserts and put them in one spot. We can also
improve the assert message to also indicate the valid range.

While we're in the area, we can collapse a few case paths
as a result of this assert movement.
2023-07-24 11:05:13 -04:00
Ryan Houdek c77ea3b392 github-actions: Adds timeout to upload results
Usually uploading of results takes about two seconds.
Sometimes github's connection to the runner flakes and it stalls out the
upload action for some reason.

Github's default timeout is SIX HOURS.
Change this to a one-minute timeout on the upload step so it quickly
goes away when Github's internet flakes out.
2023-07-21 15:32:36 -07:00
Ryan Houdek 20aaad15e4 Merge pull request #2806 from Sonicadvance1/for_2804
X86Tables: Adds spaceship operator to couple op types
2023-07-21 15:29:58 -07:00
Ryan Houdek 6f2452e2ea Merge pull request #2807 from Sonicadvance1/gvisor_socket_flakes
unittests/gvisor: Adds all socket tests to flakes
2023-07-21 15:26:51 -07:00
Ryan Houdek b34401bb33 Merge pull request #2805 from bylaws/win32-fix
FHU: Avoid calling faccessat on WIN32
2023-07-21 15:00:48 -07:00
Ryan Houdek dff9868e9a unittests/gvisor: Adds all socket tests to flakes
All of these flake so slam them all in the flake file.
2023-07-21 14:50:58 -07:00
Ryan Houdek 036a196984 X86Tables: Adds spaceship operator to couple op types
For #2804 so it can compare if GPRs match more easily.

Only adding to GPR and Literal types, since its ambiguous what this
would mean for the memory accessing types. Going to leave those other
ones alone for now.
2023-07-21 14:42:01 -07:00
Billy Laws a7ac4fa6e4 FHU: Avoid calling faccessat on WIN32 2023-07-21 22:21:15 +01:00
Ryan Houdek 82295b2943 Merge pull request #2763 from Sonicadvance1/fix_x11_variadic_thunks
Thunks/X11: Fixes variadic packing and callbacks.
2023-07-21 01:49:26 -07:00
Ryan Houdek eca80b6046 Remove inline assembly
64 should be enough for anybody.
2023-07-20 16:05:25 -07:00
Ryan Houdek 02fc02964b Thunks/LibX11: Removes variadic multiple argument copying
Just copy one at a time.
2023-07-20 13:52:30 -07:00
Ryan Houdek e8fcb070b3 Thunks/X11: Fixes variadic packing and callbacks.
Two bugs here that caused thunking X11 thunking in Wine/Proton to not
work.

The easier of the two. The various variadic functions that we thunk
actually take key:value pairs where the first is a string pointer, and
the value can be various things.

We need to handle these as true key:value pairs rather than finding the
first nullptr and dropping the remainder.

Additionally, there are 12 keys that specify a callback that FEX needs
to catch and convert to host callable. Wine is the first application
that I have seen that actually uses this. If these callbacks aren't
wired up then it it can miss events.

The harder of the two problems is the `libX11_Variadic_u64` function was
subtly incorrect. Nothing had previously truly exercised this and my
test program didn't notice anything wrong while writing it.

The first incorrect thing was that it was subtracting the nullptr ender
variable before the stack size calculation, causing the value to
overwrite the stack if the number of remaining elements was event.

Secondly the assembly that was storing two elements per step was
decrementing the counter by 8 instead of two. Didn't pick this up before
since I believe the code was only hitting the non-pair path before.

This gets Proton thunking working under FEX now.
2023-07-20 13:52:30 -07:00
Ryan Houdek 597da88035 Merge pull request #2803 from Sonicadvance1/xcb_x11_dependency
ThunksDB: Adds X11 dependency to XCB
2023-07-19 15:45:47 -07:00
Ryan Houdek 1e829a47fa ThunksDB: Adds X11 dependency to XCB
I was hitting an issue where thunking in Wine+Vulkan applications was breaking
unless both OpenGL and Vulkan was enabled.

Turns out this was because Vulkan enabled only XCB, which didn't enable
X11. So when XCB is thunked but X11 isn't, this causes weird issues
where X11 calls in to XCB functions and gets in desync'ed state. Causing
hangs to appear in xcb_take_socket.

Now we can enabled just `Vulkan` as a thunk and it'll work fine.

```
(gdb) bt
   from target:/usr/lib/fex-emu/HostThunks//libvulkan-host.so
```
2023-07-19 15:23:07 -07:00
Ryan Houdek e633ef7cf5 Merge pull request #2801 from lioncash/cvt
Arm64/Emitter: Handle SVE FP convert precision group
2023-07-19 13:45:56 -07:00
Ryan Houdek ff3b40400c Merge pull request #2795 from Sonicadvance1/remove_thunk_symbol
Thunks: Remove weak symbol definitions
2023-07-19 13:45:36 -07:00
Lioncache 801106faf4 Arm64/Emitter: Handle SVE FP convert precision group 2023-07-19 15:24:54 -04:00
Ryan Houdek 536b2ed495 Merge pull request #2783 from Sonicadvance1/optimize_loadstorecontextindexed
Arm64: Optimize {Load,Store}ContextIndexed address generation
2023-07-19 12:16:39 -07:00
Alyssa Rosenzweig 072f027885 Merge pull request #2784 from Sonicadvance1/optimize_small_rcr
OpcodeDispatcher: Optimize 8/16-bit RCR
2023-07-19 15:07:58 -04:00
Ryan Houdek 754bc18813 Merge pull request #2800 from lioncash/fpa
Arm64/Emitter: Handle SVE FP arithmetic with immediate (predicated) group
2023-07-19 11:01:02 -07:00
Lioncache ef257418d3 Arm64/Emitter: Handle SVE FP arithmetic with immediate (predicated) group 2023-07-19 13:37:33 -04:00
Ryan Houdek 842b71cf83 Merge pull request #2799 from lioncash/xar
Arm64/Emitter: Handle SVE XAR
2023-07-19 09:39:30 -07:00
Ryan Houdek 8d8b64d2b7 Merge pull request #2798 from bylaws/dealock
Jit: Add block links directly through the lookup cache on thread exit
2023-07-19 09:37:57 -07:00
Lioncache 85c6ef8097 Arm64/Emitter: Handle SVE XAR
Now that we have the helper for encoding immediate shifts,
we can trivially implement the remaining missing instruction
in the bitwise logical unpredicated group.
2023-07-19 12:26:05 -04:00
Billy Laws 2be16d9054 Jit: Add block links directly through the lookup cache on thread exit
Prevents the code invalidation mutex from being locked as shared recursively,
since it is locked before entering ThreadExitFunctionLink and would end up
being locked again by ThreadAddBlockLink.

This fixes a deadlock on Windows.
2023-07-19 17:13:17 +01:00
Ryan Houdek 41b3c52663 Merge pull request #2797 from lioncash/shift
Arm64/Emitter: Add helper for encoding SVE shift immediates
2023-07-19 09:10:21 -07:00
Lioncache f1f50b7a98 Arm64/Emitter: Add helper for encoding SVE shift immediates
This lets us deduplicate the behavior rather than open-coding it everywhere
and also makes it nicer to implement instructions that make use of this
encoding pattern.
2023-07-19 11:39:28 -04:00
Ryan Houdek d7200e2a1e Merge pull request #2796 from lioncash/unary
Arm64/Emitter: Deduplicate some opcode values
2023-07-19 06:56:12 -07:00
Lioncache 3f884fe2d0 Arm64/Emitter: Deduplicate ops in SVE Bitwise Logical - Unpredicated
Same behavior, with less code.
2023-07-19 09:27:13 -04:00
Lioncache 1d453f10a0 Arm64/Emitter: Remove unnecessary qualifiers from integer unary arith ops
While we're in the area, we can make these much quicker to read by
removing the unnecessary qualifiers.
2023-07-19 09:18:36 -04:00
Lioncache 5ebd21ca6a Arm64/Emitter: Deduplicate integer unary arithmetic instructions
We can move the opcodes into the underlying implementation along with
the asserts to deduplicate a bit of code.
2023-07-19 09:16:33 -04:00
Ryan Houdek fb34b507a1 Thunks: Remove weak symbol definitions
Fixes #2754

These panicking fallbacks are at times not ending up in as plt calls
for some reason that I haven't been able to reproduce locally.

So far the only way I can reproduce is building with Canonical's PPA
build system, since rebuilding locally didn't resolve the issue.

This will change the failure mode from these panicking asserts happening
at call time, to dlopen failing during relocation when loading the
thunk. Which LD_DEBUG=all can be used for debugging relocation failure
in that case.
2023-07-18 22:33:58 -07:00
Ryan Houdek 2c91b5cff6 Merge pull request #2788 from Sonicadvance1/enum_configs_disassembler
FEXCore/Config: Adds support for enum mask configuration array
2023-07-18 18:40:22 -07:00
Ryan Houdek 79f7dcbaa5 FEXCore/Config: Adds support for enum mask configuration array
Allows us to consume an array of strings and convert it to an mask of
enum values. This is a quality of life change that allows us to specify
a mask of options.

The first configuration option added to support this is to control the
vixl disassembler. Now by default the vixl disassembler doesn't
disassemble any blocks and needs to be enabled individually.

eg:
```
FEXLoader --disassemble=blocks <args>
FEXLoader --disassemble=dispatcher <args>
FEXLoader --disassemble=blocks,dispatcher <args>
```

Has the additional convenience option of just passing in numbers as
well.

```
FEXLoader --disassemble=2 <args>
FEXLoader --disassemble=1 <args>
FEXLoader --disassemble=3 <args>
```

Also of course all of this works through environment variables.
```
FEX_DISASSEMBLE=blocks FEXInterpreter <args>
FEX_DISASSEMBLE=dispatcher FEXInterpreter <args>
FEX_DISASSEMBLE=blocks,dispatcher FEXInterpreter <args>
```

While only used fairly sparingly now, this is likely to have some
additional configurations using this in the future. Since we already
have some configs that are basically using enums, but just by doing
string comparisons.

This was asked for by a developer, so I figured I would throw it
together quick.
2023-07-18 18:17:43 -07:00
Ryan Houdek f7b7997c77 Merge pull request #2786 from Sonicadvance1/minor_fcmov_opt2
OpcodeDispatcher: Another FCMov minor optimization
2023-07-18 18:14:17 -07:00
Ryan Houdek 0674dfab0a Merge pull request #2787 from Sonicadvance1/icache_only_code
Arm64: Only clear icache for code
2023-07-18 18:14:05 -07:00
Ryan Houdek 6979dc9c4e Merge pull request #2794 from lioncash/decvsib
OpcodeDispatcher: Handle VSIB byte
2023-07-18 11:28:32 -07:00
Lioncache fe5f17d92e OpcodeDispatcher: Handle VSIB byte
Ensures that we handle the AVX2 VSIB byte in a decent way.

As is, we can't compute the [index * scale] variant portion
of the entire address operand, since the scale needs to act
on every element of the vector after sign extension.

What we can do though, is compute the base address and add
the displacement to it ahead of time though.
2023-07-18 13:56:38 -04:00
Ryan Houdek be71886990 Merge pull request #2793 from lioncash/unusedvar
OpcodeDispatcher: Remove unused member variables and reorganize
2023-07-18 08:32:47 -07:00
Lioncache 5043e5fbc0 OpcodeDispatcher: Move ShouldDump member into private section
Like with HandledLock, we can move this into the private section
and put an API around it for consistency.
2023-07-18 11:03:25 -04:00
Lioncache 2acfde3cad OpcodeDispatcher: Move CTX member into private section
This isn't used outside of the class.
2023-07-18 11:00:04 -04:00
Lioncache c1eeeaf688 OpcodeDispatcher: Move flag-related variables into private section
These are only used within the opcode dispatcher, so they can be private.
2023-07-18 10:58:09 -04:00
Lioncache f2b3229a87 OpcodeDispatcher: Move JumpTargets into private section
This is only used within the class, so it can be made private.
2023-07-18 10:52:48 -04:00
Lioncache e20bfc0701 OpcodeDispatcher: Move HandledLock boolean into private class section
This can be trivially hidden and have an API put around it.
2023-07-18 10:43:52 -04:00
Lioncache 4caee5c9be OpcodeDispatcher: Remove unused Current_Header and Current_HeaderNode variables
These aren't used outside of being assigned to, so they can be removed.
2023-07-18 10:40:26 -04:00
Ryan Houdek 46d3f283b8 Merge pull request #2792 from lioncash/strun
Arm64/Emitter: Handle unsized contiguous STR variants
2023-07-18 05:54:32 -07:00
Lioncache cee5512a56 Arm64/Emitter: Handle unsized contiguous STR variants
We handle the unsized load variants, so we should do the same with the stores.
2023-07-18 08:38:23 -04:00
Mai 8c53a373bf Merge pull request #2791 from Sonicadvance1/fix_xcb_version_typo
Thunks/xcb: Fixes typo in version check.
2023-07-18 07:36:14 -04:00
Ryan Houdek f73176d5dc Thunks/xcb: Fixes typo in version check.
Somehow this 14 turned in to a 4.
2023-07-18 04:16:11 -07:00
Ryan Houdek 80ae3e632d Arm64: Optimize {Load,Store}ContextIndexed address generation
All of these IR operations were being fairly inefficient in their
address calculation. All of these are known using power of 2 stride
indexing. So all of these can be converted from three instructions to
one.

These are always used for x87 stack accesses so each one gets an
improvement.

Before:
```asm
0x0000ffff6a800248  d2800200    mov x0, #0x10
0x0000ffff6a80024c  9b007e80    mul x0, x20, x0
0x0000ffff6a800250  8b000380    add x0, x28, x0
0x0000ffff6a800254  fd417805    ldr d5, [x0, #752]
```

After:
```asm
0x0000ffff91e80240  8b141380    add x0, x28, x20, lsl #4
0x0000ffff91e80244  fd417805    ldr d5, [x0, #752]
```
2023-07-17 22:59:33 -07:00
Ryan Houdek ed75c19324 OpcodeDispatcher: Optimize 8/16-bit RCR
The BFI cascades in this particular instruction weren't optimal.
Biggest improvement is the 8-bit version, while the 16-bit version gets
a minor improvement.

8-bit instruction count reduced from 38 to 29.
16-bit instruction count reduced from 34 to 28.

RCL can have a similar optimization done to it.
```asm
Before 16-bit:
0x0000ffff80a801e0  10ffffe0    adr x0, #-0x4 (addr 0xffff80a801dc)
0x0000ffff80a801e4  f9005f80    str x0, [x28, #184]
0x0000ffff80a801e8  d3403cb4    uxth x20, w5
0x0000ffff80a801ec  d3403cf5    uxth x21, w7
0x0000ffff80a801f0  394b0396    ldrb w22, [x28, #704]
0x0000ffff80a801f4  12001294    and w20, w20, #0x1f
0x0000ffff80a801f8  d2800017    mov x23, #0x0
0x0000ffff80a801fc  b3403eb7    bfxil x23, x21, #0, #16
0x0000ffff80a80200  b37002d7    bfi x23, x22, #16, #1
0x0000ffff80a80204  b36f3eb7    bfi x23, x21, #17, #16
0x0000ffff80a80208  b35f02d7    bfi x23, x22, #33, #1
0x0000ffff80a8020c  aa1703e0    mov x0, x23
0x0000ffff80a80210  b35e3ea0    bfi x0, x21, #34, #16
0x0000ffff80a80214  aa0003f5    mov x21, x0
0x0000ffff80a80218  b34e02d5    bfi x21, x22, #50, #1
0x0000ffff80a8021c  9ad426b7    lsr x23, x21, x20
0x0000ffff80a80220  b3403ee7    bfxil x7, x23, #0, #16
0x0000ffff80a80224  51000698    sub w24, w20, #0x1 (1)
0x0000ffff80a80228  9ad826b5    lsr x21, x21, x24
0x0000ffff80a8022c  d34002b5    ubfx x21, x21, #0, #1
0x0000ffff80a80230  7100069f    cmp w20, #0x1 (1)
0x0000ffff80a80234  9a9622b4    csel x20, x21, x22, hs
0x0000ffff80a80238  390b0394    strb w20, [x28, #704]
0x0000ffff80a8023c  d34f3ef4    ubfx x20, x23, #15, #1
0x0000ffff80a80240  d34e3af5    ubfx x21, x23, #14, #1
0x0000ffff80a80244  ca150294    eor x20, x20, x21
0x0000ffff80a80248  390b2f94    strb w20, [x28, #715]
0x0000ffff80a8024c  58000040    ldr x0, pc+8 (addr 0xffff80a80254)
0x0000ffff80a80250  d63f0000    blr x0
0x0000ffff80a80254  967da128    bl #-0x6097b60 (addr 0xffff7a9e86f4)
0x0000ffff80a80258  0000ffff    udf #0xffff
0x0000ffff80a8025c  00010023    unallocated (Unallocated)
0x0000ffff80a80260  00000000    udf #0x0
[DEBUG] RIP: 0x10020
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 34
[DEBUG] Blow-up Amt: 34x

After 16-bit:
0x0000ffffa7c801e0  10ffffe0            adr x0, #-0x4 (addr 0xffffa7c801dc)
0x0000ffffa7c801e4  f9005f80            str x0, [x28, #184]
0x0000ffffa7c801e8  d3403cb4            uxth x20, w5
0x0000ffffa7c801ec  d3403cf5            uxth x21, w7
0x0000ffffa7c801f0  394b0396            ldrb w22, [x28, #704]
0x0000ffffa7c801f4  12001294            and w20, w20, #0x1f
0x0000ffffa7c801f8  b37002d5            bfi x21, x22, #16, #1
0x0000ffffa7c801fc  b36f42b5            bfi x21, x21, #17, #17
0x0000ffffa7c80200  b35e42b5            bfi x21, x21, #34, #17
0x0000ffffa7c80204  9ad426b7            lsr x23, x21, x20
0x0000ffffa7c80208  b3403ee7            bfxil x7, x23, #0, #16
0x0000ffffa7c8020c  51000698            sub w24, w20, #0x1 (1)
0x0000ffffa7c80210  9ad826b5            lsr x21, x21, x24
0x0000ffffa7c80214  d34002b5            ubfx x21, x21, #0, #1
0x0000ffffa7c80218  7100069f            cmp w20, #0x1 (1)
0x0000ffffa7c8021c  9a9622b4            csel x20, x21, x22, hs
0x0000ffffa7c80220  390b0394            strb w20, [x28, #704]
0x0000ffffa7c80224  d34f3ef4            ubfx x20, x23, #15, #1
0x0000ffffa7c80228  d34e3af5            ubfx x21, x23, #14, #1
0x0000ffffa7c8022c  ca150294            eor x20, x20, x21
0x0000ffffa7c80230  390b2f94            strb w20, [x28, #715]
0x0000ffffa7c80234  58000040            ldr x0, pc+8 (addr 0xffffa7c8023c)
0x0000ffffa7c80238  d63f0000            blr x0
0x0000ffffa7c8023c  bd9cc128            unallocated (Unallocated)
0x0000ffffa7c80240  0000ffff            udf #0xffff
0x0000ffffa7c80244  00010023            unallocated (Unallocated)
0x0000ffffa7c80248  00000000            udf #0x0
[DEBUG] RIP: 0x10020
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 28
[DEBUG] Blow-up Amt: 28x

Before 8-bit:
0x0000ffffa92801e0  10ffffe0            adr x0, #-0x4 (addr 0xffffa92801dc)
0x0000ffffa92801e4  f9005f80            str x0, [x28, #184]
0x0000ffffa92801e8  d3401cb4            uxtb x20, w5
0x0000ffffa92801ec  d3401cf5            uxtb x21, w7
0x0000ffffa92801f0  394b0396            ldrb w22, [x28, #704]
0x0000ffffa92801f4  12001294            and w20, w20, #0x1f
0x0000ffffa92801f8  d2800017            mov x23, #0x0
0x0000ffffa92801fc  b3401eb7            bfxil x23, x21, #0, #8
0x0000ffffa9280200  b37802d7            bfi x23, x22, #8, #1
0x0000ffffa9280204  b3771eb7            bfi x23, x21, #9, #8
0x0000ffffa9280208  b36f02d7            bfi x23, x22, #17, #1
0x0000ffffa928020c  b36e1eb7            bfi x23, x21, #18, #8
0x0000ffffa9280210  b36602d7            bfi x23, x22, #26, #1
0x0000ffffa9280214  b3651eb7            bfi x23, x21, #27, #8
0x0000ffffa9280218  b35d02d7            bfi x23, x22, #35, #1
0x0000ffffa928021c  aa1703e0            mov x0, x23
0x0000ffffa9280220  b35c1ea0            bfi x0, x21, #36, #8
0x0000ffffa9280224  aa0003f5            mov x21, x0
0x0000ffffa9280228  b35402d5            bfi x21, x22, #44, #1
0x0000ffffa928022c  9ad426b7            lsr x23, x21, x20
0x0000ffffa9280230  b3401ee7            bfxil x7, x23, #0, #8
0x0000ffffa9280234  51000698            sub w24, w20, #0x1 (1)
0x0000ffffa9280238  9ad826b5            lsr x21, x21, x24
0x0000ffffa928023c  d34002b5            ubfx x21, x21, #0, #1
0x0000ffffa9280240  7100069f            cmp w20, #0x1 (1)
0x0000ffffa9280244  9a9622b4            csel x20, x21, x22, hs
0x0000ffffa9280248  390b0394            strb w20, [x28, #704]
0x0000ffffa928024c  d3471ef4            ubfx x20, x23, #7, #1
0x0000ffffa9280250  d3461af5            ubfx x21, x23, #6, #1
0x0000ffffa9280254  ca150294            eor x20, x20, x21
0x0000ffffa9280258  390b2f94            strb w20, [x28, #715]
0x0000ffffa928025c  58000040            ldr x0, pc+8 (addr 0xffffa9280264)
0x0000ffffa9280260  d63f0000            blr x0
0x0000ffffa9280264  bf062128            unallocated (Unallocated)
0x0000ffffa9280268  0000ffff            udf #0xffff
0x0000ffffa928026c  00010022            unallocated (Unallocated)
0x0000ffffa9280270  00000000            udf #0x0
[DEBUG] RIP: 0x10020
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 38
[DEBUG] Blow-up Amt: 38x

After 8-bit:
0x0000ffff9cc801e0  10ffffe0    adr x0, #-0x4 (addr 0xffff9cc801dc)
0x0000ffff9cc801e4  f9005f80    str x0, [x28, #184]
0x0000ffff9cc801e8  d3401cb4    uxtb x20, w5
0x0000ffff9cc801ec  d3401cf5    uxtb x21, w7
0x0000ffff9cc801f0  394b0396    ldrb w22, [x28, #704]
0x0000ffff9cc801f4  12001294    and w20, w20, #0x1f
0x0000ffff9cc801f8  b37802d5    bfi x21, x22, #8, #1
0x0000ffff9cc801fc  b37722b5    bfi x21, x21, #9, #9
0x0000ffff9cc80200  b36e46b5    bfi x21, x21, #18, #18
0x0000ffff9cc80204  b3778eb5    bfi x21, x21, #9, #36
0x0000ffff9cc80208  9ad426b7    lsr x23, x21, x20
0x0000ffff9cc8020c  b3401ee7    bfxil x7, x23, #0, #8
0x0000ffff9cc80210  51000698    sub w24, w20, #0x1 (1)
0x0000ffff9cc80214  9ad826b5    lsr x21, x21, x24
0x0000ffff9cc80218  d34002b5    ubfx x21, x21, #0, #1
0x0000ffff9cc8021c  7100069f    cmp w20, #0x1 (1)
0x0000ffff9cc80220  9a9622b4    csel x20, x21, x22, hs
0x0000ffff9cc80224  390b0394    strb w20, [x28, #704]
0x0000ffff9cc80228  d3471ef4    ubfx x20, x23, #7, #1
0x0000ffff9cc8022c  d3461af5    ubfx x21, x23, #6, #1
0x0000ffff9cc80230  ca150294    eor x20, x20, x21
0x0000ffff9cc80234  390b2f94    strb w20, [x28, #715]
0x0000ffff9cc80238  58000040    ldr x0, pc+8 (addr 0xffff9cc80240)
0x0000ffff9cc8023c  d63f0000    blr x0
0x0000ffff9cc80240  b2a75128    unallocated (Unallocated)
0x0000ffff9cc80244  0000ffff    udf #0xffff
0x0000ffff9cc80248  00010022    unallocated (Unallocated)
0x0000ffff9cc8024c  00000000    udf #0x0
[DEBUG] RIP: 0x10020
[DEBUG] Guest Code instructions: 1
[DEBUG] Host Code instructions: 29
[DEBUG] Blow-up Amt: 29x
```
2023-07-17 19:13:23 -07:00
Ryan Houdek 4a7fa7f2bc Merge pull request #2790 from lioncash/vsib
Frontend: Handle VSIB byte
2023-07-17 18:41:03 -07:00
Lioncache 98f51c47fa Frontend: Handle VSIB byte
Extends handling of the SIB byte to also handle the AVX2 VSIB byte.

While we're in the area, we can set up the gather instruction flags as well.
2023-07-17 16:16:21 -04:00
Ryan Houdek 724a8e13bf Merge pull request #2789 from lioncash/scatter
Arm64/Emitter: Handle ST1{*} scatter store variants
2023-07-17 09:20:06 -07:00
Mai d9b52fd67d Merge pull request #2785 from Sonicadvance1/32bit_mov_bitmask
ArmEmitter: Support 32-bit bitmask moves
2023-07-17 12:13:27 -04:00
Mai d3a2795106 Merge pull request #2781 from Sonicadvance1/optimize_phminposuw
OpcodeDispatcher: Minor optimization to phminposuw
2023-07-17 12:12:25 -04:00
Mai 3cd6c2d91a Merge pull request #2779 from Sonicadvance1/optimize_shiftd
OpcodeDispatcher: Optimize 32/64-bit SH{L,R}D with extr
2023-07-17 12:11:51 -04:00
Mai b86abfbccf Merge pull request #2778 from Sonicadvance1/move_tls_signal_frontend
SignalDelegator: Moves last TLS variable to the frontend
2023-07-17 12:10:58 -04:00
Mai ee66985ae0 Merge pull request #2777 from Sonicadvance1/deadstore_elimination
IR/Passes: Fixes DeadStoreElimination pass
2023-07-17 12:09:54 -04:00
Mai daeba0625f Merge pull request #2776 from Sonicadvance1/fix_constprop_mask
ConstProp: Fix shift mask in const-prop
2023-07-17 12:08:33 -04:00
Lioncache 5e6af25194 Arm64/Emitter: Handle ST1{*} Vector + Imm scatter stores 2023-07-17 12:05:27 -04:00
Lioncache 7f4528a6b0 Arm64/Emitter: Handle ST1{*} Scalar + Vector scatter stores 2023-07-17 11:11:30 -04:00
Ryan Houdek 776b7674e4 Arm64: Only clear icache for code
Currently we're clearing icache including the data that lives on the
tail of the block. Instead only clear the code that the was emitted and
not tail data.

Additionally only disasm the code rather than all the tail data as well,
as it gets unwieldy if viewing.
2023-07-16 22:24:01 -07:00
Ryan Houdek 24cb2610a2 OpcodeDispatcher: Another FCMov minor optimization
If we are loading exactly the flags we need from the RFLAGS (ensuring we
don't load the reserved flag in bit 1) then we don't need to do a mask
on the result.

Additionally there is some bad code-motion around selects that was
causing SBFE operations to occur on constants. Ensure that we const-prop
any SBFE operations to clean this up.

This PR along with #2783 causes FMOV blow-up to go from 41 instruction
to 31 instructions.
2023-07-16 21:37:54 -07:00
Ryan Houdek 54b7a43b95 ArmEmitter: Support 32-bit bitmask moves
Noticed this when inspecting some code that was moving constant
`0x80808080` in to a register. Was using two move instructions when it
could have used a single bitmask move.

This now checks to see if a constant can be 32-bit encoded in a logical
bitmask move and uses that.
2023-07-16 18:52:34 -07:00
Mai 1b1e9e0fb5 Merge pull request #2782 from Sonicadvance1/minor_fild_opt
OpcodeDispatcher: Minor optimization to FILD
2023-07-16 04:49:50 -04:00
Ryan Houdek ead43c6a51 OpcodeDispatcher: Minor optimization to FILD
Removes one instruction from FILD, or two instructions if the CPU
supports the CSSC extension.

Going from 47 instructions to 46/45.
2023-07-16 01:28:57 -07:00
Ryan Houdek e49de77225 IR: Implements a 2's complement Integer absolute
Supports CSSC extension.
2023-07-16 01:28:57 -07:00
Ryan Houdek 8f4fe39b7d Emitter: Adds support for CNEG instruction alias
tests included.
2023-07-16 00:25:52 -07:00
Ryan Houdek f250509718 OpcodeDispatcher: Minor optimization to phminposuw
This instruction doesn't match ARM semantics very well since it returns
the position of the minimum element.

But at the very least the insert in to the final instruction can be a
bit more optimal, Converts an 5 inst eor+mov+mov+mov+mov in to 2 inst
mov+mov.

This works because `VUMinV` already zero extends the vector so the
position only needs to be inserted at the end.
2023-07-15 23:23:32 -07:00
Ryan Houdek 6179c5a13e OpcodeDispatcher: Optimize 32/64-bit SH{L,R}D with extr
32-bit and 64-bit SH{L,R}D matches behaviour of EXTR. Optimize to using
this op in that case.
This converts the lsl+lsr+orr sequence in to a single extr instruction.

16-bit still goes down the old path.

Weirdly this code manages to have a bad insert for no reason? But
unrelated since this happens in the old code as well.

```
  %4(GPRFixed3) i64 = LoadRegister #0x0, #0x20, GPR, GPRFixed, u8:Tmp:Size
  %5(GPR0) i64 = LoadRegister #0x0, #0x8, GPR, GPRFixed, u8:Tmp:Size
  %6(GPRFixed0) i64 = Extr %5(GPR0) i64, %4(GPRFixed3) i64, #0x3e
```

Not sure why the SRA fails on that second LoadRegister.
2023-07-15 22:05:39 -07:00
Ryan Houdek 9e14a83442 SignalDelegator: Moves last TLS variable to the frontend
There was one holdout variable that was in a TLS object in FEXCore. Move
it to the frontend with the rest of the TLS variables.

Allows us to remove "Frontend" TLS management to be the only TLS
management.
2023-07-15 20:28:58 -07:00
Ryan Houdek 95bfd003d2 IR/Passes: Fixes DeadStoreElimination pass
This pass is currently doing nothing in main.
Ever since we have enforced that LoadContext/StoreContext doesn't touch
GPRs and FPRs, this has only been eliminating flags.

Remove that usage of LoadContext/StoreContext and replace with their
their replacement of LoadRegister/StoreRegister for tracking GPR and FPR
accesses.

Stripped from #2700 since this is safe to merge.
2023-07-14 22:22:26 -07:00
Ryan Houdek 08ca43c3c4 ConstProp: Fix shift mask in const-prop
Noticed while looking at #2700.

Testing doesn't currently see this as a bug but will once #2700 starts
optimizing StoreRegister+LoadRegister pairs.

Doesn't fix the issues in that PR, but this is one.
2023-07-14 18:14:57 -07:00
Ryan Houdek 96b428dcaa Merge pull request #2774 from Sonicadvance1/xcb_version
CMake: Fix pkg version extraction for xcb
2023-07-14 17:18:29 -07:00
Ryan Houdek 421214e723 Merge pull request #2775 from lioncash/vecimm
Arm64: Emitter: Handle LD1{*}/LDFF1{*} Vector + Immediate encodings
2023-07-13 22:18:00 -07:00
Lioncache 1a18bbb966 Arm64: Emitter: Handle LD1{*}/LDFF1{*} Vector + Immediate encodings
While we're in the area implementing the Scalar + Vector variants,
we may as well cross off the Vector + Immediate variants and
complete all of the load variants for the regular LD1{*} loads
2023-07-14 00:50:48 -04:00
Ryan Houdek 699c3f5762 Merge pull request #2773 from lioncash/memop
Arm64/Emitter: Simplify SVEMemOperand data union
2023-07-13 18:24:44 -07:00
Ryan Houdek da0a1710c0 CMake: Fix pkg version extraction
Our regex would only ever capture a single digit, so versions that had
more than one digit per section would lose additional digits.

Fixes and moves the helper to a cmake file to be shared between
GuestLibs and HostLibs.

Uses the fix in xcb because Fedora ships an older version that doesn't
have some of FEX's newer symbols.
2023-07-13 16:01:14 -07:00
Lioncache c1205eb809 Arm64/Emitter: Mark SVEMemOperand Type enum as enum class
Now that we have helpers to make querying a little less verbose, we can mark
the enum as an enum class to get rid of implicit conversions.
2023-07-13 15:28:17 -04:00
Lioncache 31b7cd77e9 Arm64/Emitter: Simplify SVEMemOperand data union
We can just move the header out of the union, since it's present in all cases.
2023-07-13 15:28:14 -04:00
Ryan Houdek 9def04c705 Merge pull request #2747 from Sonicadvance1/non-multiarch-thunks
Thunks: Fixes thunks in non-multiarch
2023-07-13 12:22:30 -07:00
Ryan Houdek 68a2441e65 Merge pull request #2772 from lioncash/insrem
OpcodeDispatcher: Narrow use of LoadXMMRegister in StoreResult_WithOpSize
2023-07-13 12:07:05 -07:00
Ryan Houdek 4ac0dec568 Thunks: Fixes thunks in non-multiarch
Removes the @PREFIX_ARCH@ replacement  string in the thunks path.

The library prefix paths now get generated upfront and everything gets
replaced to handle the differences between multiarch distros.

Fixes Thunks on Arch and Fedora.
2023-07-13 11:59:13 -07:00
Ryan Houdek 1d7b4bb522 Merge pull request #2768 from alyssarosenzweig/fix/pf
OpcodeDispatcher: Fix and optimize PF calculation
2023-07-13 11:54:30 -07:00
Ryan Houdek 22f95e627d Merge pull request #2769 from random415/main
fix spelling errors
2023-07-13 11:48:39 -07:00
Ryan Houdek 599b64e975 Merge pull request #2771 from alyssarosenzweig/print/de-ssa
IR: Print SSA values as %123 instead of %ssa123
2023-07-13 11:48:28 -07:00
Ryan Houdek 5c27febbc2 Merge pull request #2770 from lioncash/gather
Arm64/Emitter: Handle LD1{*}/LDFF1{*} scalar + vector load variants
2023-07-13 11:35:22 -07:00
Lioncache 58c93568f6 OpcodeDispatcher: Narrow use of LoadXMMRegister in StoreResult_WithOpSize
This only needs to be loaded when a partial insert needs to be performed,
so we can narrow it's scope instead of always loading it in the AVX case.
2023-07-13 14:23:34 -04:00
Alyssa Rosenzweig 491e4e2c23 IR: Print SSA values as %123 instead of %ssa123
This is less noisy with no loss of clarity, and follows the notation
used by both LLVM IR and NIR. (So, it should be familiar.)

Change done with:

    sed -i -e 's/%ssa/%/g' $(git grep -l '%ssa')

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-13 11:52:36 -04:00
Lioncache 1ed2e24fba Arm64/Emitter: Handle LDFF1{*} scalar plus vector variants
These can use the same handling code that we introduced for the normal
LD1{*} gathers, so we may as well expose support for them.
2023-07-13 11:25:44 -04:00
Lioncache 1fbb8bd78f Arm64/Emitter: Handle LD1{*} Scalar + Vector variants 2023-07-13 11:25:41 -04:00
Elias James Howell b953433404 fix spelling errors
Fixing some minor spelling errors which should not affect functionality but improve the overall quality of documentation.
2023-07-13 11:23:59 -04:00
Alyssa Rosenzweig eae950be16 OpcodeDispatcher: Use vector ops for PF calculation
On current targets, popcount is a vector op. By using VPopcount
ourselves when calculating, we can reduce some pointless masking.
Before:

    and x22, x4, #0xff
    fmov d0, x22
    cnt v0.8b, v0.8b
    addv b0, v0.8b
    umov w22, v0.b[0]
    eor x22, x22, #0x1
    strb w22, [x28, #706]

After:

    eor x22, x4, #0x1
    fmov s4, w22
    cnt v4.16b, v4.16b
    umov w22, v4.b[0]
    strb w22, [x28, #706]

llvm-mca before:

    Iterations:        100
    Instructions:      700
    Total Cycles:      2002
    Total uOps:        700

    Dispatch Width:    2
    uOps Per Cycle:    0.35
    IPC:               0.35
    Block RThroughput: 3.5

llvm-mca after:

    Iterations:        100
    Instructions:      500
    Total Cycles:      1402
    Total uOps:        500

    Dispatch Width:    2
    uOps Per Cycle:    0.36
    IPC:               0.36
    Block RThroughput: 2.5

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-13 11:08:47 -04:00
Alyssa Rosenzweig 7e6bb04db1 OpcodeDispatcher: Extract CalculatePF
This does duplicate the _Constant(1) but it doesn't matter because it
gets inlined into the eor anyway. There is no functional change here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-13 10:05:18 -04:00
Alyssa Rosenzweig 68555546bc unittests: Add more PF coverage
FEX bugs folder of shame. This test fails on main, but passes with the
bug fix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-13 08:39:10 -04:00
Alyssa Rosenzweig 716cac35a8 OpcodeDispatcher: Fix PF calculation
We store garbage in the upper bits. That's ok, but it means we need to
mask on read for correct behaviour.

Closes #2767

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-13 08:38:46 -04:00
Ryan Houdek 9722c4c5a4 Merge pull request #2766 from alyssarosenzweig/flags/add-of
OpcodeDispatcher: Optimize ADD/ADC OF flag packing
2023-07-12 15:47:21 -07:00
Alyssa Rosenzweig e8c0e19afc OpcodeDispatcher: "Calculcate" -> "Calculate"
Typofix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-12 18:07:04 -04:00
Alyssa Rosenzweig c559fec959 OpcodeDispatcher: Optimize ADD/ADC OF flag packing
We can fold the Not into the And. This requires flipping the arguments
to Andn, but we do not flip the order of the assignments since that
requires an extra register in a test I'm looking at.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-12 18:06:36 -04:00
Alyssa Rosenzweig 8d2fabe705 OpcodeDispatcher: Deduplicate ADD/ADC OF generation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-07-12 18:06:36 -04:00
Mai 6455c4817a Merge pull request #2764 from Sonicadvance1/remove_unused_tls
FEXCore: Removes unused TLS variable
2023-07-12 17:25:08 -04:00
Ryan Houdek 5dbd1b8dc2 FEXCore: Removes unused TLS variable
Not sure why this still existed.
2023-07-12 13:05:47 -07:00
Mai 7765bbc7b8 Merge pull request #2762 from Sonicadvance1/FEXRootFSFetcherPercent
FEXRootFSFetcher: Make verification percent easier to read
2023-07-12 15:34:07 -04:00
Mai 2283c73fae Merge pull request #2742 from Sonicadvance1/fix_win32
FEXCore: Fixes WIN32 compiling again
2023-07-12 15:33:46 -04:00
Ryan Houdek 5fef0c29aa FEXCore: Rename Telemetry helper function GetObject
WIN32 has a define already called `GetObject` and will cause our
symbol to have an A appended to it and break linking.

Just rename it to `GetTelemetryValue`
2023-07-12 11:53:13 -07:00
Ryan Houdek d387c46aab FEXCore: Fixes WIN32 compiling again
Mostly a quick bandage while I'm setting getting ready to setup the
runners to test this for us.
2023-07-12 11:53:13 -07:00
Ryan Houdek 3bb7f9d6b5 FEXRootFSFetcher: Make verification percent easier to read
Multiply it by 100 to actually show as a percentage, only show two
digits past the decimal, and update every second to be more responsive.
2023-07-11 20:03:06 -07:00
Ryan Houdek 70d54122b2 Merge pull request #2761 from bylaws/main
FHU: Fix WIN32 getcpu implementation
2023-07-11 17:37:06 -07:00
Billy Laws 4e266cf9fd FHU: Fix WIN32 getcpu implementation
Both parameters to getcpu are nullable and FEX relies upon this for
CPUID.
2023-07-12 00:02:49 +01:00
Mai ddd6dbfdcc Merge pull request #2759 from Sonicadvance1/redundant_bfe_flags
OpcodeDispatcher: Remove spurious bfe with flag storing
2023-07-10 22:19:21 -04:00
Mai 7f2557e322 Merge pull request #2757 from Sonicadvance1/optimize_movss_reg
OpcodeDispatcher: Optimize MOVSS to register
2023-07-10 21:22:47 -04:00
Mai 810c7d926c Merge pull request #2758 from Sonicadvance1/optimize_tso_vector_loadstores
IR: Optimize vector TSO loadstore address calculation
2023-07-10 21:21:46 -04:00
Ryan Houdek 04c325661c OpcodeDispatcher: Remove spurious bfe with flag storing
Noticed during introspection that we were generating zero constants
redundantly. Bunch of single cycle hits or zero-register renames.

Every time a `SetRFLAG` helper was called, it was /always/ doing a BFE
on everything passed in to extract the lowest bit. In nearly all cases
the data getting passed in is already only the lowest bit.

Instead, stop the helper from doing this BFE, and ensure the
OpcodeDispatcher does BFE in the couple of cases it still needs to do.

As I was skimming through all these to ensure BFE isn't necessary, I did
notice that some of the BCD instructions are wrong or questionable. So I
left a comment on those so we can come back to it.
2023-07-10 18:03:23 -07:00
Mai 0121e858f2 Merge pull request #2756 from Sonicadvance1/movss_optimize
OpcodeDispatcher: Optimize MOVSS to memory destination
2023-07-10 18:48:13 -04:00
Mai 2c3361be9e Merge pull request #2755 from Sonicadvance1/stop_installing_fmt
CMake: Stop installing fmt
2023-07-10 18:47:30 -04:00
Ryan Houdek 2d800b2627 IR: Optimize vector TSO loadstore address calculation
These address calculations were failing to understand that they can be
optimized. When TSO emulation is disabled these were fine, but with TSO
we were eating one more instruction.

Before:
```
add x20, x12, #0x4 (4)
dmb ish
ldr s16, [x20]
dmb ish
```

After:
```
dmb ish
ldr s16, [x12, #4]
dmb ish
```

Also left a note that once LRCPC3 is supported in hardware that we can do a similar optimization there.
2023-07-10 15:21:46 -07:00
Ryan Houdek 55ed3e0549 OpcodeDispatcher: Optimize MOVSS to register
Easily fixed. Found through inspection.

Before:
```
eor v0.16b, v0.16b, v0.16b
mov v0.s[0], v17.s[0]
mov v4.16b, v0.16b
mov v16.s[0], v4.s[0]
```

After:
```
mov v16.s[0], v17.s[0]
```
2023-07-10 14:36:27 -07:00
Ryan Houdek 55d084ebb0 OpcodeDispatcher: Optimize MOVSS to memory destination
Easy fixed. Found through inspection.

Before:
```
eor v0.16b, v0.16b, v0.16b
mov v0.s[0], v16.s[0]
mov v4.16b, v0.16b
str s4, [x11]
```

After:
```
str s16, [x11]
```
2023-07-10 14:25:01 -07:00
Ryan Houdek 457dc5dd90 CMake: Stop installing fmt
Fixes #2751

Luckily fmt provides an option to disable this.
2023-07-10 12:26:39 -07:00
Mai 98eda5e163 Merge pull request #2749 from Sonicadvance1/optimize_away_redundant_masks
OpcodeDispatcher: Optimize some shifts size masking
2023-07-10 08:08:57 -04:00
Ryan Houdek 592935790e Merge pull request #2750 from Sonicadvance1/fix_pcmpestri
OpcodeDispatcher: Fixes bug with pcmpestri
2023-07-08 18:47:38 -07:00
Ryan Houdek 92d0344d6a OpcodeDispatcher: Fixes bug with pcmpestri
When this instruction returns the index in to the ecx register, this is
defined as a 32-bit result. This means it actually gets zero-extended to
the full 64-bit GPR size on 64-bit processes.
Previously FEX was doing a 32-bit insert which leaves garbage data in
the upper 32-bits of the RCX register.

Adds a unit test to ensure the result is zero extended.
Fixes running Java games under FEX now that SSE4.2 is exposed.
2023-07-08 18:08:47 -07:00
Ryan Houdek 9327435f97 OpcodeDispatcher: Optimize some shifts size masking
Inspired from #2561, these shifts  don't need to be masked if we know
their operating size up front.

Causes a handful of these to become more optimal.
2023-07-08 16:41:15 -07:00
Mai 573f339647 Merge pull request #2748 from Sonicadvance1/fix_missing_header
unittests: Adds missing header
2023-07-08 18:11:33 -04:00
Ryan Houdek c1f18951ab unittests: Adds missing header
Newer libstdc++ moved an internal header include and now this failed to
compile.
2023-07-08 14:50:10 -07:00
Mai 8a4bfba47c Merge pull request #2745 from Sonicadvance1/optimize_fcmov
OpcodeDispatcher: Optimize GetPackedRFLAG
2023-07-07 22:29:52 -04:00
Mai 69ea03f0eb Merge pull request #2746 from Sonicadvance1/optimize_maskmov
OpcodeDispatcher: Optimize MASKMOVDQU and MASKMOVQ
2023-07-07 22:29:37 -04:00
Mai 462feec2a6 Merge pull request #2743 from Sonicadvance1/minor_cleanup
FEXCore: Minor cleanup
2023-07-07 22:25:53 -04:00
Ryan Houdek 15f5fe658b OpcodeDispatcher: Optimize MASKMOVDQU and MASKMOVQ
This previous implementation was particularly gnarly. Because these
instructions are both weackly ordered and have implementation dependent
exception and trap behaviour these can actually be fairly conveniently
converted over to a load + cmlt + bsl + str instruction.

For the XMM variant this reduces code blowup from 80x to 15x!
For the MMX variant this reduces code blowup from 46x to 17x!

Both of these improvements are significant wins! There's still some
minor improvement that could be done with bsl that requires some
redundant moves, but since we don't have constraint support for this we
still eat two additional instructions

Before:
```asm
0x0000ffff7b800718  10ffffe0    adr x0, #-0x4 (addr 0xffff7b800714)
0x0000ffff7b80071c  f9005f80    str x0, [x28, #184]
0x0000ffff7b800720  4eb11e24    mov v4.16b, v17.16b
0x0000ffff7b800724  4eb01e05    mov v5.16b, v16.16b
0x0000ffff7b800728  aa0b03f4    mov x20, x11
0x0000ffff7b80072c  4e083c95    mov x21, v4.d[0]
0x0000ffff7b800730  4e083cb6    mov x22, v5.d[0]
0x0000ffff7b800734  d3471eb7    ubfx x23, x21, #7, #1
0x0000ffff7b800738  b4000077    cbz x23, #+0xc (addr 0xffff7b800744)
0x0000ffff7b80073c  d3401ed7    uxtb x23, w22
0x0000ffff7b800740  39000297    strb w23, [x20]
0x0000ffff7b800744  d34f3eb7    ubfx x23, x21, #15, #1
0x0000ffff7b800748  b4000077    cbz x23, #+0xc (addr 0xffff7b800754)
0x0000ffff7b80074c  d3483ed7    ubfx x23, x22, #8, #8
0x0000ffff7b800750  39000697    strb w23, [x20, #1]
0x0000ffff7b800754  d3575eb7    ubfx x23, x21, #23, #1
0x0000ffff7b800758  b4000077    cbz x23, #+0xc (addr 0xffff7b800764)
0x0000ffff7b80075c  d3505ed7    ubfx x23, x22, #16, #8
0x0000ffff7b800760  39000a97    strb w23, [x20, #2]
0x0000ffff7b800764  d35f7eb7    ubfx x23, x21, #31, #1
0x0000ffff7b800768  b4000077    cbz x23, #+0xc (addr 0xffff7b800774)
0x0000ffff7b80076c  d3587ed7    ubfx x23, x22, #24, #8
0x0000ffff7b800770  39000e97    strb w23, [x20, #3]
0x0000ffff7b800774  d3679eb7    ubfx x23, x21, #39, #1
0x0000ffff7b800778  b4000077    cbz x23, #+0xc (addr 0xffff7b800784)
0x0000ffff7b80077c  d3609ed7    ubfx x23, x22, #32, #8
0x0000ffff7b800780  39001297    strb w23, [x20, #4]
0x0000ffff7b800784  d36fbeb7    ubfx x23, x21, #47, #1
0x0000ffff7b800788  b4000077    cbz x23, #+0xc (addr 0xffff7b800794)
0x0000ffff7b80078c  d368bed7    ubfx x23, x22, #40, #8
0x0000ffff7b800790  39001697    strb w23, [x20, #5]
0x0000ffff7b800794  d377deb7    ubfx x23, x21, #55, #1
0x0000ffff7b800798  b4000077    cbz x23, #+0xc (addr 0xffff7b8007a4)
0x0000ffff7b80079c  d370ded7    ubfx x23, x22, #48, #8
0x0000ffff7b8007a0  39001a97    strb w23, [x20, #6]
0x0000ffff7b8007a4  d37ffeb5    lsr x21, x21, #63
0x0000ffff7b8007a8  b4000075    cbz x21, #+0xc (addr 0xffff7b8007b4)
0x0000ffff7b8007ac  d378fed5    lsr x21, x22, #56
0x0000ffff7b8007b0  39001e95    strb w21, [x20, #7]
0x0000ffff7b8007b4  4e183c95    mov x21, v4.d[1]
0x0000ffff7b8007b8  4e183cb6    mov x22, v5.d[1]
0x0000ffff7b8007bc  d3471eb7    ubfx x23, x21, #7, #1
0x0000ffff7b8007c0  b4000077    cbz x23, #+0xc (addr 0xffff7b8007cc)
0x0000ffff7b8007c4  d3401ed7    uxtb x23, w22
0x0000ffff7b8007c8  39002297    strb w23, [x20, #8]
0x0000ffff7b8007cc  d34f3eb7    ubfx x23, x21, #15, #1
0x0000ffff7b8007d0  b4000077    cbz x23, #+0xc (addr 0xffff7b8007dc)
0x0000ffff7b8007d4  d3483ed7    ubfx x23, x22, #8, #8
0x0000ffff7b8007d8  39002697    strb w23, [x20, #9]
0x0000ffff7b8007dc  d3575eb7    ubfx x23, x21, #23, #1
0x0000ffff7b8007e0  b4000077    cbz x23, #+0xc (addr 0xffff7b8007ec)
0x0000ffff7b8007e4  d3505ed7    ubfx x23, x22, #16, #8
0x0000ffff7b8007e8  39002a97    strb w23, [x20, #10]
0x0000ffff7b8007ec  d35f7eb7    ubfx x23, x21, #31, #1
0x0000ffff7b8007f0  b4000077    cbz x23, #+0xc (addr 0xffff7b8007fc)
0x0000ffff7b8007f4  d3587ed7    ubfx x23, x22, #24, #8
0x0000ffff7b8007f8  39002e97    strb w23, [x20, #11]
0x0000ffff7b8007fc  d3679eb7    ubfx x23, x21, #39, #1
0x0000ffff7b800800  b4000077    cbz x23, #+0xc (addr 0xffff7b80080c)
0x0000ffff7b800804  d3609ed7    ubfx x23, x22, #32, #8
0x0000ffff7b800808  39003297    strb w23, [x20, #12]
0x0000ffff7b80080c  d36fbeb7    ubfx x23, x21, #47, #1
0x0000ffff7b800810  b4000077    cbz x23, #+0xc (addr 0xffff7b80081c)
0x0000ffff7b800814  d368bed7    ubfx x23, x22, #40, #8
0x0000ffff7b800818  39003697    strb w23, [x20, #13]
0x0000ffff7b80081c  d377deb7    ubfx x23, x21, #55, #1
0x0000ffff7b800820  b4000077    cbz x23, #+0xc (addr 0xffff7b80082c)
0x0000ffff7b800824  d370ded7    ubfx x23, x22, #48, #8
0x0000ffff7b800828  39003a97    strb w23, [x20, #14]
0x0000ffff7b80082c  d37ffeb5    lsr x21, x21, #63
0x0000ffff7b800830  b4000075    cbz x21, #+0xc (addr 0xffff7b80083c)
0x0000ffff7b800834  d378fed5    lsr x21, x22, #56
0x0000ffff7b800838  39003e95    strb w21, [x20, #15]
0x0000ffff7b80083c  58000040    ldr x0, pc+8 (addr 0xffff7b800844)
0x0000ffff7b800840  d63f0000    blr x0
```

After:
```asm
0x0000ffff7ac00718  10ffffe0            adr x0, #-0x4 (addr 0xffff7ac00714)
0x0000ffff7ac0071c  f9005f80            str x0, [x28, #184]
0x0000ffff7ac00720  4e20aa24            cmlt v4.16b, v17.16b, #0
0x0000ffff7ac00724  3dc00165            ldr q5, [x11]
0x0000ffff7ac00728  4ea41c80            mov v0.16b, v4.16b
0x0000ffff7ac0072c  6e651e00            bsl v0.16b, v16.16b, v5.16b
0x0000ffff7ac00730  4ea01c04            mov v4.16b, v0.16b
0x0000ffff7ac00734  3d800164            str q4, [x11]
0x0000ffff7ac00738  58000040            ldr x0, pc+8 (addr 0xffff7ac00740)
0x0000ffff7ac0073c  d63f0000            blr x0
```
2023-07-07 18:37:17 -07:00
Ryan Houdek 052aa4317b OpcodeDispatcher: Optimize GetPackedRFLAG
Only return the particular flags that are being requested in the moment
since compacting them all when requested is fairly slow.

x87 fcmov in particular was requesting all the flags when it only needs
a couple.
This reduces a `fcmovb` instruction count blowup from 103x to 38x. Still
more room to go but this one stood out as being particularly bad.

Old:
```asm
0x0000000265a002bc  10ffffe0    adr x0, #-0x4 (addr 0x265a002b8)
0x0000000265a002c0  f9005f80    str x0, [x28, #184]
0x0000000265a002c4  d2800014    mov x20, #0x0
0x0000000265a002c8  d2800035    mov x21, #0x1
0x0000000265a002cc  d2800056    mov x22, #0x2
0x0000000265a002d0  394b0397    ldrb w23, [x28, #704]
0x0000000265a002d4  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a002d8  aa1702d6    orr x22, x22, x23
0x0000000265a002dc  394b0b97    ldrb w23, [x28, #706]
0x0000000265a002e0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a002e4  531e76f7    lsl w23, w23, #2
0x0000000265a002e8  aa1702d6    orr x22, x22, x23
0x0000000265a002ec  394b1397    ldrb w23, [x28, #708]
0x0000000265a002f0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a002f4  531c6ef7    lsl w23, w23, #4
0x0000000265a002f8  aa1702d6    orr x22, x22, x23
0x0000000265a002fc  394b1b97    ldrb w23, [x28, #710]
0x0000000265a00300  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00304  531a66f7    lsl w23, w23, #6
0x0000000265a00308  aa1702d6    orr x22, x22, x23
0x0000000265a0030c  394b1f97    ldrb w23, [x28, #711]
0x0000000265a00310  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00314  531962f7    lsl w23, w23, #7
0x0000000265a00318  aa1702d6    orr x22, x22, x23
0x0000000265a0031c  394b2397    ldrb w23, [x28, #712]
0x0000000265a00320  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00324  53185ef7    lsl w23, w23, #8
0x0000000265a00328  aa1702d6    orr x22, x22, x23
0x0000000265a0032c  394b2797    ldrb w23, [x28, #713]
0x0000000265a00330  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00334  53175af7    lsl w23, w23, #9
0x0000000265a00338  aa1702d6    orr x22, x22, x23
0x0000000265a0033c  394b2b97    ldrb w23, [x28, #714]
0x0000000265a00340  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00344  531656f7    lsl w23, w23, #10
0x0000000265a00348  aa1702d6    orr x22, x22, x23
0x0000000265a0034c  394b2f97    ldrb w23, [x28, #715]
0x0000000265a00350  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00354  531552f7    lsl w23, w23, #11
0x0000000265a00358  aa1702d6    orr x22, x22, x23
0x0000000265a0035c  394b3397    ldrb w23, [x28, #716]
0x0000000265a00360  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00364  53144ef7    lsl w23, w23, #12
0x0000000265a00368  aa1702d6    orr x22, x22, x23
0x0000000265a0036c  394b3b97    ldrb w23, [x28, #718]
0x0000000265a00370  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00374  531246f7    lsl w23, w23, #14
0x0000000265a00378  aa1702d6    orr x22, x22, x23
0x0000000265a0037c  394b4397    ldrb w23, [x28, #720]
0x0000000265a00380  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00384  53103ef7    lsl w23, w23, #16
0x0000000265a00388  aa1702d6    orr x22, x22, x23
0x0000000265a0038c  394b4797    ldrb w23, [x28, #721]
0x0000000265a00390  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a00394  530f3af7    lsl w23, w23, #17
0x0000000265a00398  aa1702d6    orr x22, x22, x23
0x0000000265a0039c  394b4b97    ldrb w23, [x28, #722]
0x0000000265a003a0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a003a4  530e36f7    lsl w23, w23, #18
0x0000000265a003a8  aa1702d6    orr x22, x22, x23
0x0000000265a003ac  394b4f97    ldrb w23, [x28, #723]
0x0000000265a003b0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a003b4  530d32f7    lsl w23, w23, #19
0x0000000265a003b8  aa1702d6    orr x22, x22, x23
0x0000000265a003bc  394b5397    ldrb w23, [x28, #724]
0x0000000265a003c0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a003c4  530c2ef7    lsl w23, w23, #20
0x0000000265a003c8  aa1702d6    orr x22, x22, x23
0x0000000265a003cc  394b5797    ldrb w23, [x28, #725]
0x0000000265a003d0  d3407ef7    ubfx x23, x23, #0, #32
0x0000000265a003d4  530b2af7    lsl w23, w23, #21
0x0000000265a003d8  aa1702d6    orr x22, x22, x23
0x0000000265a003dc  924002d6    and x22, x22, #0x1
0x0000000265a003e0  93400294    sbfx x20, x20, #0, #1
0x0000000265a003e4  934002b5    sbfx x21, x21, #0, #1
0x0000000265a003e8  f10002df    cmp x22, #0x0 (0)
0x0000000265a003ec  9a950294    csel x20, x20, x21, eq
0x0000000265a003f0  4e080e84    dup v4.2d, x20
0x0000000265a003f4  394baf94    ldrb w20, [x28, #747]
0x0000000265a003f8  91000695    add x21, x20, #0x1 (1)
0x0000000265a003fc  92400ab5    and x21, x21, #0x7
0x0000000265a00400  d2800200    mov x0, #0x10
0x0000000265a00404  9b007e80    mul x0, x20, x0
0x0000000265a00408  8b000380    add x0, x28, x0
0x0000000265a0040c  3dc0bc05    ldr q5, [x0, #752]
0x0000000265a00410  d2800200    mov x0, #0x10
0x0000000265a00414  9b007ea0    mul x0, x21, x0
0x0000000265a00418  8b000380    add x0, x28, x0
0x0000000265a0041c  3dc0bc06    ldr q6, [x0, #752]
0x0000000265a00420  4ea41c80    mov v0.16b, v4.16b
0x0000000265a00424  6e651cc0    bsl v0.16b, v6.16b, v5.16b
0x0000000265a00428  4ea01c04    mov v4.16b, v0.16b
0x0000000265a0042c  d2800200    mov x0, #0x10
0x0000000265a00430  9b007e80    mul x0, x20, x0
0x0000000265a00434  8b000380    add x0, x28, x0
0x0000000265a00438  3d80bc04    str q4, [x0, #752]
0x0000000265a0043c  58000040    ldr x0, pc+8 (addr 0x265a00444)
0x0000000265a00440  d63f0000    blr x0
```

New:
```asm
0x0000000265a002bc  10ffffe0    adr x0, #-0x4 (addr 0x265a002b8)
0x0000000265a002c0  f9005f80    str x0, [x28, #184]
0x0000000265a002c4  d2800014    mov x20, #0x0
0x0000000265a002c8  d2800035    mov x21, #0x1
0x0000000265a002cc  d2800056    mov x22, #0x2
0x0000000265a002d0  394b0397    ldrb w23, [x28, #704]
0x0000000265a002d4  330002f6    bfxil w22, w23, #0, #1
0x0000000265a002d8  924002d6    and x22, x22, #0x1
0x0000000265a002dc  93400294    sbfx x20, x20, #0, #1
0x0000000265a002e0  934002b5    sbfx x21, x21, #0, #1
0x0000000265a002e4  f10002df    cmp x22, #0x0 (0)
0x0000000265a002e8  9a950294    csel x20, x20, x21, eq
0x0000000265a002ec  4e080e84    dup v4.2d, x20
0x0000000265a002f0  394baf94    ldrb w20, [x28, #747]
0x0000000265a002f4  91000695    add x21, x20, #0x1 (1)
0x0000000265a002f8  92400ab5    and x21, x21, #0x7
0x0000000265a002fc  d2800200    mov x0, #0x10
0x0000000265a00300  9b007e80    mul x0, x20, x0
0x0000000265a00304  8b000380    add x0, x28, x0
0x0000000265a00308  3dc0bc05    ldr q5, [x0, #752]
0x0000000265a0030c  d2800200    mov x0, #0x10
0x0000000265a00310  9b007ea0    mul x0, x21, x0
0x0000000265a00314  8b000380    add x0, x28, x0
0x0000000265a00318  3dc0bc06    ldr q6, [x0, #752]
0x0000000265a0031c  4ea41c80    mov v0.16b, v4.16b
0x0000000265a00320  6e651cc0    bsl v0.16b, v6.16b, v5.16b
0x0000000265a00324  4ea01c04    mov v4.16b, v0.16b
0x0000000265a00328  d2800200    mov x0, #0x10
0x0000000265a0032c  9b007e80    mul x0, x20, x0
0x0000000265a00330  8b000380    add x0, x28, x0
0x0000000265a00334  3d80bc04    str q4, [x0, #752]
0x0000000265a00338  58000040    ldr x0, pc+8 (addr 0x265a00340)
0x0000000265a0033c  d63f0000    blr x0
```
2023-07-07 17:01:59 -07:00
Ryan Houdek debcb0e047 Arm64: Optimize BFI in the case that Dst == srcDst
ARM64 BFI doesn't allow you to encode two source registers here to match
our SSA semantics. Also since we don't support RA constraints to ensure
that these match, just do the optimal case in the backend.

Leave a comment for future RA contraint excavators to make this more
optimal
2023-07-07 16:43:41 -07:00
Ryan Houdek baf04b6a41 FEXCore: Minor cleanup
This isn't required anymore since we are exposing the virtual class
directly.
2023-07-07 15:06:14 -07:00
444 changed files with 80947 additions and 8464 deletions

No files matched your search

+12
View File
@@ -172,6 +172,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXCore APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fexcore_apitests
- name: FEXCore APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXCoreAPITests.log || true
- name: ARMEmitter tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -254,6 +265,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+12
View File
@@ -157,6 +157,17 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_APITests.log || true
- name: FEXCore APITest tests
working-directory: ${{runner.workspace}}/build
shell: bash
run: cmake --build . --config $BUILD_TYPE --target fexcore_apitests
- name: FEXCore APITest Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_FEXCoreAPITests.log || true
- name: FEXLinuxTests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -189,6 +200,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+109
View File
@@ -0,0 +1,109 @@
name: Instruction Count CI run
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=False -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_DISABLETELEMETRY: 1
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE --target CodeSizeValidation instcountci_test_files
- name: Instruction Count Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target instcountci_tests
- name: Instruction Count Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_InstCountCI.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+97
View File
@@ -0,0 +1,97 @@
name: Mingw build
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
BUILD_TYPE: Debug
FEX_ENABLEAVX: 1
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, x64, mingw], [self-hosted, ARM64, mingw]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name: Set CC x86
if: matrix.arch[1] == 'x64'
run: |
echo "MINGW_TRIPLE=x86_64-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64
if: matrix.arch[1] == 'ARM64'
run: |
echo "MINGW_TRIPLE=aarch64-w64-mingw32" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=False -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+17 -1
View File
@@ -65,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -85,6 +85,21 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: ASM Tests 128-bit
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disableavx"
FEX_FORCESVEWIDTH: "128"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test 128-bit Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM128bit.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
@@ -112,6 +127,7 @@ jobs:
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
+30
View File
@@ -0,0 +1,30 @@
# Extracts a version from the passed in version string in the form of "<Major>.<Minor>.<Patch>".
# If a part of the version is missing then it gets set as zero.
# Version variables returned in:
# ${Package}_VERSION_MAJOR
# ${Package}_VERSION_MINOR
# ${Package}_VERSION_PATCH
function(version_to_variables VERSION _Package)
string(REPLACE "." ";" VERSION_LIST "${VERSION}")
list (LENGTH VERSION_LIST VERSION_LEN)
if (${VERSION_LEN} GREATER 0)
list(GET VERSION_LIST 0 VERSION_MAJOR)
set(${_Package}_VERSION_MAJOR ${VERSION_MAJOR} PARENT_SCOPE)
else()
set(${_Package}_VERSION_MAJOR 0 PARENT_SCOPE)
endif()
if (${VERSION_LEN} GREATER 1)
list(GET VERSION_LIST 1 VERSION_MINOR)
set(${_Package}_VERSION_MINOR ${VERSION_MINOR} PARENT_SCOPE)
else()
set(${_Package}_VERSION_MINOR 0 PARENT_SCOPE)
endif()
if (${VERSION_LEN} GREATER 2)
list(GET VERSION_LIST 2 VERSION_PATCH)
set(${_Package}_VERSION_PATCH ${VERSION_PATCH} PARENT_SCOPE)
else()
set(${_Package}_VERSION_PATCH 0 PARENT_SCOPE)
endif()
endfunction()
+8 -11
View File
@@ -13,8 +13,7 @@ option(ENABLE_CLANG_FORMAT "Run clang format over the source" FALSE)
option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
option(ENABLE_LLD "Enable linking with lld" FALSE)
option(ENABLE_MOLD "Enable linking with mold" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
@@ -157,13 +156,9 @@ endif()
set (PTHREAD_LIB pthread)
if (ENABLE_LLD AND ENABLE_MOLD)
message (FATAL_ERROR "Cannot enable both lld and mold")
elseif (ENABLE_LLD)
set (LD_OVERRIDE "-fuse-ld=lld")
add_link_options(${LD_OVERRIDE})
elseif (ENABLE_MOLD)
add_link_options("-fuse-ld=mold")
if (USE_LINKER)
message(STATUS "Overriding linker to: ${USE_LINKER}")
add_link_options("-fuse-ld=${USE_LINKER}")
endif()
if (ENABLE_LIBCXX)
@@ -241,7 +236,7 @@ if (BUILD_TESTS)
endif()
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
include_directories(SYSTEM External/vixl/src/)
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
@@ -267,6 +262,8 @@ if (BUILD_TESTS)
include(Catch)
endif()
# Disable fmt install
set(FMT_INSTALL OFF)
add_subdirectory(External/fmt/)
add_subdirectory(External/imgui/)
@@ -430,7 +427,7 @@ if (BUILD_TESTS)
endif()
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(External/FEXCore)
add_subdirectory(FEXCore/)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
+66 -63
View File
@@ -6,10 +6,10 @@
"X11"
],
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1.2.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGL.so.1.7.0"
"@PREFIX_LIB@/libGL.so",
"@PREFIX_LIB@/libGL.so.1",
"@PREFIX_LIB@/libGL.so.1.2.0",
"@PREFIX_LIB@/libGL.so.1.7.0"
]
},
"GLESv2": {
@@ -18,17 +18,17 @@
"X11"
],
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libGLESv2.so.2.0.0"
"@PREFIX_LIB@/libGLESv2.so",
"@PREFIX_LIB@/libGLESv2.so.2",
"@PREFIX_LIB@/libGLESv2.so.2.0.0"
]
},
"X11": {
"Library": "libX11-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so.6",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libX11.so.6.4.0"
"@PREFIX_LIB@/libX11.so",
"@PREFIX_LIB@/libX11.so.6",
"@PREFIX_LIB@/libX11.so.6.4.0"
]
},
"Vulkan": {
@@ -37,8 +37,8 @@
"xcb"
],
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libvulkan.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libvulkan.so.1",
"@PREFIX_LIB@/libvulkan.so",
"@PREFIX_LIB@/libvulkan.so.1",
"@HOME@/.local/share/Steam/ubuntu12_32/steam-runtime/pinned_libs_64/libvulkan.so.1"
],
"Comment": [
@@ -46,139 +46,142 @@
]
},
"xcb": {
"Depends": [
"X11"
],
"Library": "libxcb-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb.so.1.1.0"
"@PREFIX_LIB@/libxcb.so",
"@PREFIX_LIB@/libxcb.so.1",
"@PREFIX_LIB@/libxcb.so.1.1.0"
]
},
"xcb-dri2": {
"Library": "libxcb-dri2-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri2.so.0.0.0"
"@PREFIX_LIB@/libxcb-dri2.so",
"@PREFIX_LIB@/libxcb-dri2.so.0",
"@PREFIX_LIB@/libxcb-dri2.so.0.0.0"
]
},
"xcb-dri3": {
"Library": "libxcb-dri3-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-dri3.so.0.0.0"
"@PREFIX_LIB@/libxcb-dri3.so",
"@PREFIX_LIB@/libxcb-dri3.so.0",
"@PREFIX_LIB@/libxcb-dri3.so.0.0.0"
]
},
"xcb-xfixes": {
"Library": "libxcb-xfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-xfixes.so.0.0.0"
"@PREFIX_LIB@/libxcb-xfixes.so",
"@PREFIX_LIB@/libxcb-xfixes.so.0",
"@PREFIX_LIB@/libxcb-xfixes.so.0.0.0"
]
},
"xcb-shm": {
"Library": "libxcb-shm-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-shm.so.0.0.0"
"@PREFIX_LIB@/libxcb-shm.so",
"@PREFIX_LIB@/libxcb-shm.so.0",
"@PREFIX_LIB@/libxcb-shm.so.0.0.0"
]
},
"xcb-sync": {
"Library": "libxcb-sync-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-sync.so.1.0.0"
"@PREFIX_LIB@/libxcb-sync.so",
"@PREFIX_LIB@/libxcb-sync.so.1",
"@PREFIX_LIB@/libxcb-sync.so.1.0.0"
]
},
"xcb-randr": {
"Library": "libxcb-randr-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-randr.so.0.1.0"
"@PREFIX_LIB@/libxcb-randr.so",
"@PREFIX_LIB@/libxcb-randr.so.0",
"@PREFIX_LIB@/libxcb-randr.so.0.1.0"
]
},
"xcb-present": {
"Library": "libxcb-present-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-present.so.0.0.0"
"@PREFIX_LIB@/libxcb-present.so",
"@PREFIX_LIB@/libxcb-present.so.0",
"@PREFIX_LIB@/libxcb-present.so.0.0.0"
]
},
"xcb-glx": {
"Library": "libxcb-glx-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxcb-glx.so.0.0.0"
"@PREFIX_LIB@/libxcb-glx.so",
"@PREFIX_LIB@/libxcb-glx.so.0",
"@PREFIX_LIB@/libxcb-glx.so.0.0.0"
]
},
"xshmfence": {
"Library": "libxshmfence-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libxshmfence.so.1.0.0"
"@PREFIX_LIB@/libxshmfence.so",
"@PREFIX_LIB@/libxshmfence.so.1",
"@PREFIX_LIB@/libxshmfence.so.1.0.0"
]
},
"drm": {
"Library": "libdrm-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libdrm.so.2.4.0"
"@PREFIX_LIB@/libdrm.so",
"@PREFIX_LIB@/libdrm.so.2",
"@PREFIX_LIB@/libdrm.so.2.4.0"
]
},
"asound": {
"Library": "libasound-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so.2",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libasound.so.2.0.0"
"@PREFIX_LIB@/libasound.so",
"@PREFIX_LIB@/libasound.so.2",
"@PREFIX_LIB@/libasound.so.2.0.0"
]
},
"Xrender": {
"Library": "libXrender-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXrender.so.1.3.0"
"@PREFIX_LIB@/libXrender.so",
"@PREFIX_LIB@/libXrender.so.1",
"@PREFIX_LIB@/libXrender.so.1.3.0"
]
},
"Xext": {
"Library": "libXext-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so.6",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXext.so.6.4.0"
"@PREFIX_LIB@/libXext.so",
"@PREFIX_LIB@/libXext.so.6",
"@PREFIX_LIB@/libXext.so.6.4.0"
]
},
"Xfixes": {
"Library": "libXfixes-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so.3",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libXfixes.so.3.1.0"
"@PREFIX_LIB@/libXfixes.so",
"@PREFIX_LIB@/libXfixes.so.3",
"@PREFIX_LIB@/libXfixes.so.3.1.0"
]
},
"OpenCL": {
"Library" : "libOpenCL-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so.1",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libOpenCL.so.1.0.0"
"@PREFIX_LIB@/libOpenCL.so",
"@PREFIX_LIB@/libOpenCL.so.1",
"@PREFIX_LIB@/libOpenCL.so.1.0.0"
]
},
"WaylandClient": {
"Library" : "libwayland-client-guest.so",
"Overlay": [
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so.0",
"@PREFIX_LIB@/@PREFIX_ARCH@-linux-gnu/libwayland-client.so.0.20.0"
"@PREFIX_LIB@/libwayland-client.so",
"@PREFIX_LIB@/libwayland-client.so.0",
"@PREFIX_LIB@/libwayland-client.so.0.20.0"
]
},
"":{}
+1
View File
@@ -6,3 +6,4 @@ mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xf
credentials yes
fix_binary yes
preserve yes
expose_interpreter optional
+1
View File
@@ -6,3 +6,4 @@ mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xf
credentials yes
fix_binary yes
preserve yes
expose_interpreter optional
+1 -1
-88
View File
@@ -1,88 +0,0 @@
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
namespace FEXCore {
namespace CPU {
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
}
return CurrentCodeBuffer;
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto &Buffer: CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
}
}
return false;
}
}
}
-23
View File
@@ -1,23 +0,0 @@
#pragma once
namespace FEXCore {
class CodeLoader;
}
namespace FEXCore::Context {
class ContextImpl;
}
namespace FEXCore::CPU {
/**
* @brief Create the CPU core backend for the context passed in
*
* @param CTX
*
* @return true if core was able to be create
*/
bool CreateCPUCore(FEXCore::Context::ContextImpl *CTX);
bool LoadCode(FEXCore::Context::ContextImpl *CTX, FEXCore::CodeLoader *Loader);
}
@@ -1,184 +0,0 @@
#pragma once
#include <FEXCore/IR/IR.h>
#define GD *GetDest<uint64_t*>(Data->SSAData, Node)
#define GDP GetDest<void*>(Data->SSAData, Node)
#define DO_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(GDP); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
*Dst_d = func(*Src1_d, *Src2_d); \
break; \
}
#define DO_SCALAR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
Dst_d[0] = func(Src1_d[0], Src2_d[0]); \
break; \
}
#define DO_VECTOR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_PAIR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i*2], Src1_d[i*2 + 1]); \
Dst_d[i+Elements] = func(Src2_d[i*2], Src2_d[i*2 + 1]); \
} \
break; \
}
#define DO_VECTOR_SCALAR_OP(size, type, func)\
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], *Src2_d); \
} \
break; \
}
#define DO_VECTOR_0SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(); \
} \
break; \
}
#define DO_VECTOR_1SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src_d[i]); \
} \
break; \
}
#define DO_VECTOR_REDUCE_1SRC_OP(size, type, func, start_val) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
type begin = start_val; \
for (uint8_t i = 0; i < Elements; ++i) { \
begin = func(begin, Src_d[i]); \
} \
Dst_d[0] = begin; \
break; \
}
#define DO_VECTOR_SAT_OP(size, type, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(type, type2, func, min, max) \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src2); \
memcpy(Dst_d, Src1, Elements * sizeof(type2));\
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i+Elements] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP_SRC(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i+Elements], min, max); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i], (type)Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP_TOP_SRC(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i+Elements], (type)Src2_d[i+Elements]); \
} \
break; \
}
struct InterpVector256 {
__uint128_t Lower;
__uint128_t Upper;
};
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::NodeID Op) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Src.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
File diff suppressed because it is too large. Load diff
@@ -1,77 +0,0 @@
/*
$info$
tags: ir|opts
desc: Sanity checking pass
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/sstream.h>
#include "Interface/IR/PassManager.h"
#include <memory>
namespace FEXCore::IR::Validation {
class PhiValidation final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
bool PhiValidation::Run(IREmitter *IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::PHIValidation");
bool HadError = false;
auto CurrentIR = IREmit->ViewIR();
fextl::ostringstream Errors;
// Walk the list and calculate the control flow
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
bool FoundNonPhi{};
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
switch (IROp->Op) {
// BEGINBLOCK doesn't matter for us
case IR::OP_BEGINBLOCK: break;
case IR::OP_PHIVALUE:
case IR::OP_PHI: {
if (FoundNonPhi) {
// If we have found a non-phi IR op and then had a Phi or PhiValue value then this is a programming mistake
// PHI values MUST be defined at the top of the block only
HadError |= true;
Errors << "Phi %ssa" << CurrentIR.GetID(CodeNode) << ": Was defined after non-phi operations. Which is invalid!" << std::endl;
}
// Check all the phi values to ensure they have the same type
break;
}
default:
FoundNonPhi = true;
break;
}
}
}
if (HadError) {
fextl::stringstream Out;
FEXCore::IR::Dump(&Out, &CurrentIR, nullptr);
Out << "Errors:" << std::endl << Errors.str() << std::endl;
LogMan::Msg::EFmt("{}", Out.str());
}
return false;
}
fextl::unique_ptr<FEXCore::IR::Pass> CreatePhiValidation() {
return fextl::make_unique<PhiValidation>();
}
}
-13
View File
@@ -1,13 +0,0 @@
#pragma once
#include <cstdint>
namespace FEXCore {
[[nodiscard]] constexpr uint64_t AlignUp(uint64_t value, uint64_t size) {
return value + (size - value % size) % size;
}
[[nodiscard]] constexpr uint64_t AlignDown(uint64_t value, uint64_t size) {
return value - value % size;
}
} // namespace FEXCore
+1 -1
+1 -1
File renamed without changes.
File renamed without changes.
File renamed without changes.
@@ -27,7 +27,9 @@ def print_header():
#ifndef OPT_STRARRAY
#define OPT_STRARRAY(group, enum, json, default) OPT_BASE(fextl::string, group, enum, json, default)
#endif
#ifndef OPT_STRENUM
#define OPT_STRENUM(group, enum, json, default) OPT_BASE(uint64_t, group, enum, json, default)
#endif
'''
output_file.write(header)
@@ -40,6 +42,7 @@ def print_tail():
#undef OPT_UINT64
#undef OPT_STR
#undef OPT_STRARRAY
#undef OPT_STRENUM
'''
output_file.write(tail)
@@ -127,12 +130,13 @@ def print_man_options(options):
short = op_vals["ShortArg"]
default = op_vals["Default"]
value_type = op_vals["Type"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
if (value_type == "str" or value_type == "strarray" or value_type == "strenum"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_option(
@@ -141,6 +145,12 @@ def print_man_options(options):
op_vals["Desc"],
default
)
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
output_man.write(".El\n")
@@ -150,12 +160,13 @@ def print_man_environment(options):
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
value_type = op_vals["Type"]
# Textual default rather than enum based
if ("TextDefault" in op_vals):
default = op_vals["TextDefault"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
if (value_type == "str" or value_type == "strarray" or value_type == "strenum"):
# Wrap the string argument in quotes
default = "'" + default + "'"
print_man_env_option(
@@ -165,6 +176,13 @@ def print_man_environment(options):
False
)
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
print_man_environment_tail()
output_man.write(".El\n")
@@ -334,7 +352,7 @@ def print_argloader_options(options):
for op_key, op_vals in group_vals.items():
default = op_vals["Default"]
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray"):
if (op_vals["Type"] == "str" or op_vals["Type"] == "strarray" or op_vals["Type"] == "strenum"):
# Wrap the string argument in quotes
default = "\"" + default + "\""
@@ -382,7 +400,10 @@ def print_parse_argloader_options(options):
# boolean values need a decimal specifier. Otherwise fmt prints strings.
conversion_func = "fextl::fmt::format(\"{:d}\", "
if (value_type == "strarray"):
if (value_type == "strenum"):
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key, op_key))
elif (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
output_argloader.write("\tfor (auto iter = Array.begin(); iter != Array.end(); ++iter) {\n")
@@ -407,13 +428,56 @@ def print_parse_envloader_options(options):
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key, op_key))
output_argloader.write("}\n")
if ("ArgumentHandler" in op_vals):
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = {0}(Value);\n".format(conversion_func))
output_argloader.write("Value = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
def print_parse_enum_options(options):
output_argloader.write("#ifdef ENUMDEFINES\n")
output_argloader.write("#undef ENUMDEFINES\n")
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
if (op_vals["Type"] == "strenum"):
output_argloader.write("enum class {} : uint64_t {{\n".format(op_key))
Enums = op_vals["Enums"]
i = 0
# Always have an OFF.
output_argloader.write("\tOFF = 0,\n")
for enum_op_key, enum_op_vals in Enums.items():
output_argloader.write("\t{} = 1ULL << {},\n".format(enum_op_key.upper(), i))
i += 1
output_argloader.write("};\n")
output_argloader.write("FEX_DEF_NUM_OPS({})\n".format(op_key))
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
if (op_vals["Type"] == "strenum"):
Enums = op_vals["Enums"]
output_argloader.write("using {}ConfigPair = std::pair<std::string_view, FEXCore::Config::{}>;\n".format(op_key, op_key))
output_argloader.write("constexpr static std::array<{}ConfigPair, {}> {}_EnumPairs = {{{{\n".format(op_key, len(Enums) + 1, op_key))
i = 0
# Always have an OFF.
output_argloader.write("\t{{ \"off\", FEXCore::Config::{}::OFF }},\n".format(op_key))
for enum_op_key, enum_op_vals in Enums.items():
output_argloader.write("\t{{ \"{}\", FEXCore::Config::{}::{} }},\n".format(enum_op_vals, op_key, enum_op_key.upper()))
i += 1
output_argloader.write("}};\n")
output_argloader.write("#endif\n")
def check_for_duplicate_options(options):
short_map = []
long_map = []
@@ -492,4 +556,7 @@ print_parse_argloader_options(options);
# Generate environment loader code
print_parse_envloader_options(options);
# Generate enum variable options
print_parse_enum_options(options);
output_argloader.close()
@@ -251,11 +251,8 @@ def parse_ops(ops):
# Print out enum values
def print_enums():
if len(IROps) > 255:
ExitError("We have more than uint8_t ops. We have {}. Time to upgrade to uint16_t".format(len(IROps)))
output_file.write("#ifdef IROP_ENUM\n")
output_file.write("enum IROps : uint8_t {\n")
output_file.write("enum IROps : uint16_t {\n")
for op in IROps:
output_file.write("\tOP_{},\n" .format(op.Name.upper()))
@@ -282,7 +279,6 @@ def print_ir_structs(defines):
output_file.write("\tIROps Op;\n\n")
output_file.write("\tuint8_t Size;\n")
output_file.write("\tuint8_t ElementSize;\n")
output_file.write("\tuint8_t _pad;\n")
output_file.write("\ttemplate<typename T>\n")
output_file.write("\tT const* C() const { return reinterpret_cast<T const*>(Data); }\n")
@@ -679,7 +675,7 @@ def print_ir_allocator_helpers():
output_file.write("\t\t#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED\n")
for Validation in op.EmitValidation:
output_file.write("\t\tassert({});\n".format(Validation))
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"\");\n".format(Validation))
output_file.write("\t\t#endif\n")
output_file.write("\t\treturn Op;\n")
@@ -1,4 +1,4 @@
set (MAN_DIR ${CMAKE_INSTALL_PREFIX}/share/man CACHE PATH "MAN_DIR")
set (MAN_DIR share/man CACHE PATH "MAN_DIR")
set (FEXCORE_BASE_SRCS
Interface/Config/Config.cpp
@@ -133,11 +133,11 @@ set (SRCS
Interface/IR/Passes/DeadCodeElimination.cpp
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
File renamed without changes.
Loaded 100 of 444 files, more files were not shown because too many files have changed in this diff. Show more