Compare commits

..
593 Commits
Author SHA1 Message Date
Ryan Houdek bf66cac272 Docs: Update for release FEX-2309 2023-09-05 22:16:35 -07:00
Ryan Houdek 016c3c0f07 Merge pull request #3057 from Sonicadvance1/support_procfs_interpreter
FEXInterpreter: Supports procfs/interpreter
2023-09-05 21:53:45 -07:00
Ryan Houdek 09a49a3420 FEXInterpreter: Supports procfs/interpreter
This is a new procfs symlink path that changes behaviour of binfmt_misc
when exposed. We need to check both procfs/exe and procfs/interpreter
and see if they exist AND also differ.

Once/if they do then we can disable a bunch of checking of paths once
they do. The fallback when none of this is supported has the same
behaviour has previously where it still does all the regular checking.

During binfmt_misc install cmake will check the kernel version for the
raw binfmt_misc writing. Which will never pass until we have a real
kernel version that it is upstreamed in.

For update-binfmts we add a new optional argument where the tool will
drop the flag if the host kernel version isn't new enough to handle the
option.
2023-09-05 21:30:47 -07:00
Ryan Houdek 16f826c18e Merge pull request #3064 from alyssarosenzweig/ir/rm-phi
IR: Remove phi nodes
2023-09-05 19:43:52 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Ryan Houdek 04228952fa Merge pull request #3063 from alyssarosenzweig/flag/defer-pf-completely
Defer PF calculation completely
2023-09-05 13:18:22 -07:00
Alyssa Rosenzweig 5a3dc8c2ab InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:00:10 -04:00
Alyssa Rosenzweig 8efe2eeef6 OpcodeDispatcher: Defer PF invert
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.

This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:

* Zeroing PF. This now requires writing 1 instead of 0, which may require an
  extra move for the constant. Some of this will go away when we merge PF+AF
  into a single register, which is next up on the list. In that case, the
  two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
  x86 view of PF), that will get absorbed into the or, if we also write AF. If
  we leave AF undefined and need to write a 1, that's a single mov instruction
  and we couldn't do better anyway if not inverted (since we'd still have a mov
  wzr even then). So in view of the future work, this isn't something I'm
  concerned about.

* Float comparisons that put Unordered into PF. These require an extra invert to
  match the new convention. These are already so unnecessarily bloated that I'm
  not convinced I'm making things materially worse here. But we realistically
  need multiple destination support in the IR to fix this particular mess.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 15:58:18 -04:00
Ryan Houdek 8184c55424 Merge pull request #3062 from alyssarosenzweig/flag/no-pf
Remove ABINoPF option
2023-09-05 12:26:45 -07:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Ryan Houdek a8bc6bbb2e Merge pull request #3059 from alyssarosenzweig/flag/defer-af-xor
Defer second XOR for AF
2023-09-05 11:38:47 -07:00
Alyssa Rosenzweig 305bb98cf8 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig 80ac824dd9 Unittests: Fix bogus lahf tests
Logical ops leave AF undefined so we can't expect it to be zero after. Mask the
result of lahf to avoid testing UB. These unit tests would regress from the work
in this MR otherwise.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Ryan Houdek c27f69dd6b Merge pull request #3061 from alyssarosenzweig/flag/cmc
OpcodeDispatcher: Optimize CMC
2023-09-05 11:06:21 -07:00
Alyssa Rosenzweig b6462ee854 OpcodeDispatcher: Optimize CMC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 13:55:05 -04:00
Ryan Houdek 252ca88b3e Merge pull request #3058 from alyssarosenzweig/flag/undef
Stop zeroing undefined flags
2023-09-05 10:49:02 -07:00
Alyssa Rosenzweig 8ba1e91699 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig ff0b514da8 OpcodeDispatcher: Invalidate PF/AF in more cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig 240260576b OpcodeDispatcher: Stop zeroing so many flags
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:10:29 -04:00
Ryan Houdek 486f0ba1e3 Merge pull request #3038 from alyssarosenzweig/flag/af
Defer AF extract
2023-09-05 08:54:47 -07:00
Ryan Houdek 67a26a0e98 Merge pull request #3056 from Sonicadvance1/constpool_heuristic
ConstProp: Adds constpool distance heuristic
2023-09-05 08:46:22 -07:00
Alyssa Rosenzweig a1e3ce3cdf InstCountCI: Update for deferred AF extract
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:39:32 -04:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00
Alyssa Rosenzweig f46ba52e0e OpcodeDispatcher: Use LoadAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Alyssa Rosenzweig c88b022d9f OpcodeDispatcher: Add AF accessors
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Ryan Houdek a24680b0a1 ConstProp: Adds constpool distance heuristic
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.

This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.

FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.

To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.
2023-09-03 18:27:43 -07:00
Mai 2d78b1fbca Merge pull request #3055 from Sonicadvance1/defer_disasm_init
Arm64: Only allocate vixl::Decoder if enabled
2023-09-03 18:58:42 -04:00
Ryan Houdek ca1c33047c Arm64: Only allocate vixl::Decoder if enabled
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.

Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
2023-09-03 12:53:18 -07:00
Ryan Houdek b18592f153 Merge pull request #3050 from Sonicadvance1/fix_flag_reconstruction
OpcodeDispatcher: Fixes NZCV and PF flag compacting
2023-09-03 10:15:44 -07:00
Mai 8017a91e52 Merge pull request #3054 from Sonicadvance1/remove_small_ir_size_assumptions
OpcodeDispatcher: Remove final assumptions about small IR operating sizes
2023-09-03 05:47:38 -04:00
Ryan Houdek a235c3d81a InstCountCI: Update for small operator changes and opt 2023-09-03 02:26:23 -07:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Mai 62fcf6cbd0 Merge pull request #3053 from Sonicadvance1/rflag_handling_32bit
OpcodeDispatcher: Cleans up RFLAGS size handling
2023-09-03 05:13:42 -04:00
Ryan Houdek cb8f183e9f InstCountCI: Update for flags cleanup 2023-09-03 01:41:52 -07:00
Ryan Houdek 44a14e7fd0 OpcodeDispatcher: Cleans up RFLAGS size handling
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.

Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
2023-09-03 01:41:04 -07:00
Mai 2d22176699 Merge pull request #3052 from Sonicadvance1/remove_todo_shlimm
OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
2023-09-03 04:38:52 -04:00
Mai 338cb199a0 Merge pull request #3051 from Sonicadvance1/remove_todo_shiftleft
OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront
2023-09-03 04:38:00 -04:00
Ryan Houdek 16839671e5 InstcountCI: Update for SHLImm changes 2023-09-03 00:28:36 -07:00
Ryan Houdek c425db7284 OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
This one was easy, barely anything changes behaviour, as seen by
InstCountCI changes.
2023-09-03 00:27:28 -07:00
Ryan Houdek ba422c1dc9 InstCountCI: Update for ShiftLeft changes 2023-09-03 00:16:46 -07:00
Ryan Houdek d0595b5f13 OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront 2023-09-03 00:16:26 -07:00
Ryan Houdek 8bae58dcc3 Merge pull request #3049 from bylaws/x87
FEXCore: Rework X87 tag word handling
2023-09-03 00:00:12 -07:00
Ryan Houdek cdfa6939b3 FEXLinuxTests: Adds unit tests to ensure we set EFLAGS correctly
Tests all five of the flags that need specific handling. Without the
prior patch FEX would fail these.
2023-09-02 23:10:57 -07:00
Ryan Houdek cb5d665046 OpcodeDispatcher: Fixes NZCV and PF flag compacting
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.

This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
2023-09-02 23:10:57 -07:00
Billy Laws 393b8c657a InstCountCI: Update for x87 changes 2023-09-02 09:17:33 -07:00
Billy Laws 3792f707dc FEXLoader: Convert between abridged/full tag fmts in signal dispatch
X86 fpstate expects FTW to be saved in the FSAVE format, whereas X64
fpstate expects it to be saved in the abridged format used by FXSAVE.
2023-09-02 09:17:33 -07:00
Billy Laws 13b8f95f85 X87: Switch all stack pointer accesses to 32-bit OpSize 2023-09-02 09:17:33 -07:00
Billy Laws cb49373f47 FEXCore: Rework X87 tag word handling
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
2023-09-02 09:17:33 -07:00
Ryan Houdek ea965810e5 Merge pull request #3048 from Sonicadvance1/construct_eflags
Context: Adds helper to reconstruct and consume packed EFLAGS
2023-09-02 08:14:58 -07:00
Ryan Houdek 435f03c703 Context: Adds helper to reconstruct and consume packed EFLAGS
Currently FEX's internal EFLAGS representation is a perfect 1:1 mapping
between bit offset and byte offset. This is going to change with #3038.
There should be no reason that the frontend needs to understand how to
reconstruct the compacted flags from the internal representation.

Adds context helpers and moves all the logic to FEXCore. The locations
that previously needed to handle this have been converted over to use
this.
2023-09-02 07:05:54 -07:00
Ryan Houdek 81046efcd3 Merge pull request #3047 from alyssarosenzweig/build/no-telem
SignalDelegator: Fix build with telemetry disabled
2023-09-02 05:47:33 -07:00
Alyssa Rosenzweig 69b0747281 SignalDelegator: Fix build with telemetry disabled
/home/alyssa/FEX/Source/Tools/FEXLoader/LinuxSyscalls/SignalDelegator.cpp:1546:7: error: use of undeclared identifier 'CrashMask'
      CrashMask |= (1ULL << Signal);

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-02 08:09:05 -04:00
Mai f0ab9603e1 Merge pull request #3046 from Sonicadvance1/fix_telemetry_exit_group
Syscalls: Fix telemetry with exit_group
2023-09-01 05:15:56 -04:00
Ryan Houdek aea4a88e00 Syscalls: Fix telemetry with exit_group
Picked up on more games exiting with exit_group that I want to ensure
their telemetry data gets saved. Implement support for this.
2023-09-01 01:53:07 -07:00
Ryan Houdek ee8092bdc8 Merge pull request #3035 from alyssarosenzweig/flag/opts
Optimize ADD flag calculation
2023-08-31 16:30:01 -07:00
Alyssa Rosenzweig 31c92a422b InstCountCI: Update for add/sub work
Big Delete The Code energy.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:04 -04:00
Alyssa Rosenzweig 2edce18a59 ConstProp: Optimize XOR with 0
This cleans up the AF flag calculation for `neg`.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:04 -04:00
Alyssa Rosenzweig 659568ec1b ConstProp: Propagate 0 to first argument of SubNZCV
This allows inlining a constant into the comparison for Neg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:01 -04:00
Alyssa Rosenzweig dcb09b085b OpcodeDispatcher: Use AddNZCV/SubNZCV
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:41 -04:00
Alyssa Rosenzweig 45d99a2cce OpcodeDispatcher: Fix source sizes for Sub flags
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:39 -04:00
Alyssa Rosenzweig f688364bf2 IR: Add AddNZCV/SubNZCV op
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.

Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 07:52:12 -04:00
Mai 7c81a0d4fe Merge pull request #3042 from Sonicadvance1/optimize_call
OpcodeDispatcher: Optimize calls with push
2023-08-31 02:23:06 -04:00
Mai 7f997387b0 Merge pull request #3044 from Sonicadvance1/optimize_aes
Arm64: Optimize AES operations by caching a zero register
2023-08-31 00:20:10 -04:00
Ryan Houdek cfe6b929bb InstCountCI: Updates for aes optimization
XMM AES operation is now optimal.
Most other changes are just because of RA changing around the zero
register.
2023-08-30 20:43:11 -07:00
Ryan Houdek 61df7a576a OpcodeDispatcher: Be super defensive when starting a new block
Ensure all cached data is correct.
2023-08-30 20:43:11 -07:00
Ryan Houdek ffaa908475 OpcodeDispatcher: Use the named zero register for each usage
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
2023-08-30 20:43:11 -07:00
Ryan Houdek a2307f28d6 Arm64: Optimize AES operations by caching a zero register
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.

Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.

AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
2023-08-30 19:00:43 -07:00
Ryan Houdek 1446d4fe12 IR: Adds support for named vector zero
This is useful for caching a zero register vector which we use in
various locations. This will be abused soon.
2023-08-30 18:59:38 -07:00
Mai 9fb8c95ef4 Merge pull request #3043 from Sonicadvance1/optimize_blendv
OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
2023-08-30 20:36:14 -04:00
Ryan Houdek 58caee3614 InstCountCI: Update for optimized blendv 2023-08-30 17:13:05 -07:00
Ryan Houdek 24215f7ad0 OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
2023-08-30 17:09:34 -07:00
Ryan Houdek b20c518bf0 OpcodeDispatcher: Optimize calls with push
InstCountCI doesn't cover branch instructions so needs manual
inspection.
2023-08-30 16:25:01 -07:00
Ryan Houdek d738538a34 Merge pull request #3041 from Sonicadvance1/telem_runtime
FEXCore: Allows disabling telemetry at runtime
2023-08-30 13:30:51 -07:00
Ryan Houdek 48df17b836 InstCountCI: Update for changed tests
Since the telemetry array is always in the corestate now some numbers
change.
2023-08-30 13:01:27 -07:00
Ryan Houdek 81a32c3998 FEXCore: Allows disabling telemetry at runtime
This is useful for InstCountCI so you can disable the telemetry
gathering even if enabled so it doesn't affect the CI system.
2023-08-30 12:59:41 -07:00
Mai 02b891c0fe Merge pull request #3040 from Sonicadvance1/optimize_aeskeygen
Arm64: Optimize AESKeyGenAssist
2023-08-30 15:58:29 -04:00
Mai 012750f2bb Merge pull request #3037 from Sonicadvance1/remove_debug_log
OpcodeDispatcher: Removes erroneous debug log
2023-08-30 15:57:31 -04:00
Ryan Houdek b40147a269 InstCountCI: Update for aeskeygenassist 2023-08-30 12:17:07 -07:00
Ryan Houdek d8f131fa3d Arm64: Optimize AESKeyGenAssist
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.

I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
2023-08-30 12:15:09 -07:00
Ryan Houdek 2c2081c977 Merge pull request #3039 from Sonicadvance1/print_OpSize
IR: Adds printer for OpSize
2023-08-30 11:50:13 -07:00
Ryan Houdek 572cc57aa3 IR: Adds printer for OpSize 2023-08-30 11:32:43 -07:00
Ryan Houdek 9bfd4b650f OpcodeDispatcher: Removes erroneous debug log 2023-08-30 11:12:12 -07:00
Ryan Houdek e1c2033fa9 Merge pull request #3036 from Sonicadvance1/OpSize_IRDump_Parse
IR: Adds IR::OpSize to IRDumper
2023-08-30 11:01:06 -07:00
Ryan Houdek d3ee794dcb IR: Adds IR::OpSize to IRDumper
This was missing before.
2023-08-30 10:43:23 -07:00
Mai b40f784da0 Merge pull request #3034 from Sonicadvance1/remove_implicit_add
IR: Removes implicit sized add
2023-08-30 08:50:49 -04:00
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Mai 351c0eee42 Merge pull request #3033 from Sonicadvance1/remove_implicit_bfe
IR: Removes implicit sized bfe
2023-08-29 23:02:02 -04:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Mai 8239f8aa27 Merge pull request #3031 from Sonicadvance1/remove_implicit_and
IR: Removes implicit sized and
2023-08-29 02:32:40 -04:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek 59a4d15907 Merge pull request #3030 from Sonicadvance1/remove_implicit_sub
IR: Removes implicit sized sub
2023-08-28 22:20:30 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Mai 227ba9f9fc Merge pull request #3029 from Sonicadvance1/remove_variable_bfi
IR: Removes bfi from variable size
2023-08-29 00:50:45 -04:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00
Mai 9b55a34e75 Merge pull request #3028 from Sonicadvance1/remove_implicit_xor
IR: Removes implicit sized xor
2023-08-28 23:59:39 -04:00
Ryan Houdek 8bc84c202c IR: Removes implicit sized xor 2023-08-28 19:51:14 -07:00
Ryan Houdek e9a3848602 Merge pull request #3027 from Sonicadvance1/remove_implicit_andn
IR: Removes implicit sized andn
2023-08-28 19:39:49 -07:00
Ryan Houdek 516f27bff5 Merge pull request #3026 from Sonicadvance1/remove_implicit_or
IR: Removes implicit sized or
2023-08-28 19:39:42 -07:00
Ryan Houdek 1699ec9a76 IR: Removes implicit sized andn 2023-08-28 19:16:16 -07:00
Ryan Houdek db6c8852fc IR: Removes implicit sized or 2023-08-28 19:06:05 -07:00
Ryan Houdek 24a9254e37 Merge pull request #3025 from Sonicadvance1/remove_implicit_lshr
IR: Removes implicit sized lshr
2023-08-28 19:05:47 -07:00
Ryan Houdek 65dc6f3e90 IR: Removes implicit sized lshr 2023-08-28 18:16:56 -07:00
Mai 8534d3dfbf Merge pull request #3024 from Sonicadvance1/remove_implicit_lshl
IR: Removes implicit sized lshl
2023-08-28 21:08:05 -04:00
Ryan Houdek 60c4438780 IR: Removes implicit sized lshl 2023-08-28 17:50:41 -07:00
Ryan Houdek 915e52046c Merge pull request #3023 from Sonicadvance1/remove_mul
IR: Removes implicit sized mul ops
2023-08-28 17:49:32 -07:00
Ryan Houdek 898ce1ce8f IR: Removes implicit sized UMulH 2023-08-28 17:20:55 -07:00
Ryan Houdek aa8dfd6af1 IR: Removes implicit sized UMul 2023-08-28 17:20:55 -07:00
Ryan Houdek 6a6d808b0d IR: Removes implicit sized MulH 2023-08-28 17:20:55 -07:00
Ryan Houdek fac5b2ac72 IR: Removes implicit sized Mul 2023-08-28 17:20:55 -07:00
Ryan Houdek a55616e4db Merge pull request #3022 from Sonicadvance1/remove_sext
IR: Removes sext IR helper
2023-08-28 17:20:39 -07:00
Ryan Houdek b9e4a1423f IR: Removes sext IR helper
You hold no power here IR operation.
2023-08-28 17:03:38 -07:00
Ryan Houdek 6f661534a2 Merge pull request #3021 from Sonicadvance1/remove_implicit_op_part_move
IR: Removes implicit sized {Create,Extract}ElementPair
2023-08-28 17:03:12 -07:00
Ryan Houdek c0bb6a053f IR: Removes implicit sized {Create,Extract}ElementPair 2023-08-28 16:50:00 -07:00
Ryan Houdek bc1e89d91d Merge pull request #3020 from Sonicadvance1/remove_implicit_ops_pt_atomic
Remove implicit sized IR ops part atomic
2023-08-28 07:27:47 -07:00
Ryan Houdek c1b4c11e54 Merge pull request #3018 from Sonicadvance1/opcodedispatcher_sizeopt
OpcodeDispatcher: Optimize Get{Src,Dst}Size
2023-08-28 07:13:41 -07:00
Ryan Houdek a36427d01e IR: Removes implicit sized CAS 2023-08-28 07:12:31 -07:00
Ryan Houdek 594baff705 IR: Removes implicit sized CASPair 2023-08-28 07:12:31 -07:00
Ryan Houdek e9f2bc037f IR: Removes non-opsize AtomicAdd/Sub/And/Or
These were unused
2023-08-28 07:12:31 -07:00
Ryan Houdek 6dcfd6eb73 IR: Removes non-opsize AtomicXor 2023-08-28 07:12:31 -07:00
Ryan Houdek 4f8a63459c IR: Removes non-opsize AtomicSwap 2023-08-28 07:12:31 -07:00
Ryan Houdek ab230bf527 IR: Removes non-opsize AtomicFetchAdd 2023-08-28 07:12:31 -07:00
Ryan Houdek 8ba9613972 IR: Removes non-opsize AtomicFetchSub 2023-08-28 07:12:31 -07:00
Ryan Houdek 9370e30af4 IR: Removes non-opsize AtomicFetchAnd 2023-08-28 07:12:31 -07:00
Ryan Houdek 25ce57ef92 IR: Removes non-opsize AtomicFetchOr 2023-08-28 07:12:31 -07:00
Ryan Houdek 436e0f6f86 IR: Removes non-opsize AtomicFetchXor 2023-08-28 07:12:31 -07:00
Ryan Houdek f44c6f394f IR: Removes non-opsize AtomicFetchNeg 2023-08-28 07:12:31 -07:00
Ryan Houdek 5415e4c95f Merge pull request #3017 from Sonicadvance1/remove_implicit_ops_pt1
Remove implicit sized IR ops part 1
2023-08-28 07:11:44 -07:00
Ryan Houdek bb1362b2bf OpcodeDispatcher: Optimize Get{Src,Dst}Size
These functions are called a lot....a lot a lot.
Optimize these in to a couple of ALU operations instead of a whole
table lookup. Confirming with output assembly that this becomes more
optimal.
2023-08-28 05:29:56 -07:00
Ryan Houdek 51baf594c9 OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek 48669b7006 IR: Removes implicit sized orlshl/orlshr 2023-08-28 05:04:32 -07:00
Ryan Houdek 16481c0e55 IR: Removes implicit sized rev 2023-08-28 05:04:32 -07:00
Ryan Houdek b00310674a IR: Removes implicit sized not 2023-08-28 05:04:32 -07:00
Ryan Houdek 5768444ce9 IR: Removes implicit sized abs 2023-08-28 05:04:32 -07:00
Ryan Houdek 1f1473eb74 IR: Removes implicit sized neg 2023-08-28 05:04:32 -07:00
Ryan Houdek cccd7001cb IR: Removes implicit sized ashr 2023-08-28 05:04:32 -07:00
Ryan Houdek ce8392d5ae IR: Removes implicit sized ror 2023-08-28 05:02:01 -07:00
Ryan Houdek 386cf36cfd IR: Removes implicit sized sbfe
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.

We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek 0ddc23a5c9 IR: Removes implicit sized Popcount 2023-08-28 05:02:01 -07:00
Ryan Houdek b95648a4ab IR: Removes implicit sized FindMSB 2023-08-28 05:02:01 -07:00
Ryan Houdek bf18672999 IR: Removes implicit sized FindLSB 2023-08-28 05:02:01 -07:00
Ryan Houdek f405c1be69 IR: Removes inverted EntrypointOffset/InlineEntrypointOffset 2023-08-28 05:02:01 -07:00
Ryan Houdek f3cd115fa7 IR: Removes implicit sized CountLeadingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek a14720e130 IR: Removes implicit sized FindTrailingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek 86ed909de8 IR: Removes implicit sized DIV/REM 2023-08-28 05:02:00 -07:00
Ryan Houdek ac7e75c06b IR: Removes implicit sized UDIV/UREM 2023-08-28 05:02:00 -07:00
Ryan Houdek ec3e7ceeb5 IR: Removes implicit sized EXTR 2023-08-28 05:02:00 -07:00
Ryan Houdek c8c8ddbd4f IR: Removes implicit sized PDEP/PEXT 2023-08-28 05:02:00 -07:00
Ryan Houdek 35013bda37 IR: Adds helper to convert between an integer size and IR::OpSize
This is a nop operation and will get optimized away in release builds.
2023-08-28 05:02:00 -07:00
Ryan Houdek 62a9a075b7 Arm64: Leave a comment that 32-bit division shouldn't leave garbage in upper 64-bits 2023-08-28 05:02:00 -07:00
Ryan Houdek ea6d068cc5 IR: Removes implicit sized LDIV/LREM 2023-08-28 05:02:00 -07:00
Ryan Houdek 5013473ec0 IR: Removes implicit sized LUDIV/LUREM 2023-08-28 05:02:00 -07:00
Alyssa Rosenzweig b801bacb36 Merge pull request #3016 from Sonicadvance1/update_instcountci
InstCountCI: Update for previous changes
2023-08-28 07:42:20 -04:00
Ryan Houdek f02d291da1 InstCountCI: Update for previous changes
Missed thsi in the SRA PR
2023-08-27 23:41:41 -07:00
Ryan Houdek 6f2b3e76ac Merge pull request #3013 from Sonicadvance1/32bit_sra
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek 1d7c280367 Merge pull request #3012 from Sonicadvance1/optimize_movmskps
OpcodeDispatcher: Optimizes SSE movmaskps
2023-08-27 21:29:04 -07:00
Ryan Houdek d6864871ed InstCountCI: Update for movmskps optimization 2023-08-27 21:07:20 -07:00
Ryan Houdek 514a8223d9 OpcodeDispatcher: Optimizes SSE movmaskps
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.

I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek 8d110738ac IR: Add option to disable vector shift range clamping
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.

Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.

When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Mai 590345b125 Merge pull request #3014 from Sonicadvance1/remove_implicit_alu_ops
IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
2023-08-27 20:25:11 -04:00
Mai e2ac97f4db Merge pull request #3015 from Sonicadvance1/instcount_ci_rorx_abovemax
InstCountCI: Test rorx at max mask size
2023-08-27 20:21:45 -04:00
Ryan Houdek a216ea465c InstCountCI: Test rorx at max mask size
Just to ensure we capture the nop nature of passing in the rotate amount
of the operating size.
2023-08-27 01:43:47 -07:00
Ryan Houdek e4bb0df486 IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.

This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.

Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek 7146691360 IR/Passes/RA: Enable SRA for 32-bit GPRs
Noticed that we hadn't ever enabled this, which was a concern when our
GPR operations weren't as strict about leaving garbage in the upper bits
when operating as a 32-bit operation.

Now that our ALU operations are more strict about enforcing upper bit
zeroing we can enable this.

This causes Half-Life: Source FPS to get to > 200FPS finally. Causes
significant performance improvements for 32-bit games because we're no
longer redundantly moving registers before and after every operation.
Causing a bunch of 3-4 instruction sequences to convert to 1.
2023-08-26 18:22:50 -07:00
Ryan Houdek 6cb1afc2ad unittests/asm: Adds ADC/SBB test 2023-08-26 18:22:50 -07:00
Ryan Houdek 572d6cd3e6 OpcodeDispatcher: Fixes ADC and SBB 2023-08-26 18:22:50 -07:00
Ryan Houdek fcc37bf6a8 OpcodeDispatcher: Fixes RCR and ADOX 32-bit
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek 8f7925d06f Arm64: Simple typo fix 2023-08-26 18:22:50 -07:00
Ryan Houdek eace648fa9 OpcodeDispatcher: Fixes bug in UMUL
This was trying to operating on a 32-bit value but BFE the upper
32-bits.

Actually fixes this so it is operating on the 64-bit multiply result.
2023-08-26 18:22:50 -07:00
Ryan Houdek a4ac21a4e4 OpcodeDispatcher: Fixes bug in GetRFLAG with CachedNZCV
This was only operating at byte size but it was attempting to get bit
offsets at greater than operating size.
Change operating size over to 64-bit.
2023-08-26 18:22:50 -07:00
Ryan Houdek a01e69092d Arm64: Ensure Bfe and Sbfe operate at 32-bit or 64-bit op size
For Sbfe at least it ensures the upper bits don't get filled with
garbage.
Bfe it doesn't change behaviour but best to be correct.
2023-08-26 18:22:50 -07:00
Ryan Houdek 2fde2140ef Arm64: Ensure assert is testing correct array 2023-08-26 18:22:50 -07:00
Ryan Houdek 547daf8b6e Merge pull request #2857 from Sonicadvance1/fix_ra_validation
IR: Fixes RAValidation for 32-bit applications
2023-08-26 18:22:18 -07:00
Ryan Houdek e10afefb2b IR: Fixes RAValidation for 32-bit applications
RAValidation was making an assumption that GPR register class would only
have up to 16 registers for either SRA or dynamic registers.

When running a 32-bit application we allow 17 GPRs to be dynamically
allocated, since we can take 8 back from SRA in that case.

Just split the two classes in the RAValidation pass since they will
never overlap their allocation.

Fixes validation in `32Bit_Secondary/15_XX_0.asm` locally that changed
behaviour due to tinkering.
2023-08-26 16:18:03 -07:00
Mai 0195bb6e5a Merge pull request #3010 from Sonicadvance1/exit_group_quit
Linux: Call exit_group when application tries
2023-08-25 21:13:59 -04:00
Ryan Houdek edd27a2723 Linux: Call exit_group when application tries
When the application calls exit_group we no longer need to care about
cleanup because the entire process group is leaving.

Just immediately call exit group and get out. Might revisit this in the
future.

Fixes #2752
2023-08-25 17:41:57 -07:00
Mai 7c6660f634 Merge pull request #3009 from Sonicadvance1/fix_frint64
X8764: Ensure frndint uses host rounding mode
2023-08-25 20:01:01 -04:00
Ryan Houdek 9ba46f429e X8764: Ensure frndint uses host rounding mode
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.

This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.

In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.

[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00
Mai db3dc3edf4 Merge pull request #3003 from Sonicadvance1/optimize_pshuf
OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
2023-08-25 16:36:52 -04:00
Ryan Houdek c26798af83 InstCountCI: Update for shuffles! 2023-08-25 12:59:40 -07:00
Ryan Houdek a76c2c57b0 OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek 7f63d87295 IR: Adds support for new LoadNamedVectorIndexedConstant IR 2023-08-25 12:59:40 -07:00
Mai bf12f08218 Merge pull request #3002 from Sonicadvance1/optimize_movmaskpd
OpcodeDispatcher: Optimize 128-bit movmaskpd
2023-08-25 08:50:28 -04:00
Mai 1f7d138d2a Merge pull request #3008 from Sonicadvance1/optimize_movddup
OpcodeDispatcher: Optimize movddup from register
2023-08-25 08:48:40 -04:00
Mai f36f07055a Merge pull request #3007 from Sonicadvance1/optimize_cvtdq2pd
OpcodeDispatcher: Optimize cvtdq2pd from register source
2023-08-25 08:47:47 -04:00
Mai 30a1a382c4 Merge pull request #3006 from Sonicadvance1/optimize_movq
OpcodeDispatcher: Optimizes movq
2023-08-25 08:46:02 -04:00
Mai 631655dd81 Merge pull request #3005 from Sonicadvance1/nontemporalmoves
OpcodeDispatcher: Optimize nontemporal moves
2023-08-25 08:44:48 -04:00
Mai 0ef439f830 Merge pull request #3004 from Sonicadvance1/optimize_scalar_cvt
OpcodeDispatcher: Generate more optimal code for scalar GPR converts
2023-08-25 08:43:56 -04:00
Ryan Houdek 2aba628a24 InstCountCI: Update for movddup optimization 2023-08-25 03:44:30 -07:00
Ryan Houdek 2fbcf2e4a9 OpcodeDispatcher: Optimize movddup from register
This is now optimal
2023-08-25 03:44:10 -07:00
Ryan Houdek 7367f166b0 InstCountCI: Update for cvtdq2pd optimization 2023-08-25 03:39:45 -07:00
Ryan Houdek 00124205e5 OpcodeDispatcher: Optimize cvtdq2pd from register source
This is now optimal
2023-08-25 03:39:24 -07:00
Ryan Houdek 1a3a59dd38 InstCountCI: Update for movq optimization 2023-08-25 03:30:52 -07:00
Ryan Houdek 81281e2115 OpcodeDispatcher: Optimizes movq
Removes a redundant move between registers and makes it optimal.
Also removes a couple redundant moves on the avx version.
2023-08-25 03:29:58 -07:00
Ryan Houdek a901e5f3df InstCountCI: Update for optimized cvtsi2s{s,d} 2023-08-25 03:19:11 -07:00
Ryan Houdek 1cb2b084b3 OpcodeDispatcher: Generate more optimal code for scalar GPR converts
1) In the case that we are converted a GPR, don't zero extend it first.
2) In the case that the scalar comes from memory, load it first in an
   FPR and converted it in-place.

These are now optimal in the case of AFP is unsupported.
2023-08-25 03:19:11 -07:00
Ryan Houdek 189b0da68f JIT/Int: Add support for scalar conversion as well 2023-08-25 03:19:11 -07:00
Ryan Houdek 9fa877e512 InstCountCI: Update for move temporal optimization 2023-08-25 03:13:28 -07:00
Ryan Houdek f3679a99ec OpcodeDispatcher: Optimize nontemporal moves
These are now optimal.
2023-08-25 03:13:10 -07:00
Ryan Houdek 62156f2152 ARM64JIT: Adds support for scalar cvt 2023-08-25 02:34:30 -07:00
Ryan Houdek 3808f97283 InstCountCI: Update for movmaskpd optimization 2023-08-24 17:27:54 -07:00
Ryan Houdek 3a6d25f56a unittests/ASM: Update movmskpd test to include an edge of garbage but no sign bit 2023-08-24 17:27:54 -07:00
Ryan Houdek 9a54898429 OpcodeDispatcher: Optimize 128-bit movmaskpd
I'd consider this optimal now.

Thanks to @dougallj for the optimization idea again!
2023-08-24 17:27:54 -07:00
Ryan Houdek 80d871fb18 Merge pull request #3001 from Sonicadvance1/optimize_cvtps2pd
OpcodeDispatcher: Optimize cvtps2pd
2023-08-24 16:09:12 -07:00
Ryan Houdek 4d58ec1025 Merge pull request #3000 from Sonicadvance1/optimize_cvt
OpcodeDispatcher: Optimize MMX conversion operation
2023-08-24 16:09:04 -07:00
Ryan Houdek 60ad76732a InstCountCI: Update for cvtps2pd optimization 2023-08-24 15:56:02 -07:00
Ryan Houdek 3731e6d88b OpcodeDispatcher: Optimize cvtps2pd
SSE version is now optimal and AVX version gets rid of a redundant move.
2023-08-24 15:55:11 -07:00
Ryan Houdek e3812f9c3c InstCountCI: Update tests for mmx optimization 2023-08-24 15:46:54 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek 72ce7ddf2d Arm64: Optimize CVT operations for 64-bit variants
Using 128-bit converts for 64-bit versions cuts their throughput in half
on Cortex. Ensure we use the 64-bit version when possible.
2023-08-24 15:45:07 -07:00
Ryan Houdek e025d32531 Merge pull request #2999 from lioncash/catch
Externals: Update Catch2 to v2.13.10
2023-08-24 15:19:23 -07:00
Ryan Houdek 200dbdd0e6 Merge pull request #2998 from Sonicadvance1/optimize_addsub_fcma
OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
2023-08-24 15:19:06 -07:00
Lioncache 989fe22e2d Externals: Update Catch2 to v2.13.10
Updates it to the latest v2 branch tag
2023-08-24 18:04:59 -04:00
Ryan Houdek 4e7adeec85 InstCountCI: Adds new files for FCMA 2023-08-24 15:00:42 -07:00
Ryan Houdek a1210f892a OpcodeDispatcher: Optimize addsubp{s,d} using fcadd
This extension was added with seemingly Cortex-A710 and turns this
instruction in to two instructions which is quite good.

Needs #2994 merged first.

Huge thanks to @dougallj for the optimization idea!
2023-08-24 15:00:41 -07:00
Ryan Houdek ba01eac467 IR: Adds support for ARM's FCMA FCADD instruction 2023-08-24 15:00:41 -07:00
Ryan Houdek c5d147322f HostFeatures: Adds support for FCMA 2023-08-24 15:00:41 -07:00
Ryan Houdek df99b7b9b6 Merge pull request #2994 from Sonicadvance1/cache_namedvectorconstants
OpcodeDispatcher: Cache named vector constants in the block
2023-08-24 15:00:05 -07:00
Ryan Houdek 565b30e15e OpcodeDispatcher: Cache named vector constants in the block
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.

Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.

Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Mai ab83ab42dd Merge pull request #2993 from Sonicadvance1/addsub_opt
OpcodeDispatcher: Optimize AddSubP{S,D}
2023-08-24 17:44:41 -04:00
Ryan Houdek b0ec4197ba Merge pull request #2997 from lioncash/fmt
Externals: Update fmt to 10.1.0
2023-08-24 14:19:18 -07:00
Lioncache 9035a29906 Externals: Update fmt to 10.1.0
Updates fmt to the latest version.
2023-08-24 17:01:40 -04:00
Ryan Houdek b547550442 InstCountCI: Update for addsubp 2023-08-23 20:33:55 -07:00
Ryan Houdek f300196d90 OpcodeDispatcher: Optimize AddSubP{S,D}
Use a named constant for loading the sign inversion, then EOR the second
source and just FAdd it all.
In a vacuum it isn't a significant improvement, but as soon as more than
one instruction is in a block it will eventually get optimized with
named constant caching and be a significant win.

Thanks to @rygorous for the idea!
2023-08-23 20:32:51 -07:00
Ryan Houdek c5f358b47a Merge pull request #2992 from lioncash/doc
x86_64/MemoryOps: Fix mislabeled IR op messages
2023-08-23 20:05:39 -07:00
Lioncache 42ccc18606 x86_64/MemoryOps: Fix mislabeled IR op messages 2023-08-23 22:54:36 -04:00
Mai 66c6f96120 Merge pull request #2990 from Sonicadvance1/optimize_pmulh
OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
2023-08-23 22:06:14 -04:00
Ryan Houdek 1aa2c534f7 Merge pull request #2991 from lioncash/pcl
OpcodeDispatcher: Remove redundant moves from PCLMULQDQ and AES operations
2023-08-23 18:47:52 -07:00
Ryan Houdek e9f292462a InstCountCI: Update for pmulh{u,}w optimization 2023-08-23 18:38:05 -07:00
Ryan Houdek 77b6d854b9 OpcodeDispatcher: Optimize PMULH{U,}W using new IR operations
SSE implementations are now optimal.
SVE-128bit operation makes it more optimal.
2023-08-23 18:38:05 -07:00
Ryan Houdek 05b9651279 IR: Implements new vector multiply returning high bits
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 8a622a3c1a OpcodeDispatcher: Remove redundant moves from VAESEnc
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:30:06 -04:00
Lioncache f4848fd1a7 OpcodeDispatcher: Remove redundant moves from VAESEncLast
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:28:14 -04:00
Lioncache d37ce08ae9 OpcodeDispatcher: Remove redundant move from VAESDec
Zero-extension will occur automatically upon storing if necessary.
2023-08-23 21:26:49 -04:00
Lioncache a6f1a9f8e8 OpcodeDispatcher: Remove redundant moves from VAESDecLast
Zero-extension will occur upon storing if necessary.
2023-08-23 21:25:15 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache 410e99ba09 OpcodeDispatcher: Remove redundant moves from VPCLMULQDQOp
Zero-extension will occur if necessary upon storing.
2023-08-23 21:18:39 -04:00
Ryan Houdek 6e4765d48b Merge pull request #2989 from lioncash/ins
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
2023-08-23 18:09:15 -07:00
Ryan Houdek 172c8f3ba6 Merge pull request #2988 from lioncash/half
Arm64/ConversionOps: Add missing half-precision conversions to scalar functions
2023-08-23 17:56:23 -07:00
Lioncache 203a2b1105 Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
If Dst and DestVector alias one another, then we don't need to
move the vector unnecessarily.
2023-08-23 20:50:21 -04:00
Ryan Houdek 4297e13fcf Merge pull request #2986 from lioncash/ext
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
2023-08-23 17:40:50 -07:00
Ryan Houdek 5b8a0f1e0d Merge pull request #2987 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from VSQXTN2/VSQXTUN2/VSQSHL/VSRSHR
2023-08-23 17:40:39 -07:00
Lioncache 5ad56ad52e Arm64/ConversionOps: Add missing half-precision operations to Float_FromGPR_S
Provides parity with vector operations.
2023-08-23 20:36:34 -04:00
Lioncache 24e7baf28f Arm64/ConversionOps: Add missing half-precision conversions to Float_FToF
Provides parity with the vector conversion operations.
2023-08-23 20:31:49 -04:00
Lioncache b248ae4c04 Arm64/VectorOps: Remove redundant moves from SVE SQSHL
We don't need to emit a move if the destination and source alias.
2023-08-23 20:12:09 -04:00
Lioncache 95bea864cf Arm64/VectorOps: Remove redundant moves from SVE SRSHR
We don't need to perform a move is the destination aliases
the source vector to be shifted.
2023-08-23 20:10:51 -04:00
Lioncache 47c4507bb6 Arm64/VectorOps: Remove redundant moves from SVE VSQXTUN2
We don't need to perform a move if the destination aliases the lower vector.
2023-08-23 20:10:29 -04:00
Lioncache 5ea0b6db28 Arm64/VectorOps: Remove redundant moves from SVE VSQXTN2
We don't need to perform a move if the destination aliases the
lower vector.
2023-08-23 20:02:28 -04:00
Mai ee10153d14 Merge pull request #2984 from Sonicadvance1/optimize_pack
OpcodeDispatcher: Use new IR ops for pack instructions
2023-08-23 20:02:16 -04:00
Lioncache d0d94adabe Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
We don't need to do any moves here is the destination aliases the
lower bits.
2023-08-23 19:56:10 -04:00
Ryan Houdek 926b8c2c97 Merge pull request #2985 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from SVE variable/immediate/vector shifts when possible
2023-08-23 16:41:39 -07:00
Lioncache 18ebcdc9de Arm64/VectorOps: Remove redundant moves in VUshrNI2
If the destination and VectorLower alias, then we don't need
to emit a movprfx.
2023-08-23 18:52:32 -04:00
Lioncache f31a9a52e6 Arm64/VectorOps: Remove redundant moves from SVE immediate vector shifts when possible
If the destination and source vector alias one another, then the
operation can largely be done in place.
2023-08-23 18:36:21 -04:00
Lioncache 03504a5f8c Arm64/VectorOps: Remove redundant moves from SVE vector shifts when possible
If the destination and the vector to be shifted alias, then we can
avoid needing to move some data around.
2023-08-23 18:24:58 -04:00
Lioncache d29b4de1ee Arm64/VectorOps: Remove redundant moves from SVE variable vector register shifts when possible
In the event that the destination and the vector to be shifted
alias one another, then we can skip the movprfx, since it's not
necessary.
2023-08-23 18:24:53 -04:00
Ryan Houdek 5e20be756e InstCountCI: Update for pack instruction optimization 2023-08-23 15:16:54 -07:00
Ryan Houdek fc4559d3c4 OpcodeDispatcher: Use new IR ops for pack instructions
The MMX and SSE versions of these instructions are now optimal.
2023-08-23 15:14:38 -07:00
Ryan Houdek c508570da0 IR: Implements VSQXT{U,}NPair operations
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Ryan Houdek ec6548e302 Merge pull request #2983 from lioncash/bsl
Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
2023-08-23 15:11:34 -07:00
Lioncache d5e145c4b0 Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
If the destination and true vector alias one another, then we can
perform the operation in place instead of moving data around.
2023-08-23 17:54:10 -04:00
Ryan Houdek 350bca97c6 Merge pull request #2982 from lioncash/imin
Arm64/VectorOps: Remove redundant moves from SVE V{S,U}Min/V{S,U}Max when possible
2023-08-23 14:53:15 -07:00
Ryan Houdek 226405880f Merge pull request #2981 from lioncash/fmin
Arm64/VectorOps: Remove redundant moves from SVE VFMin/VFMax when possible
2023-08-23 14:46:03 -07:00
Lioncache 37a8cb6821 Arm64/VectorOps: Remove redundant moves from SVE VSMax when possible
When the destination and first source alias one another, then we
can perform the operation in place instead of moving data around.
2023-08-23 17:34:52 -04:00
Lioncache fe2c7dbf97 Arm64/VectorOps: Remove redundant moves from SVE VUMax when possible
When the destination and source alias one another, then we
can perform the operation in place without needing to move
data around.
2023-08-23 17:32:17 -04:00
Lioncache 787b4f37fb Arm64/VectorOps: Remove redundant moves from SVE VSMin when possible
When the destination and first source alias one another, then we can
perform the operation in place without moving any data.
2023-08-23 17:30:08 -04:00
Lioncache c3faa019f5 Arm64/VectorOps: Remove redundant moves from SVE VUMin when possible
If the destination and first source alias, then we can perfom the operation
in place.
2023-08-23 17:27:52 -04:00
Ryan Houdek da098d8204 Merge pull request #2979 from lioncash/div
Arm64/VectorOps: Remove moves from SVE VFDiv if possible
2023-08-23 14:22:46 -07:00
Lioncache 149852b122 Arm64/VectorOps: Remove redundant moves from SVE VFMax if possible
If Dst and Vector1 alias one another, then the operation can be
performed in place instead of moving data around.
2023-08-23 17:18:41 -04:00
Lioncache ecf02846e6 Arm64/VectorOps: Remove redundant moves from SVE VFMin is possible
If Dst and Vector1 alias one another, then we can do the merging move
in place instead of shuffling data around.
2023-08-23 17:16:25 -04:00
Ryan Houdek 2501ebc1cd Merge pull request #2980 from lioncash/avg
Arm64/VectorOps: Remove redundant moves from SVE VURAvg if possible
2023-08-23 14:05:51 -07:00
Lioncache a8f7529847 Arm64/VectorOps: Remove moves from SVE VFDiv if possible
Given the operation is:

Dst = Vector1 / Vector2

If Dst and Vector1 alias one another, then we can just perform
the division as is without any moving of data around.
2023-08-23 16:53:56 -04:00
Lioncache 8431ab43a0 Arm64/VectorOps: Remove redundant moves from SVE VURAvg if possible
If Dst and Vector1 alias one another, then we can perform the operation
without needing to move any data around.
2023-08-23 16:51:17 -04:00
Mai 4b06069c0d Merge pull request #2972 from Sonicadvance1/optimize_scalar_mov
OpcodeDispatcher: Optimizes scalar movd/movq
2023-08-23 16:30:45 -04:00
Ryan Houdek b646f4b781 Merge pull request #2978 from lioncash/misc
OpcodeDispatcher: Remove redundant moves from remaining AVX ops
2023-08-23 13:18:25 -07:00
Ryan Houdek 38853c2a9b InstCountCI: Update for vmovd/vmovq optimization 2023-08-23 12:56:11 -07:00
Ryan Houdek 8836ab8988 OpcodeDispatcher: Optimizes scalar movd/movq
MMX and SSE versions are now optimal.
2023-08-23 12:56:11 -07:00
Ryan Houdek a40526a541 Merge pull request #2977 from lioncash/pack
OpcodeDispatcher: Remove redundant moves from VPACKUSOP/VPACKSSOp
2023-08-23 12:33:05 -07:00
Lioncache ea9747289a OpcodeDispatcher: Remove redundant moves from remaining AVX ops
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 15:31:59 -04:00
Lioncache 735e2060a3 OpcodeDispatcher: Remove redundant moves from VPACKUSOP/VPACKSSOp
Zero-extension will occur automatically if necessary.
2023-08-23 15:09:57 -04:00
Ryan Houdek 86ef6fe48d Merge pull request #2976 from lioncash/mov
OpcodeDispatcher: Remove unnecessary moves from AVX move ops where applicable
2023-08-23 12:02:05 -07:00
Lioncache 8e7e91d61f OpcodeDispatcher: Remove redundant moves in VMOVLPOp
Zero-extension will automatically occur upon storing if necessary.
2023-08-23 14:23:57 -04:00
Lioncache bcba3700c8 OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
Zero-extension will automatically occur if necessary upon storing.

We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache 7d05797e82 OpcodeDispatcher: Remove redundant moves from VMOVHPOp
Zero-extension will automatically occur if necessary upon storing.
2023-08-23 14:14:35 -04:00
Lioncache c409ea78bc OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Ryan Houdek a62ba75ede Merge pull request #2975 from lioncash/scalar
Arm64/ConversionOps: Add scalar support to Vector_FToI
2023-08-23 10:57:38 -07:00
Lioncache 4a7ef3da13 OpcodeDispatcher: Remove unnecessary moves in AVXVectorRound
Zero-extension will occur automatically if necessary upon storing.
2023-08-23 13:41:56 -04:00
Lioncache 990b70dcd6 OpcodeDispatcher: Use scalar rounding for scalar round instructions 2023-08-23 13:34:01 -04:00
Lioncache 393cea2e8b Arm64/ConversionOps: Correct AdvSIMD round-to-nearest Vector_FToI case
This was previously using frinti, which uses the host rounding mode, rather
than round to nearest.
2023-08-23 13:27:44 -04:00
Ryan Houdek 6624f50abf Merge pull request #2974 from lioncash/extend
OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
2023-08-23 10:25:56 -07:00
Ryan Houdek cd1f401363 Merge pull request #2973 from lioncash/vfcmp
OpcodeDispatcher: Remove unnecessary moves in AVXVFCMPOp
2023-08-23 10:25:16 -07:00
Lioncache e89321dc60 Arm64/ConversionOps: Add scalar support to Vector_FToI
This can be used for the scalar conversions instead of always using the
vector variants.
2023-08-23 13:23:07 -04:00
Lioncache d99bcbf01b OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
Zero-extension will already occur if necessary upon storing.

Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Mai 0819338dbf Merge pull request #2970 from Sonicadvance1/optimize_pminmax
Arm64: Optimize VFMin/VFMax
2023-08-23 12:39:37 -04:00
Lioncache f516aed4b7 OpcodeDispatcher: Remove unnecessary moves in AVXVFCMPOp
Zero-extension will already occur if necessary upon storing.
2023-08-23 12:37:45 -04:00
Ryan Houdek 76430baf88 Merge pull request #2971 from lioncash/blend
OpcodeDispatcher: Remove redundant moves in AVX blend special cases
2023-08-22 21:42:07 -07:00
Lioncache 3858e4124b OpcodeDispatcher: Remove redundant moves in AVX blend special cases
Zero-extension will happen if necessary upon storing.
2023-08-23 00:08:33 -04:00
Mai 819fe110da Merge pull request #2967 from Sonicadvance1/optimize_storeelement
OpcodeDispatcher: Optimize MOVHP{S,D}
2023-08-23 00:01:25 -04:00
Ryan Houdek 5db5944ad2 InstCountCI: Update for min/max optimization 2023-08-22 21:00:27 -07:00
Ryan Houdek f0b1030e54 Arm64: Optimize VFMin/VFMax
We can be more optimal on the selects. This makes 3DNow! and SSE packed
min/max operations optimal.
2023-08-22 20:59:11 -07:00
Ryan Houdek adfd6787c0 Merge pull request #2969 from lioncash/insert
OpcodeDispatcher: Remove unnecessary moves from AVX inserts
2023-08-22 20:51:55 -07:00
Ryan Houdek b35ad8d8ed InstCountCI: Update for movhp{s,d} optimization 2023-08-22 20:42:24 -07:00
Ryan Houdek 0ee2579a5e OpcodeDispatcher: Optimize MOVHP{S,D}
Loads can turn in to element Loads.
Stores can turn in to element stores.

These four instruction variants are now optimal.
2023-08-22 20:42:24 -07:00
Ryan Houdek 6aa2cab41c IR: Implement support for vector store element
Matches ARM64 ST1 semantics
2023-08-22 20:42:24 -07:00
Mai bb2f7107cd Merge pull request #2963 from Sonicadvance1/optimize_loadelement
OpcodeDispatcher: Optimize MOVLP{S,D} loads
2023-08-22 23:41:59 -04:00
Lioncache c33f3ff8df OpcodeDispatcher: Remove unnecessary moves from AVX inserts
We already zero-extend on stores when necessary.
2023-08-22 23:29:40 -04:00
Ryan Houdek 5d44a445dd Merge pull request #2968 from lioncash/shift2
OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
2023-08-22 20:23:12 -07:00
Ryan Houdek 519c670374 InstCountCI: Update for load element optimization
Adds movhlps special case which was missed before.
2023-08-22 20:15:16 -07:00
Ryan Houdek de239cde67 OpcodeDispatcher: Optimize MOVLP{S,D} loads
This now uses the new load element IR operation and makes these
instructions optimal.

LRPCPC3 will introduce instructions in the future for TSO emulation to
help these operations, but that doesn't exist today.
2023-08-22 20:15:16 -07:00
Ryan Houdek 5e57ec94cf IR: Implement support for vector load element
Matches Arm64 LD1 semantics.
2023-08-22 20:15:16 -07:00
Lioncache 2f5fae7677 OpcodeDispatcher: Remove unnecessary moves from AVX register shifts
Zero-extension will occur automatically when necessary upon storing.
2023-08-22 23:06:53 -04:00
Ryan Houdek fb65fb29c7 Merge pull request #2966 from lioncash/shift
OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
2023-08-22 20:02:14 -07:00
Lioncache 8f8062eb4e OpcodeDispatcher: Remove redundant moves from AVX immediate shifts
These zero-extensions will occur automatically when applicable.
2023-08-22 22:50:10 -04:00
Ryan Houdek 36a54183f5 Merge pull request #2965 from lioncash/mov
OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
2023-08-22 19:38:46 -07:00
Lioncache e5f5629ffc OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Ryan Houdek 4443c667ec Merge pull request #2964 from Sonicadvance1/missed_optimal
InstCountCI: Fix some mislabeled instructions
2023-08-22 19:05:14 -07:00
Ryan Houdek ead141fd90 Merge pull request #2962 from lioncash/variable
OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
2023-08-22 18:52:47 -07:00
Ryan Houdek 2c64523317 InstCountCI: Fix some mislabeled instructions
These were all optimal. Fixed now.
2023-08-22 18:47:17 -07:00
Lioncache 2b071e282e OpcodeDispatcher: Remove unnecessary moves from AVXVariableShiftImpl
We already zero-extend on a store if necessary.
2023-08-22 21:20:34 -04:00
Ryan Houdek 14144523f7 Merge pull request #2961 from lioncash/minpos
OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
2023-08-22 18:19:46 -07:00
Ryan Houdek 1f2c5fc6c6 Merge pull request #2960 from lioncash/index
Arm64: Optimize SVE VInsElement
2023-08-22 18:19:13 -07:00
Mai 42200bf7b6 Merge pull request #2959 from Sonicadvance1/movlpd_store
X86Tables: Optimize MOVLPD stores
2023-08-22 21:02:23 -04:00
Lioncache fa17d9fae9 OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
We already do a zero-extend if necessary in StoreResult.

This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Lioncache 398a70312e Arm64: Optimize SVE VInsElement
This can be done without storing any data to memory and also
compressing the amount of instructions being used.

Thanks to @dougallj for the optimization suggestions.
2023-08-22 20:36:29 -04:00
Ryan Houdek 76afc653e8 InstCountCI: Update for movlpd store optimization 2023-08-22 17:34:15 -07:00
Ryan Houdek d3ed9766e8 X86Tables: Optimize MOVLPD stores
Just use the full register size and store the lower bits.
2023-08-22 17:33:34 -07:00
Ryan Houdek ed7f1b017d Merge pull request #2958 from Sonicadvance1/optimize_phminpos
OpcodeDispatcher: Optimize phminposuw
2023-08-22 16:58:27 -07:00
Ryan Houdek fb60f9e406 InstCountCI: Update for phminposuw optimization 2023-08-22 16:29:06 -07:00
Ryan Houdek c795d42d21 OpcodeDispatcher: Optimize phminposuw
I would now consider the XMM version of this to be optimal.

Thanks to @rygorous for giving the idea for how to optimize this!
2023-08-22 16:29:06 -07:00
Ryan Houdek bbf9cb9d52 IR: Implements new VRev32 and LoadNamedVectorConstant ops
VRev32 matches Arm64 semantics directly.
LoadNamedVectorConstant allows FEX to quickly load "named constants".
This will allow us to have specific hardcoded vector constant values
that we can load with a ldr(State)+ldr(Value) and will be more abused in
the future.
This also allows us to do a very simple optimization in the future where
we can optimize away redundant loads of these loads if they are used
multiple times in the same block. (Not implemented here).
2023-08-22 16:29:06 -07:00
Mai 6c7933e7b1 Merge pull request #2957 from Sonicadvance1/optimize_pfnacc
OpcodeDispatcher: Optimize PFNACC
2023-08-22 10:10:11 -04:00
Ryan Houdek 364f084604 Merge pull request #2956 from Sonicadvance1/optimize_hsubp
OpcodeDispatcher: Optimize hsubp
2023-08-21 20:47:50 -07:00
Ryan Houdek ffa8f1e3dc Merge pull request #2955 from lioncash/sign
OpcodeDispatcher: Remove redundant move from VPSIGN
2023-08-21 20:47:41 -07:00
Ryan Houdek be5d5b06f8 InstCountCI: Update for pfnacc 2023-08-21 20:38:25 -07:00
Ryan Houdek ad6738939b OpcodeDispatcher: Optimize PFNACC
Turns out this can be even more optimal.
2023-08-21 20:38:18 -07:00
Mai 9df94d8a93 Merge pull request #2954 from Sonicadvance1/optimize_pmuludq
OpcodeDispatcher: Optimize pmuludq
2023-08-21 23:33:45 -04:00
Ryan Houdek 9ae85b2251 InstCountCI: Update for hsubp 2023-08-21 20:24:37 -07:00
Ryan Houdek dcb3e4ee86 OpcodeDispatcher: Optimize hsubp
This makes the SSE version optimal.
This dramatically improves the AVX version as well.
2023-08-21 20:22:14 -07:00
Lioncache dbbe6288de OpcodeDispatcher: Remove redundant move from VPSIGN
StoreResult will already zero-extend if the vector is 128-bit.
2023-08-21 23:11:07 -04:00
Ryan Houdek de1f75f7a5 InstCountCI: Update for pmuludq 2023-08-21 20:08:19 -07:00
Ryan Houdek 1563398d2c OpcodeDispatcher: Optimize pmuludq
MMX version was already optimal, SSE version is now also.
AVX version is significantly improved.
2023-08-21 20:07:35 -07:00
Ryan Houdek 71984fc0ea Merge pull request #2953 from lioncash/scalar
OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
2023-08-21 19:51:23 -07:00
Lioncache 920a0fb132 OpcodeDispatcher: Remove redundant move in AVXVectorScalarALUOpImpl
Our store will already zero-extend if the vector is 128-bit.
2023-08-21 22:39:07 -04:00
Ryan Houdek 3c88671cca Merge pull request #2952 from lioncash/alu
OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
2023-08-21 19:26:22 -07:00
Lioncache ce8169794f OpcodeDispatcher: Remove redundant moves in AVXVectorALUOp
We already zero-extend on a store if we have 256-bit vectors and the stored
vector is 128-bit.
2023-08-21 22:13:29 -04:00
Mai 185e3bfcb6 Merge pull request #2950 from Sonicadvance1/optimize_pmaddwd
OpcodeDispatcher: Optimize pmaddwd
2023-08-21 21:39:30 -04:00
Mai 3c49b3238a Merge pull request #2949 from Sonicadvance1/optimize_phsub
OpcodeDispatcher: Optimize phsub
2023-08-21 21:39:02 -04:00
Ryan Houdek 2d7a3a578e Merge pull request #2931 from Sonicadvance1/optimize_psign
Optimize PSIGN and VBSL
2023-08-21 18:30:55 -07:00
Ryan Houdek 3a2a576c35 Merge pull request #2951 from lioncash/shift
OpcodeDispatcher: Handle zero immediate shifts better
2023-08-21 18:20:18 -07:00
Ryan Houdek fe4de26250 InstCountCI: Updates tests for pmaddwd optimization
MMX and SSE implementations are now optimal.
2023-08-21 17:55:20 -07:00
Ryan Houdek 869136b907 OpcodeDispatcher: Optimize pmaddwd
This is actually fairly trivial looking at it.
2023-08-21 17:53:32 -07:00
Lioncache af8b6766d8 OpcodeDispatcher: Handle zero immediate shifts better
In the SSE and lower cases, we don't need to do anything,
since the value is already in the destination.
2023-08-21 20:46:41 -04:00
Ryan Houdek d283d1ba11 InstCountCI: Update for phsub optimization
MMX and SSE now optimal
2023-08-21 17:38:24 -07:00
Ryan Houdek 30e9beba51 JIT64: Fixes i32v2 unzips
This has just been broken since it was implemented. Turns out we had
never used these with XMM operations before today.
2023-08-21 17:38:24 -07:00
Ryan Houdek 8ee6262e5e Merge pull request #2946 from Sonicadvance1/optimize_mpsad
OpcodeDispatcher: Optimizes mpsadbw
2023-08-21 17:33:03 -07:00
Ryan Houdek 637a5d3b18 InstCountCI: Updates results from mpsadbw optimization
Pretty sure the SSE versions are optimal implementations.
The AVX versions still have spurious moves that can probably get
removed.
2023-08-21 16:57:13 -07:00
Ryan Houdek a29076244f OpcodeDispatcher: Optimizes mpsadbw
Two optimizations here:
1) The final VInsElement was generating three instructions
   - This itself could have been change to vzip, which would have
     removed two instructions.
2) Optimize how the pairwise elements are calculated to shave one
   instruction off the calculation.
   - addp odd elements and even elemnts first
   - Then transpose those elements
   - Then use one final addp to generate the result in the correct
     order.

The ext+uabdl+addp pairs of operations could be reordered to shave off
one temporary register usage if we really care later.
2023-08-21 16:56:41 -07:00
Ryan Houdek 445c43792b OpcodeDispatcher: Optimize phsub
This was...surprisingly bad. I blame myself entirely.
2023-08-21 16:54:35 -07:00
Mai f4f9b20f32 Merge pull request #2948 from Sonicadvance1/fix_newline
InstCountCI: Add newline to end of file
2023-08-21 19:50:53 -04:00
Mai 34a7feffe8 Merge pull request #2947 from Sonicadvance1/optimize_pmaddubsw
OpcodeDispatcher: Optimize SSE/AVX pmaddubsw
2023-08-21 19:50:10 -04:00
Ryan Houdek b559982515 InstCountCI: Update for newlines 2023-08-21 16:26:46 -07:00
Ryan Houdek ca96a25a7a InstCountCI: Add newline to end of file
This way these endlines don't constantly keep getting toggled.
2023-08-21 16:26:20 -07:00
Ryan Houdek 3f82c8cfe5 InstCountCI: Update for pmaddubsw
Not quite optimal but a heck of a lot better.
2023-08-21 16:22:00 -07:00
Ryan Houdek b35f4df798 OpcodeDispatcher: Optimize SSE pmaddubsw
Can be slightly more optimal with a slightly change algorithm but will
require implementing some IR ops which can be put off. It's only about
an instruction savings.
2023-08-21 16:21:01 -07:00
Ryan Houdek e765bd8986 Merge pull request #2945 from lioncash/interpdata
Interpreter: Tie SSA data elements to supported vector width
2023-08-21 11:52:01 -07:00
Lioncache f1d020ce95 Interpreter: Tie SSA data elements to supported vector width
Now, if we ever increase our vector sizes, the allocated data elements
will follow suit without needing to remember to handle this as well.
2023-08-21 14:41:00 -04:00
Ryan Houdek 8b9ee997b2 Merge pull request #2944 from lioncash/interparray
Interpreter: Use alias for temporary vector data
2023-08-21 11:28:24 -07:00
Lioncache a9a7cbce21 Interpreter: Use alias for temporary vector data
Lets us extract the size into one location for easy size
changes in the future if necessary.
2023-08-21 14:12:55 -04:00
Ryan Houdek 084d102c9c Merge pull request #2943 from Sonicadvance1/wide_shifts
IR: Implements support for wide scalar shifts
2023-08-21 09:14:38 -07:00
Ryan Houdek 214dad25b5 InstCountCI: Updates instruction tables for wide shifts
Even on platforms without SVE these have improved slightly but the real
improvement comes from SVE.

Adds some new InstCountCI files for SVE128 enabled testing.

Also enables SVE128 in the VEX maps. Host features should probably
enable SVE128 when SVE256/AVX is enabled, but that isn't the case today.
2023-08-20 19:19:15 -07:00
Ryan Houdek a3bf952f2b IR: Implements support for wide scalar shifts
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.

With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).

This is a significant improvement even on platforms that only support
128-bit SVE.
2023-08-20 19:16:40 -07:00
Mai 5cc30bdf10 Merge pull request #2942 from Sonicadvance1/fix_nontelemetry_compilation
FEXInterpeter: Fixes compilation when telemetry is disabled
2023-08-20 17:16:13 -04:00
Ryan Houdek c3b4bfbc8d FEXInterpeter: Fixes compilation when telemetry is disabled
Oops, this has been broken for a while now.
2023-08-20 13:50:33 -07:00
Ryan Houdek b39e8ed5f0 InstCountCI: Update for improvements to psign/vbsl
Quite a few improvements here.
2023-08-20 13:46:01 -07:00
Ryan Houdek ac77986c44 IR: Implements support for saturating/rounding vector shifts 2023-08-20 13:38:23 -07:00
Ryan Houdek c5f5a03c68 OpcodeDispatcher: Optimize PSIGN
This dramatically improves the performance of the PSIGN instructions.
2023-08-20 13:38:23 -07:00
Ryan Houdek 1ed9ec63be Arm64: Optimize BSL when possible.
With ASIMD FEX would never optimize BSL out of fear if some registers
overlapped it would break things. So it had previously always moved to a
temporary first and then moved the result back out when done.

Now instead check upfront if any of the source registers overlap the
destination. If the destination register overlaps any of the three
sources we can bsl, bit, or bif depending on which register gets
overlapped.

Worst case the destination doesn't overlap any of the source registers
and still needs these moves.
2023-08-20 13:38:23 -07:00
Ryan Houdek a523858f66 Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek 9e4888c6a1 Merge pull request #2930 from Sonicadvance1/support_push
OpcodeDispatcher: Implement support for push IR operation
2023-08-20 10:25:26 -07:00
Ryan Houdek 277345d2a5 Merge pull request #2941 from lioncash/sha1
OpcodeDispatcher: Improve SHA1MSG1 output
2023-08-20 10:18:10 -07:00
Lioncache 8167626a07 x86_64/VectorOps: Properly handle VExtr element sizes other than bytes
We need to convert the index into a byte index.
2023-08-20 13:03:27 -04:00
Lioncache c3778a9729 OpcodeDispatcher: Improve SHA1MSG1 output
We can simplify these inserts down to a single EXT
2023-08-20 12:50:07 -04:00
Ryan Houdek 34722348e8 Merge pull request #2912 from Sonicadvance1/optimize_flag_clearing
OpcodeDispatcher: Minor optimization around clearing flags
2023-08-19 23:03:40 -07:00
Ryan Houdek 4286d44e92 Merge pull request #2940 from lioncash/crypto
Arm64/EncryptionOps: Use MOVI reg, #0 to zero vectors
2023-08-19 20:58:35 -07:00
Lioncache 1fbf193739 Arm64/EncryptionOps: Use MOVI reg, #0 to zero vectors
This is a little more optimal than XORing the vector by itself.
2023-08-19 23:20:36 -04:00
Ryan Houdek 8302ef8c22 InstCountCI: Update changed instructions due to flag clearing improvements
Some instructions reordered, but a bunch of operations had their number of instructions reduced as well
2023-08-19 20:19:39 -07:00
Ryan Houdek 92c3014aaa OpcodeDispatcher: Minor optimization around clearing flags
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.

Back to back bfi is actually less optimal due to dependency tracking.

With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
2023-08-19 20:14:37 -07:00
Mai 6960fca256 Merge pull request #2929 from Sonicadvance1/signaldelegator_getconfig
SignalDelegator: Allow getting the internal configuration
2023-08-19 23:12:06 -04:00
Mai affbcd2241 Merge pull request #2928 from Sonicadvance1/remove_x18_saving
Arm64Emitter: Stop saving and restoring platform register
2023-08-19 23:11:44 -04:00
Mai a2e5c231ae Merge pull request #2908 from Sonicadvance1/optimize_stc_clc
IR/ConstProp: Ensure that BFI with constant bitfields can optimize to Andn or Or
2023-08-19 23:10:48 -04:00
Ryan Houdek b973c193be Merge pull request #2939 from lioncash/round
OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
2023-08-19 19:29:01 -07:00
Ryan Houdek 4c409ea47d Merge pull request #2938 from lioncash/fcmp
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
2023-08-19 19:27:24 -07:00
Ryan Houdek ac53913c37 Merge pull request #2937 from lioncash/scalarunary
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
2023-08-19 19:22:10 -07:00
Ryan Houdek 2224c23c79 Merge pull request #2934 from lioncash/scalarfp
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
2023-08-19 19:20:20 -07:00
Ryan Houdek 6ce380d7a9 Merge pull request #2936 from lioncash/scalaralu
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorScalarALUOp
2023-08-19 19:09:06 -07:00
Lioncache 5152854b98 OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
2023-08-19 22:07:25 -04:00
Ryan Houdek 8dade7eea1 Merge pull request #2935 from lioncash/scalarfp2
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
2023-08-19 19:05:52 -07:00
Ryan Houdek add5baeba5 Merge pull request #2933 from lioncash/movss
OpcodeDispatcher: Remove some extraneous MOVs from VMOVSD/VMOVSS
2023-08-19 19:03:24 -07:00
Lioncache 6ba42e5cf1 OpcodeDispatcher: Eliminate redundant moves in {AVX}VectorRound
When dealing with scalar source registers, we can opt to not zero-extend
the vector and just perform the scalar operation and then insert the result.
2023-08-19 21:27:59 -04:00
Lioncache 343b00818d OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
We dealing with scalar vector sources, we don't need to zero-extend
the vector, and we can just use it as is.
2023-08-19 21:20:03 -04:00
Lioncache 09addb217a OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
2023-08-19 20:09:19 -04:00
Lioncache 4d1f002dea OpcodeDispatcher: Remove unnecessary moves in AVXVectorScalarALUOp
Same thing as the SSE variant, but for AVX.
2023-08-19 19:22:42 -04:00
Lioncache 1158ad7b2a OpcodeDispatcher: Remove unnecessary moves in VectorScalarALUOp
We can explicitly specify the vector width when working with a
vector source, so that we don't do any unnecessary zero-extending
on the element.
2023-08-19 19:15:13 -04:00
Lioncache 6907fdca6b OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
2023-08-19 18:46:08 -04:00
Lioncache c31329609f OpcodeDispatcher: Unify handling code for MOVSD and MOVSS
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
2023-08-19 17:44:45 -04:00
Lioncache 1fe8470933 OpcodeDispatcher: Remove extraneous moves from VMOVSS/VMOVSD xmm to mem case
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
2023-08-19 17:35:52 -04:00
Lioncache 99b5aaa426 OpcodeDispatcher: Remove extraneous moves in VMOVSS/VMOVSD register case
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
2023-08-19 17:35:12 -04:00
Ryan Houdek db60a2fd4b Merge pull request #2932 from lioncash/perm
OpcodeDispatcher: Handle broadcasting cases in VPERMQ/VPERMPD
2023-08-18 22:48:30 -07:00
Lioncache 84f228a75a x86_64/VectorOps: Simplify index handling in VInsElement
While we're in the area, we can simplify these cases down a little.
2023-08-19 01:28:45 -04:00
Lioncache 83a330b039 x86_64/VectorOps: Fix insertion bugs in VDupElement for 256-bit
Previously we weren't hitting this because we were never broadcasting
from the upper lane with VDupElement.
2023-08-19 01:21:12 -04:00
Lioncache bbed4d73ed OpcodeDispatcher: Improve VPERMQ/VPERMPD broadcast cases
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
2023-08-19 01:08:30 -04:00
Ryan Houdek a7e81ca731 InstCountCI: Update for push optimization
Some of these instructions improved a decent amount. Some even ending up
as being optimal.
2023-08-18 14:28:03 -07:00
Ryan Houdek 8b051b5e63 OpcodeDispatcher: Implement support for push IR operation
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.

While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
2023-08-18 14:19:16 -07:00
Ryan Houdek 1fdc4d2c62 IR: Implement support for a push operation
This is a bit of tricky operation where due to our our usage of SSA, the
incoming source isn't guaranteed to end its live-range at this
instruction.

This gives us a behaviour where to be optimal we need to take different
paths depending on if the incoming address register is the same as the
destination node.
Once we have form of RA constraints or non-SSA IR form that can
guarantee this restriction then this will go away.
2023-08-18 14:14:38 -07:00
Ryan Houdek ea5c67da80 SignalDelegator: Allow getting the internal configuration
Not used by FEX today but will be used by the WINE integration.
2023-08-18 11:56:52 -07:00
Ryan Houdek 0373826f46 Arm64Emitter: Stop saving and restoring platform register
FEX doesn't use the platform register on wine platforms so there is no
reason to save  and restore it.

On Linux we can still use it at some point but for now it isn't part of
our RA.
2023-08-18 11:49:44 -07:00
Ryan Houdek f912715df3 InstCountCI: STC and CLC is now optimal
Interestingly the segment register move instructions were previously
considered optimal on accident. They are actually optimal now which is
funny.
2023-08-18 11:41:21 -07:00
Ryan Houdek 7db2e487c3 IR/ConstProp: Remove some UBSAN behaviour
Changes the idiom used for constant mask generation to a ternary.
This pattern is definitely used elsewhere in code but we can get rid of
all instances here.
2023-08-18 11:41:11 -07:00
Ryan Houdek 6e5111b876 IR/ConstProp: Ensure that BFI with constant bitfields can optimize to Andn or Or
This optimizes the clc and stc instructions for flag setting and
clearing.
2023-08-18 11:33:11 -07:00
Ryan Houdek d502ad63f4 IR/ConstProp: Ensure ANDN is optimized 2023-08-18 11:27:06 -07:00
Ryan Houdek fc84f6b345 Merge pull request #2927 from bylaws/interrupt
FEXCore: Allow for interrupting the JIT on block entry
2023-08-18 06:14:24 -07:00
Billy Laws de63fd05d0 FEXCore: Allow for interrupting the JIT on block entry
This takes a similar approach to deferred signal handling and allows any given
thread to be interrupted while running JIT code by protecting the appropriate
page as RO. When the thread then enters a new block, it will try to acccess
that page and segfault. This is safer than just sending a signal to the thread
as that could stop in a place where JIT context couldn't be recovered correctly.
2023-08-18 05:58:51 -07:00
Ryan Houdek d3f0c7e969 Merge pull request #2925 from bylaws/winfile
Support for Config.json loading on WIN32
2023-08-18 05:00:22 -07:00
Ryan Houdek f1aa6208eb Merge pull request #2926 from bylaws/mingw
CMake: Add mingw toolchain file
2023-08-18 05:00:09 -07:00
Ryan Houdek b4d172649e Merge pull request #2924 from bylaws/logs
LogMan: Commonise log level to string conversion
2023-08-18 04:57:15 -07:00
Billy Laws 00556023c2 Remove unnecessary WIN32 file handling TODOs
With WOW, all allocations from 64-bit code use the full address space
and limiting is handled on the syscall thunk side so theres need to
worry about STL allocations stealing AS.
2023-08-18 04:37:40 -07:00
Billy Laws 5de0714766 FileLoading: Fix handling of non-existent files on WIN32 2023-08-18 04:37:40 -07:00
Billy Laws b862203491 Config: Add windows config loading support
This relies on wine's behaviour passing through linux paths and env vars,
so that the config in the user's home directory can be accessed outside
of the wine prefix.
2023-08-18 04:37:40 -07:00
Billy Laws bbfd15f801 LogMan: Commonise log level to string conversion 2023-08-18 04:36:31 -07:00
Billy Laws 0954c7eb9f AllocatorHooks: Add C++17 aligned new/delete functions 2023-08-18 04:32:16 -07:00
Billy Laws 8b2be809ff CI: Update to use mingw toolchain file 2023-08-18 04:31:41 -07:00
Billy Laws 193157812f CMake: Add mingw toolchain file 2023-08-18 04:31:41 -07:00
Ryan Houdek f09d9af3db Merge pull request #2922 from lioncash/psrld
OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
2023-08-17 17:05:35 -07:00
Ryan Houdek d19e2507e5 FEXCore: Adds telemetry around legacy segment register setting
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.

Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.

It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.

Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.

A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.

[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
   - 3.6 - Restricted Subset of Segmentation
      - `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
        registers; the base for CS, DS, ES, and SS is ignored for 32-bit
        mode, same as 64-bit mode (treated as zero).`
   - 4.2.17 - MOV to Segment Register
      - Will fault if SS is written (Breaking anything that writes to
        SS).
      - Will not fault if CS, DS, ES are written (Thus it sets the
        segment but gets ignored due to 3.6).
2023-08-17 17:00:41 -07:00
Lioncache 9e54ec2724 OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
2023-08-17 19:33:09 -04:00
Ryan Houdek 461ca6fe7c Merge pull request #2921 from lioncash/shift
OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
2023-08-17 16:03:16 -07:00
Lioncache 5a1f32c339 OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
PSLLIImpl already checks for and handles a shift value of zero.
2023-08-17 18:47:50 -04:00
Ryan Houdek a4a68b47ce Merge pull request #2920 from lioncash/ddup
OpcodeDispatcher: Improve VMOVDDUP output
2023-08-17 15:47:06 -07:00
Lioncache 01515cea2c OpcodeDispatcher: Improve VMOVDDUP output
We can make use of TRN1 here to collapse a bunch of these moves.
2023-08-17 18:30:44 -04:00
Ryan Houdek f70b6f37a2 Merge pull request #2919 from lioncash/sldup
OpcodeDispatcher: Improve output of {V}MOVSLDUP/{V}MOVSHDUP
2023-08-17 14:49:09 -07:00
Lioncache 764c844225 OpcodeDispatcher: Improve output of {V}MOVSHDUP 2023-08-17 17:01:50 -04:00
Lioncache 31719aac6a OpcodeDispatcher: Improve output of {V}MOVSLDUP 2023-08-17 17:01:46 -04:00
Ryan Houdek c9856daaee Merge pull request #2891 from alyssarosenzweig/move-fexcore
Move External/FEXCore/ to FEXCore/
2023-08-17 13:56:07 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00
Ryan Houdek 425b0347a9 Merge pull request #2918 from lioncash/quad
IR: Allow 128-bit broadcasts in VBroadcastFromMem
2023-08-17 11:32:04 -07:00
Lioncache e14e71aaff IR: Allow 128-bit broadcasts in VBroadcastFromMem
Now all vbroadcast implementations go down the more optimal path.

For non-SVE 128-bit cases where we only have 128-bit wide registers,
we behave like ld1rqb and just act as a normal 128-bit load for
interface convenience.
2023-08-17 14:20:12 -04:00
Ryan Houdek a9dea29f03 Merge pull request #2917 from lioncash/quad
ARMEmitter: Handle SVE load and broadcast quadword groups
2023-08-17 10:59:27 -07:00
Lioncache f97df2a40f ARMEmitter: Handle SVE load and broadcast quadword (scalar plus scalar) group 2023-08-17 13:32:20 -04:00
Lioncache 6f9bc1e2fe ARMEmitter: Handle SVE load and broadcast quadword (scalar plus imm) category 2023-08-17 13:32:17 -04:00
Ryan Houdek 49b8b7cd2c Merge pull request #2916 from Sonicadvance1/128bit_predicate
Arm64Emitter: Ensure that 128-bit predicate is generated with SVE
2023-08-17 09:56:56 -07:00
Ryan Houdek cfc6368064 Merge pull request #2914 from lioncash/broad
IR: Add VBroadcastFromMem opcode
2023-08-17 09:45:11 -07:00
Ryan Houdek 059c022255 Arm64Emitter: Ensure that 128-bit predicate is generated with SVE
In the case of running on a 128-bit SVE system this predicate wasn't
setup. Since we never had any predicate usage before this wasn't an
issue. Now that #2914 is using the 128-bit predicate we need to make
sure that we are generating it.
2023-08-17 09:37:55 -07:00
Lioncache 25708be807 IR: Add TSO handling to VBroadcastFromMem 2023-08-17 12:30:57 -04:00
Lioncache c8e3ca481f OpcodeDispatcher: Remove explicit zero-extending in VBROADCASTOp
Since the implementations zero the upper lanes when appropriate, we can
remove the unnecessary explicit move.
2023-08-17 12:17:24 -04:00
Lioncache 879bc5176e IR: Add VBroadcastFromMem opcode
Allows the implementations of the vbroadcast instructions to perform the
load and broadcast in one operation as opposed to doing the load and then
broadcast separately.

Notably, the broadcasting loads can also be used on systems that have SVE 128-bit
support as well, not only 256-bit.

On non-SVE systems, we use the equivalent AdvSIMD instructions.
2023-08-17 12:17:24 -04:00
Ryan Houdek 6d562f8b3b Merge pull request #2911 from Sonicadvance1/stop_abusing_orr
Arm64: Stop abusing orr in LoadConstant
2023-08-17 09:13:36 -07:00
Ryan Houdek 1029bb1fae Merge pull request #2910 from Sonicadvance1/minor_bfi_opt
Arm64: Optimize non-optimal BFI move case
2023-08-17 09:12:37 -07:00
Ryan Houdek 1343c14db0 Merge pull request #2909 from Sonicadvance1/optimize_clzero_clear
Arm64: Optimize CacheLine{Clear,Clean}
2023-08-17 09:10:45 -07:00
Ryan Houdek 6580796169 Merge pull request #2915 from lioncash/unused
OpcodeDispatcher: Remove unused variable in AVXVectorUnaryOpImpl
2023-08-17 08:44:18 -07:00
Lioncache 77c64285cb OpcodeDispatcher: Remove unused variable in AVXVectorUnaryOpImpl
Forgot to remove this when getting rid of the unnecessary
explicit zero-extending behavior
2023-08-17 11:26:41 -04:00
Ryan Houdek 9f3730f251 Merge pull request #2913 from Sonicadvance1/string_view_self
Filemanagement: Optimize GetSelf using a string_view
2023-08-17 06:32:51 -07:00
Ryan Houdek 2be8e22e14 Filemanagement: Optimize GetSelf using a string_view
Instead of allocating a temporary copy of the string, return a view of
it instead. Should improve the performance of system calls that take
file paths. Since it was allocating a string for every single syscall
that uses them in this case.
2023-08-16 21:39:05 -07:00
Ryan Houdek fe37c89109 FEXCore/Config: Stop making temporary string copies
For config values that were string objects we were unnecessary creating
copies each time the string was accessed.

Convert the () operator over to returning a reference.
2023-08-16 21:35:13 -07:00
Ryan Houdek d11563f48f InstCountCI: Update instructions for data movement
Instruction counts don't change at all here, just the instructions being
used.
2023-08-16 19:36:12 -07:00
Ryan Houdek 23fd79a3b3 Arm64: Stop abusing orr in LoadConstant
The current implementation uses orr excessively. This has FEX missing
hardware optimization opportunities where some CPU cores will zero-cycle
move constants that fit in to the 16-bits of movz/movk.

First evaluate up front if the number of 16-bit segments is > 1, in
those cases we should check if it is a bitfield that can be moved in one
instruction with orr.

After that point we will use movz for 16-bit constant moves.

Additionally this optimizes the case where a constant of zero is loaded
to be a `mov <reg>, zr` which gets renamed in most hardware.
2023-08-16 19:35:15 -07:00
Ryan Houdek 13a4009bd3 InstCountCI: Update changed operations due to bfi operation
No instruction count changes here, just moving from lsr to mov.
2023-08-16 14:37:36 -07:00
Ryan Houdek a3b40c37c2 Arm64: Optimize non-optimal BFI move case
Commonly we are doing a BFI into a 32-bit register, which is hitting the
ubfx (lsr alias) path.

In the case of 32-bit destination we can also do a regular move, which
will take advantage of CPU's rename functionality and give a minor speed
boost.
2023-08-16 14:35:41 -07:00
Ryan Houdek 6cb0f52e94 Merge pull request #2907 from Sonicadvance1/fix_dumpir_bug
FEXCore/IR: Fixes bug in IRDumper without specification
2023-08-16 14:24:33 -07:00
Ryan Houdek f283ba4b05 InstCountCI: clwb/clfush/clflushopt is now optimal
One or two instructions depending.
2023-08-16 14:20:22 -07:00
Ryan Houdek 4522a766e0 Arm64: Optimize CacheLine{Clear,Clean}
When the cacheline size matches the expected x86 cacheline size then we
can remove the spurious move + add.
2023-08-16 14:20:22 -07:00
Ryan Houdek dc0cf98a81 Merge pull request #2898 from Sonicadvance1/instcount_asm
InstCountCI: Support encoding expected Arm64 ASM in JSON
2023-08-16 13:51:43 -07:00
Ryan Houdek fc12958095 FEXCore/IR: Fixes bug in IRDumper without specification
Didn't notice this in the previous PR, When DUMPIR=stderr without and
selection of where to place it in PASSMANAGERDUMPIR it was supposed to
put the dumper at the end of the passes.

We need to make sure that it it placed at the end of the passes rather
than current `it`.
2023-08-16 13:51:03 -07:00
Ryan Houdek fd40e058e8 InstCountCI: Update tests with inline asm
This will result in a decent amount of data churn but it's all
automated so it is a non-issue.
2023-08-16 13:38:45 -07:00
Ryan Houdek 8706f895d0 InstCountCI: Sanitize out vixl address calculations 2023-08-16 13:38:45 -07:00
Ryan Houdek 9753ebdebc InstCountCI: Support encoding expected Arm64 ASM in JSON
This will allow investigating the Arm64 directly next to the test, plus
publicly linking directly to badly behaving tests.

Perfect for nerdsniping implementations.
2023-08-16 01:50:06 -07:00
Mai f95ef7092c Merge pull request #2905 from Sonicadvance1/instcountci_updatetests
InstCountCI: Update tests from actual ARM64 device
2023-08-16 04:13:05 -04:00
Ryan Houdek d4d5bc9635 InstCountCI: Update tests from actual ARM64 device
Previous numbers were from the simulator which were slightly different.
2023-08-15 14:39:37 -07:00
Mai df3d4efc80 Merge pull request #2904 from Sonicadvance1/instcountci_only_arm
GIthub: Only enable InstCountCI on an ARM platform
2023-08-15 17:33:10 -04:00
Mai dac220a6ff Merge pull request #2903 from Sonicadvance1/instcountci_rng
InstCountCI: Adds RNG support
2023-08-15 17:21:03 -04:00
Ryan Houdek 9ba7f2fd0e InstCountCI: Fixes instcountci_tests depends 2023-08-15 14:14:24 -07:00
Ryan Houdek 1441cb76b9 HostFeatures: Adds support for overriding ARMv8.1 LSE atomics
Always enable it on the InstCountCI.
2023-08-15 14:12:27 -07:00
Ryan Houdek 2d20513e34 GIthub: Only enable InstCountCI on an ARM platform
simulator generates some instruction count differences that we don't
care about. Just run this on ARM platforms only instead.
2023-08-15 14:12:27 -07:00
Mai da334fe7d5 Merge pull request #2901 from Sonicadvance1/instcountci_only_disassembler
InstCountCI: Disables tests with unsupported configurations
2023-08-15 16:01:47 -04:00
Ryan Houdek 04f1f073c8 InstCountCI: Adds RNG support
Some instructions in the SecondaryGroup need RNG support for testing.
2023-08-15 12:58:53 -07:00
Ryan Houdek 0e52158cef Merge pull request #2902 from lioncash/unary
OpcodeDispatcher: Eliminate unnecessary moves in AVXVectorUnaryOpImpl
2023-08-15 12:54:47 -07:00
Lioncache 17956eac5f OpcodeDispatcher: Eliminate unnecessary moves in AVXVectorUnaryOpImpl
We no longer need to do any manual zero-extending here, since this
will occur automatically on hardware with SVE when 128-bit AdvSIMD
is used.
2023-08-15 15:43:34 -04:00
Ryan Houdek 6ad053a1e6 Merge pull request #2900 from lioncash/sqrt
Arm64/VectorOps: Remove redundant move in VFRSqrt SVE path
2023-08-15 12:39:02 -07:00
Ryan Houdek b18e5e2f63 InstCountCI: Disables tests with unsupported configurations
Need to have the vixl disassembler enabled for instcountci.

Also need to make sure the host platform is using the ARM64 JIT.
2023-08-15 12:27:32 -07:00
Lioncache 2708374d95 Arm64/VectorOps: Remove redundant move in VFRSqrt SVE path
We can perform the SQRT first and then broadcast 1.0 into the destination
since all the intermediary work is done, meaning we don't have to worry
about Dst and Vector aliasing one another.
2023-08-15 15:22:21 -04:00
Mai 3c99fb84e2 Merge pull request #2899 from Sonicadvance1/instcountci_fix_asm_name
InstCountCI: Ensure output nasm name doesn't conflict
2023-08-15 14:44:44 -04:00
Ryan Houdek 710a3928ff Merge pull request #2897 from lioncash/broad
ARMEmitter: Handle SVE load and broadcast element group
2023-08-15 11:36:20 -07:00
Ryan Houdek ac1d2ec1d9 InstCountCI: Ensure output nasm name doesn't conflict
Pretty sure this is why CI is unhappy. If a test in a different file has
the same name then it is highly likely to conflict when nasm is
generating files and will overwrite and erase, causing CI to break.

Include the incoming json filename as part of the asm keys so it can't
conflict here.
2023-08-15 11:29:39 -07:00
Lioncache 6acce60855 ARMEmitter: Handle SVE load and broadcast element group
These can be used to improve vbroadcast implementations from
doing a mem load+dup in the non-GPR case into just directly
loading into the destination.
2023-08-15 13:47:12 -04:00
Mai 63f28eae4f Merge pull request #2896 from Sonicadvance1/primary_table
InstCountCI: Adds primary tables
2023-08-15 13:07:35 -04:00
Ryan Houdek eda67eb0a5 Merge pull request #2895 from lioncash/scalar
ARMEmitter: Handle load/store multiple structures (scalar plus scalar) groups
2023-08-15 09:12:40 -07:00
Ryan Houdek 21cac6ef0b InstCountCI: Adds primary tables
Surprisingly few instructions are optimal here.
This is all the instruction tables completed now!
2023-08-15 09:11:39 -07:00
Ryan Houdek 1193b55150 InstCountCI: Script auto change line 2023-08-15 09:11:29 -07:00
Ryan Houdek 27a53280bd InstCountCI: Fixes bitness in script 2023-08-15 09:11:07 -07:00
Mai 135b9ac425 Merge pull request #2894 from Sonicadvance1/instcount_secondary_prefixes
InstCountCI: Adds secondary prefix tables
2023-08-15 10:22:28 -04:00
Lioncache 81115f64f6 ARMEmitter: Handle SVE Store Multiple Structures (scalar plus scalar) 2023-08-15 10:18:37 -04:00
Lioncache 0176efa3bb ARMEmitter: Handle SVE Load Multiple Structures (scalar plus scalar) group 2023-08-15 10:01:15 -04:00
Ryan Houdek 2508274ddb InstCountCI: Adds secondary prefix tables
Most of these 128-bit vector ops are looking pretty good. Scalar and
edge case ops aren't always optimal though.
2023-08-15 06:57:47 -07:00
Mai 24d01cd8d2 Merge pull request #2893 from Sonicadvance1/instcount_secondary
InstCountCI: Adds secondary tables
2023-08-15 04:24:40 -04:00
Ryan Houdek 2fb72f822d InstCountCI: Adds secondary tables
A decent number of instructions that are optimal but still quite a lot
that aren't.
2023-08-14 16:04:31 -07:00
Ryan Houdek a2c2b042b2 CodeSizeValidation: Fixes nullptr dereference 2023-08-14 16:04:20 -07:00
Ryan Houdek 398e76be89 X86Tables: Fixes typo 2023-08-14 16:04:05 -07:00
Ryan Houdek 3885bc42f6 Merge pull request #2892 from Sonicadvance1/minor_config_changes
Config: Minor changes
2023-08-14 15:10:42 -07:00
Ryan Houdek c5439b294c TestHarnessRunner: InitializeConfig paths
Just removes an erroneous message in the TestHarnessRuner on each
execution.
2023-08-14 12:37:35 -07:00
Ryan Houdek f248e7f3e7 Config: If DumpIR is enabled, default enable a passmanager option
If DumpIR is enabled but the PassManagerDumpIR option isn't enabled then
this currently does nothing.

As a convenience, enable dumping the final optimized IR if an option
hasn't been specified.
2023-08-14 12:29:56 -07:00
Ryan Houdek e51606c669 Config: Fixes mixup in PassManagerDumpIR
The opt and pass options were inverted in PassManager.
Renames the enum to make this more clear.
2023-08-14 12:28:37 -07:00
Ryan Houdek 648d8aeb65 Config: Adds missing server option to DumpIR description
This was accepted but I failed to describe it when added.
2023-08-14 12:22:35 -07:00
Ryan Houdek 112c463655 Config: Ensure OutputLog to server doesn't try to expand path
"server" isn't a path, this was missed when it was added.
2023-08-14 12:20:58 -07:00
Ryan Houdek 2f0c690d54 Merge pull request #2890 from alyssarosenzweig/constprop/set-but-not
ConstProp: Fix set-but-not-used mask variable
2023-08-14 10:17:42 -07:00
Alyssa Rosenzweig 7ecbbd6c04 ConstProp: Fix set-but-not-used mask variable
I think this was the intended logic?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-14 11:59:59 -04:00
Mai 103e6044fa Merge pull request #2889 from Sonicadvance1/instcount_primarygroup
InstCountCI: Adds Primary group tables
2023-08-13 23:07:25 -04:00
Mai e334278f80 Merge pull request #2887 from Sonicadvance1/instcount_vex_map3
InstCountCI: Adds VEX map3 tables
2023-08-13 23:06:55 -04:00
Mai 7fe2d3b305 Merge pull request #2888 from Sonicadvance1/instcount_vex_map_group
InstCountCI: Adds VEX map group tables
2023-08-13 23:06:06 -04:00
Ryan Houdek 51d4bd5f05 InstCountCI: Adds Primary group tables
A lot of things fairly close to optimal but not quite there.
2023-08-13 17:26:09 -07:00
Ryan Houdek 2a871957f8 InstCountCI: Adds VEX map3 tables
Pretty much no instructions are optimal here. 64-bit RORX is the
outlier.
2023-08-13 13:00:00 -07:00
Ryan Houdek d9266167e1 InstCountCI: Adds VEX map group tables
No instructions are optimal in this table.
2023-08-13 12:59:11 -07:00
Ryan Houdek cbfea75037 Increase maximum instruction name to 128-bytes
Some long name instructions are going beyond 32-bytes. Obviously not
hurting to encode a few additional bytes.
2023-08-13 12:32:14 -07:00
Ryan Houdek ef3887ca4f X86Tables: Fixes typo in VEX table 2023-08-13 12:32:14 -07:00
Mai 871434dab9 Merge pull request #2886 from Sonicadvance1/instcount_vex_map2
InstCountCI: Adds FEX map2 tables
2023-08-13 07:42:01 -04:00
Ryan Houdek 33b84337ea InstCountCI: Adds FEX map2 tables
A handful of instructions are optimal, a few random ones are close, but
a lot aren't close.
2023-08-13 04:29:19 -07:00
Ryan Houdek 1e2b5ecd5d InstCountParser: Allow the ability to skip instructions 2023-08-13 04:29:07 -07:00
Mai a4647982cd Merge pull request #2885 from Sonicadvance1/vex_map1
InstCountCI: Adds VEX map1 tables
2023-08-13 07:05:52 -04:00
Ryan Houdek 047bbca5bb InstCountCI: Adds VEX map1 tables
None of the instructions here are optimal, a couple are close.

Will be splitting each of the map tables in to their own json files
since each one can get fairly large.
2023-08-13 02:15:53 -07:00
Mai 1146c42044 Merge pull request #2884 from Sonicadvance1/instcount_x87
InstCountCI: Adds x87 table
2023-08-12 11:12:01 -04:00
Ryan Houdek fc85431a0e InstCountCI: Adds x87 table
Most of these instructions aren't optimal.
Even outside of the ones that jump out of the JIT they have a issues in
most cases.
2023-08-12 06:53:30 -07:00
Mai a15934bd55 Merge pull request #2883 from Sonicadvance1/instcount_h0f38
InstCountCI: Adds H0F38 table
2023-08-12 09:05:27 -04:00
Ryan Houdek e1ef84f933 InstCountCI: Adds H0F38 table
Some of these instructions haven't been fully audited, which is why they
are classified as Unknown.
2023-08-12 05:20:44 -07:00
Ryan Houdek 139dd4ccbc Merge pull request #2882 from lioncash/adr
ARMEmitter: Handle SVE ADR
2023-08-11 20:52:42 -07:00
Lioncache 8fd810c3c0 ARMEmitter: Migrate adr off SVEMemOperand
We need to move the modifier enum out of the SVEMemOperand class
since it's also used with adr. Plus, this can also be convenient
not being tied down to the class itself.

This also makes accessing modifiers less noisy, since the class
2023-08-11 22:17:46 -04:00
Lioncache 068db933bf ARMEmitter: Handle SVE ADR 2023-08-11 19:50:21 -04:00
Ryan Houdek 72357e50a2 Merge pull request #2881 from lioncash/imm8
ARMEmitter: Handle SVE CPY (immediate)
2023-08-11 15:50:39 -07:00
Lioncache 0bf74a1f3e ARMEmitter: Use signed imm8 handler with dup_imm
Lets us deduplicate the behavior used for dup and cpy.
2023-08-11 18:30:45 -04:00
Lioncache 78f06c7fcb ARMEmitter: Handle SVE CPY (immediate)
Also adds the relevant aliases.
2023-08-11 18:24:26 -04:00
Ryan Houdek 6d1fcfce09 Merge pull request #2877 from Sonicadvance1/classification_adds
InstructionCountCI: Adds three more instruction tables
2023-08-11 15:09:29 -07:00
Ryan Houdek 5a0a6dd0ca Merge pull request #2880 from lioncash/vfp
ARMEmitter: Migrate off vixl float utils
2023-08-11 14:35:23 -07:00
Mai da17e24996 Merge pull request #2879 from Sonicadvance1/fix_adcx
FEXCore: Fixes bug with 32-bit adcx
2023-08-11 17:27:00 -04:00
Lioncache aef0795dc8 ARMEmitter: Make FloatToEquivalentUInt a little more robust
Rather than compare sizes, we should be comparing the types directly,
prevents any shenanigans from happening if interface changes occur to
Float16.
2023-08-11 17:16:14 -04:00
Lioncache f63498a558 ARMEmitter: Migrate off vixl float utils 2023-08-11 17:12:34 -04:00
Ryan Houdek 9aa3fde174 FEXCore: Fixes bug with 32-bit adcx
When a 32-bit adcx instruction was encountered, it was getting treated
as a 16-bit adcx instruction instead. This is because of the 0x66 prefix
required to handle this instruction.

Adds a unit test to ensure it doesn't break again.
2023-08-11 14:09:31 -07:00
Ryan Houdek 8fce13386a Merge pull request #2878 from lioncash/fcpy
ARMEmitter: Handle SVE FCPY (predicated)
2023-08-11 13:58:31 -07:00
Lioncache 247c7ce784 ARMEmitter: Handle SVE FCPY (predicated)
While we're at it, we can reduce our dependence on vixl's utils by
implementing our own based off the pseudocode of VFPExpandImm.
2023-08-11 16:18:20 -04:00
Ryan Houdek 7e12b056dc InstructionCountCI: Adds three more instruction tables
3DNow! table, H0F3A table, and SecondaryModRM table.

H0F3A table skips the SSE4.2 string instructions because they're
nightmares for now.
2023-08-11 11:13:42 -07:00
Ryan Houdek e8e52af2e8 CodeSizeValidation: Adds support for overriding CLZero support
Vixl simulator by default doesn't support this.
2023-08-11 11:12:46 -07:00
Ryan Houdek 6c7371af60 Merge pull request #2872 from Sonicadvance1/instruction_count_ci
FEX: Adds instruction count CI
2023-08-11 11:12:00 -07:00
Ryan Houdek acc7f2fa8f FEX: Adds instruction count CI
Implements CI for tracking instruction counts for generate blocks of
code when transforming from x86 to ARM64 assembly.

This will end up encompassing every instruction in our instruction
tables similarly to how our assembly tests try to test everything in our
instruction tables.

Incidentally, the data for this CI is generated using our assembly
tests. By enabling disassembly and instruction stats when executing a
suite of instructions, this gives the stats that can be added to a json
file.

The current implementation only implements the SecondGroup table of
instructions because it is a relatively small table and has known
inefficiencies in the instruction implementations. As this gets merged I
will be adding more tables of instructions to additional json files for
testing.

These JSON files will support adjusting CPU features regardless of the
host features so it can test implementations depending on different CPU
features. This will let us test things like one instruction having
different "optimal" implementations depending on if it supports SVE128,
SVE256, SVEI8MM, etc.

This initial instruction auditing is what found the bug in our vector
shift instructions by size of zero. If inspecting the result of the CI
run, you can tell that these instructions still aren't "optimal" because
they are doing loads and stores that can be eliminated.

The "Optimal" in the JSON is purely for human readable and grepping
ability to see what is optimal versus not. Same with the "Comment"
section.

According to my auditing spreadsheet, the total number of instructions
that will end up in these json files will be about 1000, but we will
likely end up with more since there will be edge cases that can be more
optimal depending on arguments.
2023-08-11 09:10:36 -07:00
Ryan Houdek b8b4dd8008 FEXCore/Utils: Add the ability to write a fextl::string 2023-08-11 09:08:54 -07:00
Ryan Houdek ec8855f8fb Arm64: Consolidate simulator and diassembler code in to Arm64Emitter
This was confusingly split between Arm64Emitter, Arm64Dispatcher, and
Arm64JIT.

- Arm64JIT objects were unnecessary and free to be deleted.
- Arm64Dispatcher simulator and decoder moved to Arm64Emitter
- Arm64Emitter disassembler and decoder renamed
  - Dropped usage of the PrintDisassembler since it is hardcoded to go
    through a FILE* type
  - We instead want its output to go through LogMan, which means using a
    split Decoder+Disassembler object pair.
  - Can't reuse the object from the vixl simulator since the simulator
    registers the decoder as a visitor, causing the simulator to execute
    while disassembling instructions if reused.
- Disassembly output for blocks and dispatcher now output through Logman
  - Blocks wrapped in Begin/End text for tracking purposes for CI.
2023-08-11 08:05:10 -07:00
Ryan Houdek 969ad9b3b0 Merge pull request #2869 from Sonicadvance1/sve_128bit_ci
Github: Adds a CI runner for 128-bit SVE testing
2023-08-11 07:37:47 -07:00
Ryan Houdek 5de7eeea20 Merge pull request #2876 from lioncash/comment
ARMEmitter: Remove resolved TODO comment
2023-08-11 07:22:19 -07:00
Ryan Houdek 0109e88082 Merge pull request #2875 from lioncash/ff
ARMEmitter: Handle contiguous first fault load (scalar plus scalar) group
2023-08-11 07:09:32 -07:00
Lioncache 73288f377f ARMEmitter: Remove resolved TODO comment
I forgot to remove this after implementing the normal gather instruction handling.
2023-08-11 10:03:10 -04:00
Lioncache 0aaa9503c9 ARMEmitter: Add missing ld1w (scalar plus scalar) tests
ld1sw was mistakenly tested twice.

Also groups the tests by data sizes.
2023-08-11 09:51:04 -04:00
Lioncache 14cc23b6c3 ARMEmitter: Handle contiguous first fault load (scalar plus scalar) group
Adds the only missing implementation category for the first-faulting loads,
making the interface more consistent.
2023-08-11 09:46:40 -04:00
Ryan Houdek 5404dba360 Github: Adds a CI runner for 128-bit SVE testing
We don't currently have a device in CI that can run SVE with 128-bit
width registers. Until we have a device with this, make sure the vixl
simulator is also running the ASM tests in this width.
2023-08-10 22:27:59 -07:00
Ryan Houdek 833c07e9e2 CoreState: Zero initialize some important members
This was causing test failure locally where some values were set to
uninitialized data. Ensure that gregs, YMM, and MMX registers are all
zero initialized.
2023-08-10 22:27:59 -07:00
Ryan Houdek 186ec201aa Config: Stop passing a temporary std::string_view outside of scope
Was causing strenum variables to be parsed, then leaving scope would
break the string.
2023-08-10 22:27:59 -07:00
Ryan Houdek 887c47c451 Config: Adds an option to override SVE width for CI 2023-08-10 21:25:57 -07:00
Ryan Houdek 0f3460e025 Config: Fixes typo in HostFeatures disable{sve,avx} 2023-08-10 21:25:57 -07:00
Ryan Houdek fadba9a3e1 External: Update vixl
Fixes simulator bug
2023-08-10 21:25:57 -07:00
Ryan Houdek 9d26af95ab Merge pull request #2873 from neobrain/refactor_warning_fixes
Various warning fixes
2023-08-10 16:58:14 -07:00
Tony Wasserka aed4dda3e4 Arm64: Remove unused function 2023-08-10 18:45:14 +02:00
Tony Wasserka e0d21e61cc Syscalls: Fix warnings about unused variables in Release builds 2023-08-10 18:45:14 +02:00
Tony Wasserka 45d0f0d349 ARMEmitter: Fix warnings about unused variables in Release builds 2023-08-10 18:45:14 +02:00
Tony Wasserka f1cc76614b Include VIXL as a system library
This suppresses warnings from VIXL headers.
2023-08-10 18:45:14 +02:00
Ryan Houdek 099f29f1ed Merge pull request #2871 from Sonicadvance1/fix_stats_missing_member
FEXCore: Fixes Arm64 stats disassembly
2023-08-10 06:23:02 -07:00
Ryan Houdek b1a3f82923 FEXCore: Fixes Arm64 stats disassembly
Requires the IR headerop to house the number of host instructions this
code is translating for the stats.

Fixes compiling with disassembly enabled, will be used with the
instruction count CI.
2023-08-10 03:23:25 -07:00
Ryan Houdek 6f4a23dd15 Merge pull request #2870 from lioncash/indexed
ARMEmitter: Handle SVE FP multiply-add long groups
2023-08-09 21:35:33 -07:00
Ryan Houdek f3182036bc Merge pull request #2867 from Sonicadvance1/dummy_thin_handlers
FEX: Create a CommonTools static library
2023-08-09 21:34:46 -07:00
Lioncache 444961ad79 ARMEmitter: Handle SVE FP multiply-add long group 2023-08-09 15:20:04 -04:00
Lioncache 48a3271fbc ARMEmitter: Handle SVE FP multiply-add long (indexed) group 2023-08-09 15:19:49 -04:00
Mai ea8fbc61c2 Merge pull request #2868 from Sonicadvance1/irdumper_passmanager
IR: Adds Option to run the IRDumper with more configurations
2023-08-09 10:28:25 -04:00
Ryan Houdek 35e97ec9bc IR: Adds Option to run the IRDumper with more configurations
This is incredibly useful and I find myself hacking this feature in
every time I am optimizing IR. Adds a new configuration option which
allows dumping IR at various times.

Before any optimization passes has happened
After all optimizations passes have happened
Before and After each IRPass to see what is breaking something.

Needs #2864 merged first
2023-08-09 05:58:20 -07:00
Ryan Houdek 53ac8abce9 Merge pull request #2863 from Sonicadvance1/stats
Arm64: Adds stats to the disassembly
2023-08-09 04:06:22 -07:00
Ryan Houdek fe351353f6 Merge pull request #2865 from Sonicadvance1/first_sve_opt
Arm64: Implement first SVE-128bit optimization
2023-08-09 04:06:05 -07:00
Ryan Houdek a23cb0447b Arm64: Implement first SVE-128bit optimization
This is a /very/ simple optimization purely because of a choice that ARM
made with SVE in latest Cortex.

Cortex-A715:
   - sxtl/sxtl2/uxtl/uxtl2 can execute 1 instruction per cycle.
   - sunpklo/sunpkhi/uunpklo/uunpkhi can execute 2 instructions per cycle.

Cortex-X3:
   - sxtl/sxtl2/uxtl/uxtl2 can execute 2 instruction per cycle.
   - sunpklo/sunpkhi/uunpklo/uunpkhi can execute 4 instructions per cycle.

This is fairly quirky since this optimization only works on SVE systems
with 128-bit Vector length. Which since it is all of the current
consumer platforms, it will work.
2023-08-09 03:51:57 -07:00
Ryan Houdek f2aa2ce4bb Arm64: Rename HostSupportsSVE
We need to know the difference between the host supporting SVE with
128-bit registers versus 256-bit registers. Ensure we know the
difference.

No functional change here.
2023-08-09 03:51:56 -07:00
Ryan Houdek cf93652708 Config: Adds support for overriding host features
This allows use to both enable and disable regardless of what the host
supports. This replaces the old `EnableAVX` option.

Unlike the old EnableAVX option which was a binary option which could
only disable, each of these options are technically trinary states.
Not setting an option gives you the default detection, while explicitly
enabling or disabling will toggle the option regardless of what the host
supports.

This will be used by the instruction count CI in the future.
2023-08-09 03:51:37 -07:00
Ryan Houdek eaed5c4704 Merge pull request #2862 from Sonicadvance1/optimize_vector_zero
ARM64: Optimize vector zeroing
2023-08-09 03:51:04 -07:00
Mai c77ed78f5a Merge pull request #2861 from Sonicadvance1/fix_vector_shift_by_zero
FEXCore: Fixes vector shifts by zero
2023-08-09 05:52:10 -04:00
Ryan Houdek 348844a95b FEX: Create a CommonTools static library
Moves the dummy handlers over to this library. This will end up getting
used for more than the mingw test harness runner once the instruction
count CI is operational.
2023-08-09 02:27:13 -07:00
Ryan Houdek e8fb322025 unittests: Adds tests for vector shifts with zero immediate
To ensure FEX doesn't encounter the encoding bug again.
2023-08-09 02:16:17 -07:00
Ryan Houdek 5f0efda8fe ARM64: Fixes shift by immediate zero
These would emit invalid instructions in most cases. Turn in to a move
or a no-op if the shift is zero.
2023-08-09 02:16:17 -07:00
Ryan Houdek d198d701aa OpcodeDispatcher: Fixes vector shifts by immediate zero
pslldq logic was wrong in the case of zero shift.
The rest should just return their source in the case of zero shift.
2023-08-09 02:16:17 -07:00
Mai c4c7620ed5 Merge pull request #2866 from Sonicadvance1/remove_unnecessary_loadconstant
Arm64: Remove erroneous LoadConstant
2023-08-09 05:10:11 -04:00
Ryan Houdek bf5719770e Arm64: Remove erroneous LoadConstant
This was a debug LoadConstant that would load the entry in to a temprary
register to make it easier to see what RIP a block was in.

This was implemented when FEX stopped storing the RIP in the CPU state
for every block. This is now no longer necessary since FEX stores the
in the tail data of the block.

This was affecting instructioncountci when in a debug build.
2023-08-08 22:56:36 -07:00
Ryan Houdek 0f6a268243 Arm64: Adds stats to the disassembly
I use this locally when looking for optimization opportunities in the
JIT.
The instruction count CI in the future will use this as well.
Just get it upstreamed right away.
2023-08-08 22:28:52 -07:00
Ryan Houdek e0461497a0 ARM64: Optimize vector zeroing
`eor <reg>, <reg>, <reg>` is not the optimal way to zero a vector
register on ARM CPUs. Instead we should move by constant or zero
register to take advantage of zero-latency moves.
2023-08-08 22:24:11 -07:00
412 changed files with 72100 additions and 4493 deletions

No files matched your search

+109
View File
@@ -0,0 +1,109 @@
name: Instruction Count CI run
on:
push:
branches:
- main
pull_request:
branches:
- main
env:
# Customize the CMake build type here (Release, Debug, RelWithDebInfo, etc.)
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64]]
fail-fast: false
steps:
- uses: actions/checkout@v3
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS_PATH=$HOME/Rootfs/" >> $GITHUB_ENV
echo "FEX_ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
echo "ROOTFS=$HOME/Rootfs/" >> $GITHUB_ENV
- name: Update RootFS cache
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
run: $GITHUB_WORKSPACE/Scripts/CI_FetchRootFS.py
- name : submodule checkout
# Need to update submodules
run: |
git submodule sync --recursive
git submodule update --init --depth 1
- name: Clean Build Environment
run: rm -Rf ${{runner.workspace}}/build
- name: Create Build Environment
# Some projects don't allow in-source building, so create a separate build directory
# We'll use this as our working directory for all subsequent commands
run: cmake -E make_directory ${{runner.workspace}}/build
- name: Configure CMake
# Use a bash shell so we can use the same syntax for environment variable
# access regardless of the host operating system
shell: bash
working-directory: ${{runner.workspace}}/build
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=False -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
- name: Build
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_DISABLETELEMETRY: 1
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE --target CodeSizeValidation instcountci_test_files
- name: Instruction Count Tests
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target instcountci_tests
- name: Instruction Count Test Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_InstCountCI.log || true
- name: Truncate test results
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
# Cap out the log files at 20M in case something crash spins and dumps fault text
# ASM tests get quite close to 10MB
run: truncate --size=<20M ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log || true
- name: Set runner name
if: ${{ always() }}
run: echo "runner_name=$(hostname)" >> $GITHUB_ENV
- name: Upload results
if: ${{ always() }}
uses: 'actions/upload-artifact@v3'
timeout-minutes: 1
with:
name: Results-${{ env.runner_name }}
path: ${{runner.workspace}}/build/Testing/Temporary/LastTest_*.log
retention-days: 3
+6 -5
View File
@@ -26,17 +26,18 @@ jobs:
- name: Set runner label
run: echo "runner_label=${{ matrix.arch[1] }}" >> $GITHUB_ENV
- name: Add MingGW to PATH
run: echo "$HOME/llvm-mingw/build/bin/" >> $GITHUB_PATH
- name: Set CC x86
if: matrix.arch[1] == 'x64'
run: |
echo "CC=$HOME/llvm-mingw/build/bin/x86_64-w64-mingw32-clang" >> $GITHUB_ENV
echo "CXX=$HOME/llvm-mingw/build/bin/x86_64-w64-mingw32-clang++" >> $GITHUB_ENV
echo "MINGW_TRIPLE=x86_64-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64
if: matrix.arch[1] == 'ARM64'
run: |
echo "CC=$HOME/llvm-mingw/build/bin/aarch64-w64-mingw32-clang" >> $GITHUB_ENV
echo "CXX=$HOME/llvm-mingw/build/bin/aarch64-w64-mingw32-clang++" >> $GITHUB_ENV
echo "MINGW_TRIPLE=aarch64-w64-mingw32" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
@@ -73,7 +74,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=False -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DENABLE_INTERPRETER=False -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+16 -1
View File
@@ -65,7 +65,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -85,6 +85,21 @@ jobs:
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
- name: ASM Tests 128-bit
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disableavx"
FEX_FORCESVEWIDTH: "128"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test 128-bit Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM128bit.log || true
- name: IR Tests
working-directory: ${{runner.workspace}}/build
shell: bash
+2 -2
View File
@@ -236,7 +236,7 @@ if (BUILD_TESTS)
endif()
add_subdirectory(External/vixl/)
include_directories(External/vixl/src/)
include_directories(SYSTEM External/vixl/src/)
if (CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
# This means we were attempted to get compiled with GCC
@@ -427,7 +427,7 @@ if (BUILD_TESTS)
endif()
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(External/FEXCore)
add_subdirectory(FEXCore/)
# Binfmt_misc files must be installed prior to Source/ installs
add_subdirectory(Data/binfmts/)
+1
View File
@@ -6,3 +6,4 @@ mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xf
credentials yes
fix_binary yes
preserve yes
expose_interpreter optional
+1
View File
@@ -6,3 +6,4 @@ mask \xff\xff\xff\xff\xff\xfe\xfe\x00\x00\x00\x00\xff\xff\xff\xff\xff\xfe\xff\xf
credentials yes
fix_binary yes
preserve yes
expose_interpreter optional
+1 -1
-88
View File
@@ -1,88 +0,0 @@
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include <FEXCore/Core/CPUBackend.h>
namespace FEXCore {
namespace CPU {
CPUBackend::CPUBackend(FEXCore::Core::InternalThreadState *ThreadState, size_t InitialCodeSize, size_t MaxCodeSize)
: ThreadState(ThreadState), InitialCodeSize(InitialCodeSize), MaxCodeSize(MaxCodeSize) {}
CPUBackend::~CPUBackend() {
for (auto CodeBuffer : CodeBuffers) {
FreeCodeBuffer(CodeBuffer);
}
CodeBuffers.clear();
}
auto CPUBackend::GetEmptyCodeBuffer() -> CodeBuffer * {
if (ThreadState->CurrentFrame->SignalHandlerRefCounter == 0) {
if (CodeBuffers.empty()) {
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
} else {
if (CodeBuffers.size() > 1) {
// If we have more than one code buffer we are tracking then walk them and delete
// This is a cleanup step
for (size_t i = 1; i < CodeBuffers.size(); i++) {
FreeCodeBuffer(CodeBuffers[i]);
}
CodeBuffers.resize(1);
}
// Set the current code buffer to the initial
CurrentCodeBuffer = &CodeBuffers[0];
if (CurrentCodeBuffer->Size != MaxCodeSize) {
FreeCodeBuffer(*CurrentCodeBuffer);
// Resize the code buffer and reallocate our code size
CurrentCodeBuffer->Size *= 1.5;
CurrentCodeBuffer->Size = std::min(CurrentCodeBuffer->Size, MaxCodeSize);
*CurrentCodeBuffer = AllocateNewCodeBuffer(CurrentCodeBuffer->Size);
}
}
} else {
// We have signal handlers that have generated code
// This means that we can not safely clear the code at this point in time
// Allocate some new code buffers that we can switch over to instead
auto NewCodeBuffer = AllocateNewCodeBuffer(InitialCodeSize);
EmplaceNewCodeBuffer(NewCodeBuffer);
}
return CurrentCodeBuffer;
}
auto CPUBackend::AllocateNewCodeBuffer(size_t Size) -> CodeBuffer {
CodeBuffer Buffer;
Buffer.Size = Size;
Buffer.Ptr = static_cast<uint8_t *>(
FEXCore::Allocator::VirtualAlloc(Buffer.Size, true));
LOGMAN_THROW_AA_FMT(!!Buffer.Ptr, "Couldn't allocate code buffer");
if (static_cast<Context::ContextImpl*>(ThreadState->CTX)->Config.GlobalJITNaming()) {
static_cast<Context::ContextImpl*>(ThreadState->CTX)->Symbols.RegisterJITSpace(Buffer.Ptr, Buffer.Size);
}
return Buffer;
}
void CPUBackend::FreeCodeBuffer(CodeBuffer Buffer) {
FEXCore::Allocator::VirtualFree(Buffer.Ptr, Buffer.Size);
}
bool CPUBackend::IsAddressInCodeBuffer(uintptr_t Address) const {
for (auto &Buffer: CodeBuffers) {
auto start = (uintptr_t)Buffer.Ptr;
auto end = start + Buffer.Size;
if (Address >= start && Address < end) {
return true;
}
}
return false;
}
}
}
@@ -1,184 +0,0 @@
#pragma once
#include <FEXCore/IR/IR.h>
#define GD *GetDest<uint64_t*>(Data->SSAData, Node)
#define GDP GetDest<void*>(Data->SSAData, Node)
#define DO_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(GDP); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
*Dst_d = func(*Src1_d, *Src2_d); \
break; \
}
#define DO_SCALAR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
Dst_d[0] = func(Src1_d[0], Src2_d[0]); \
break; \
}
#define DO_VECTOR_COMPARE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type2*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_PAIR_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i*2], Src1_d[i*2 + 1]); \
Dst_d[i+Elements] = func(Src2_d[i*2], Src2_d[i*2 + 1]); \
} \
break; \
}
#define DO_VECTOR_SCALAR_OP(size, type, func)\
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], *Src2_d); \
} \
break; \
}
#define DO_VECTOR_0SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(); \
} \
break; \
}
#define DO_VECTOR_1SRC_OP(size, type, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src_d[i]); \
} \
break; \
}
#define DO_VECTOR_REDUCE_1SRC_OP(size, type, func, start_val) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type*>(Src); \
type begin = start_val; \
for (uint8_t i = 0; i < Elements; ++i) { \
begin = func(begin, Src_d[i]); \
} \
Dst_d[0] = begin; \
break; \
}
#define DO_VECTOR_SAT_OP(size, type, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type*>(Src1); \
auto *Src2_d = reinterpret_cast<type*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = func(Src1_d[i], Src2_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_NOSIZE(type, type2, func, min, max) \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i], min, max); \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src2); \
memcpy(Dst_d, Src1, Elements * sizeof(type2));\
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i+Elements] = (type)func(Src_d[i], min, max); \
} \
break; \
}
#define DO_VECTOR_1SRC_2TYPE_OP_TOP_SRC(size, type, type2, func, min, max) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src_d = reinterpret_cast<type2*>(Src); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func(Src_d[i+Elements], min, max); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i], (type)Src2_d[i]); \
} \
break; \
}
#define DO_VECTOR_2SRC_2TYPE_OP_TOP_SRC(size, type, type2, func) \
case size: { \
auto *Dst_d = reinterpret_cast<type*>(Tmp); \
auto *Src1_d = reinterpret_cast<type2*>(Src1); \
auto *Src2_d = reinterpret_cast<type2*>(Src2); \
for (uint8_t i = 0; i < Elements; ++i) { \
Dst_d[i] = (type)func((type)Src1_d[i+Elements], (type)Src2_d[i+Elements]); \
} \
break; \
}
struct InterpVector256 {
__uint128_t Lower;
__uint128_t Upper;
};
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::OrderedNodeWrapper Op) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetDest(void* SSAData, FEXCore::IR::NodeID Op) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Op.Value];
return reinterpret_cast<Res>(DstPtr);
}
template<typename Res>
Res GetSrc(void* SSAData, FEXCore::IR::OrderedNodeWrapper Src) {
auto DstPtr = &reinterpret_cast<InterpVector256*>(SSAData)[Src.ID().Value];
return reinterpret_cast<Res>(DstPtr);
}
@@ -1,77 +0,0 @@
/*
$info$
tags: ir|opts
desc: Sanity checking pass
$end_info$
*/
#include <FEXCore/IR/IR.h>
#include <FEXCore/IR/IREmitter.h>
#include <FEXCore/IR/IntrusiveIRList.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/sstream.h>
#include "Interface/IR/PassManager.h"
#include <memory>
namespace FEXCore::IR::Validation {
class PhiValidation final : public FEXCore::IR::Pass {
public:
bool Run(IREmitter *IREmit) override;
};
bool PhiValidation::Run(IREmitter *IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::PHIValidation");
bool HadError = false;
auto CurrentIR = IREmit->ViewIR();
fextl::ostringstream Errors;
// Walk the list and calculate the control flow
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
bool FoundNonPhi{};
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
switch (IROp->Op) {
// BEGINBLOCK doesn't matter for us
case IR::OP_BEGINBLOCK: break;
case IR::OP_PHIVALUE:
case IR::OP_PHI: {
if (FoundNonPhi) {
// If we have found a non-phi IR op and then had a Phi or PhiValue value then this is a programming mistake
// PHI values MUST be defined at the top of the block only
HadError |= true;
Errors << "Phi %" << CurrentIR.GetID(CodeNode) << ": Was defined after non-phi operations. Which is invalid!" << std::endl;
}
// Check all the phi values to ensure they have the same type
break;
}
default:
FoundNonPhi = true;
break;
}
}
}
if (HadError) {
fextl::stringstream Out;
FEXCore::IR::Dump(&Out, &CurrentIR, nullptr);
Out << "Errors:" << std::endl << Errors.str() << std::endl;
LogMan::Msg::EFmt("{}", Out.str());
}
return false;
}
fextl::unique_ptr<FEXCore::IR::Pass> CreatePhiValidation() {
return fextl::make_unique<PhiValidation>();
}
}
+1 -1
+1 -1
File renamed without changes.
File renamed without changes.
File renamed without changes.
@@ -402,7 +402,7 @@ def print_parse_argloader_options(options):
if (value_type == "strenum"):
output_argloader.write("\tfextl::string UserValue = Options[\"{0}\"];\n".format(op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key))
output_argloader.write("\tSet(FEXCore::Config::ConfigOption::CONFIG_{}, FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, UserValue));\n".format(op_key.upper(), op_key, op_key, op_key))
elif (value_type == "strarray"):
# these need a bit more help
output_argloader.write("\tauto Array = Options.all(\"{0}\");\n".format(op_key))
@@ -431,13 +431,13 @@ def print_parse_envloader_options(options):
value_type = op_vals["Type"]
if (value_type == "strenum"):
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = FEXCore::Config::EnumParser(FEXCore::Config::{}_EnumPairs, Value);\n".format(op_key, op_key))
output_argloader.write("Value = FEXCore::Config::EnumParser<FEXCore::Config::{}ConfigPair>(FEXCore::Config::{}_EnumPairs, Value_View);\n".format(op_key, op_key, op_key))
output_argloader.write("}\n")
if ("ArgumentHandler" in op_vals):
conversion_func = "FEXCore::Config::Handler::{0}".format(op_vals["ArgumentHandler"])
output_argloader.write("else if (Key == \"FEX_{0}\") {{\n".format(op_key.upper()))
output_argloader.write("Value = {0}(Value);\n".format(conversion_func))
output_argloader.write("Value = {0}(Value_View);\n".format(conversion_func))
output_argloader.write("}\n")
output_argloader.write("#endif\n")
@@ -447,7 +447,7 @@ def print_parse_enum_options(options):
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
if (op_vals["Type"] == "strenum"):
output_argloader.write("enum {} : uint64_t {{\n".format(op_key))
output_argloader.write("enum class {} : uint64_t {{\n".format(op_key))
Enums = op_vals["Enums"]
i = 0
# Always have an OFF.
@@ -457,6 +457,8 @@ def print_parse_enum_options(options):
i += 1
output_argloader.write("};\n")
output_argloader.write("FEX_DEF_NUM_OPS({})\n".format(op_key))
for op_group, group_vals in options.items():
for op_key, op_vals in group_vals.items():
@@ -133,11 +133,11 @@ set (SRCS
Interface/IR/Passes/DeadCodeElimination.cpp
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/PhiValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
File renamed without changes.
File renamed without changes.
Loaded 100 of 412 files, more files were not shown because too many files have changed in this diff. Show more