Commit Graph
320 Commits
Author SHA1 Message Date
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek d5c3036bc2 JITx86: Fixes VREV64 with 32-bit element size.
This has been incorrect since it has been implemented.
Noticed when implementing optimizations.
2023-09-12 19:23:59 -07:00
Mai ebdca02218 Merge pull request #3084 from Sonicadvance1/optimize_bswap
OpcodeDispatcher: Optimize 32-bit bswap
2023-09-12 20:09:00 -04:00
Ryan Houdek c362d3a9d8 OpcodeDispatcher: Optimize 32-bit bswap
Removes a redundant move, making it optimal now.
2023-09-12 16:19:10 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Mai 90f7937146 Merge pull request #3079 from Sonicadvance1/recover_two_temps
Arm64: Recover two unused vector vector temporary registers
2023-09-11 22:06:03 -04:00
Ryan Houdek b5a1d323c2 Arm64: Recover two unused vector vector temporary registers
This leaves us with two temporary vectors that the JIT can use.
As of last month we stopped using v2 and v3 as temporaries and these can
now be given back to the JIT.

Ensures that the registers are still sequentially ordered and adds
support for spilling the FPR counts that are aligned by 2 instead of 4.
Adds a couple of instructions to filling and spilling but isn't that big
of an issue.

InstcountCI has some ridiculously large changes just because RA is
starting at a new register number.
2023-09-11 16:48:25 -07:00
Ryan Houdek b453439968 HostFeatures: Detect FlagM/2
Currently unused but at least detect the feature so that our Arm64 JIT
can use it in the future.
2023-09-11 16:41:30 -07:00
Mai 48521a4416 Merge pull request #3075 from Sonicadvance1/optimize_bt_ops
OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
2023-09-11 16:05:33 -04:00
Mai 6fe643d270 Merge pull request #3076 from Sonicadvance1/enable_enhanced_rep_movs
CPUID: Enabled Enhanced REP MOVSB/STOSB
2023-09-11 15:35:56 -04:00
Mai 6d9b52452e Merge pull request #3072 from Sonicadvance1/crc32_is_fixed_size
IR: Changes crc32 operation to always return a 32-bit result.
2023-09-11 15:34:37 -04:00
Mai 950007c815 Merge pull request #3071 from Sonicadvance1/update_rcl_opsize
OpcodeDispatcher: Update 32/64-bit RCL for operating size
2023-09-11 15:34:06 -04:00
Mai d029394c27 Merge pull request #3070 from Sonicadvance1/update_rcr_opsize
OpcodeDispatcher: Update 32/64-bit RCR for operating size
2023-09-11 15:33:34 -04:00
Ryan Houdek 2f77982b54 CPUID: Enabled Enhanced REP MOVSB/STOSB
Missed with #2490.
This changes behaviour of glibc's memmove slightly, seems to recover a
bit of performance on Half-Life 2's title screen.
2023-09-10 20:49:08 -07:00
Ryan Houdek 4feb059f51 OpcodeDispatcher: Optimize the case of all flags invalidated
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
2023-09-10 20:16:29 -07:00
Ryan Houdek 3d1bbe505d OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
2023-09-10 20:16:29 -07:00
Ryan Houdek 315d1855de IR: Changes crc32 operation to always return a 32-bit result.
CRC32 is always a 32-bit sized operation even with a 64-bit source
value.
This doesn't change any InstCountCI results.
2023-09-09 10:02:30 -07:00
Ryan Houdek 6c62691af0 OpcodeDispatcher: Update 32/64-bit RCL for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:40:12 -07:00
Ryan Houdek 47f50a7008 OpcodeDispatcher: Update 32/64-bit RCR for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:33:27 -07:00
Ryan Houdek 636f8aa4a7 Arm64: Fix undefined behaviour in Push operation
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.

This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.

Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
2023-09-07 17:38:39 -07:00
Ryan Houdek 22ca46a227 Arm64: Fixes SVE V{S,U}MulH
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```

We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```

As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.

InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
2023-09-07 16:47:08 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Ryan Houdek c62b5a3103 IR:RCLSE: Partially reenables the RCLSE pass
This is taking steps to start fixing RCLSE which was started by #2700.
Same situation as that PR, since #2170 when we converted
{Load,Store}Context in to {Load,Store}Register we broke this pass
entirely. It hasn't been doing anything for redundant GPRs and FPRs
since at least November of last year.

Technically it was potentially still optimizing redundant MMX
accesses, but it is so broken that it doesn't matter.

Instead of going all in like #2700 did, tear down the pass and start
again. We are now /only/ optimizing redundant context/register loads.
This fixes an issue that comes up commonly where the same register used
as sources was getting loaded twice, causing redundant moves.

`packsswb xmm0, xmm0` for example was generating a four instruction
sequence instead of three instructions because we weren't eliminating
the redundant load.

Going to take reimplementing all the optimizations that this pass does
in steps. This way we can track any regression in the independent steps
unlike what happened in #2700.

Confirmed that Proton/Sonic Mania still works after this.
2023-09-07 15:50:27 -07:00
Ryan Houdek 3c729bcacb IR/RCLSE: Removes unused CalculateControlFlowInfo
This is unused and this only optimizes inside of a block.
2023-09-07 15:49:29 -07:00
Ryan Houdek 16f826c18e Merge pull request #3064 from alyssarosenzweig/ir/rm-phi
IR: Remove phi nodes
2023-09-05 19:43:52 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Alyssa Rosenzweig 8efe2eeef6 OpcodeDispatcher: Defer PF invert
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.

This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:

* Zeroing PF. This now requires writing 1 instead of 0, which may require an
  extra move for the constant. Some of this will go away when we merge PF+AF
  into a single register, which is next up on the list. In that case, the
  two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
  x86 view of PF), that will get absorbed into the or, if we also write AF. If
  we leave AF undefined and need to write a 1, that's a single mov instruction
  and we couldn't do better anyway if not inverted (since we'd still have a mov
  wzr even then). So in view of the future work, this isn't something I'm
  concerned about.

* Float comparisons that put Unordered into PF. These require an extra invert to
  match the new convention. These are already so unnecessarily bloated that I'm
  not convinced I'm making things materially worse here. But we realistically
  need multiple destination support in the IR to fix this particular mess.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 15:58:18 -04:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Ryan Houdek c27f69dd6b Merge pull request #3061 from alyssarosenzweig/flag/cmc
OpcodeDispatcher: Optimize CMC
2023-09-05 11:06:21 -07:00
Alyssa Rosenzweig b6462ee854 OpcodeDispatcher: Optimize CMC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 13:55:05 -04:00
Alyssa Rosenzweig ff0b514da8 OpcodeDispatcher: Invalidate PF/AF in more cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig 240260576b OpcodeDispatcher: Stop zeroing so many flags
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:10:29 -04:00
Ryan Houdek 486f0ba1e3 Merge pull request #3038 from alyssarosenzweig/flag/af
Defer AF extract
2023-09-05 08:54:47 -07:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00
Alyssa Rosenzweig f46ba52e0e OpcodeDispatcher: Use LoadAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Alyssa Rosenzweig c88b022d9f OpcodeDispatcher: Add AF accessors
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Ryan Houdek a24680b0a1 ConstProp: Adds constpool distance heuristic
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.

This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.

FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.

To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.
2023-09-03 18:27:43 -07:00
Ryan Houdek ca1c33047c Arm64: Only allocate vixl::Decoder if enabled
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.

Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
2023-09-03 12:53:18 -07:00
Ryan Houdek b18592f153 Merge pull request #3050 from Sonicadvance1/fix_flag_reconstruction
OpcodeDispatcher: Fixes NZCV and PF flag compacting
2023-09-03 10:15:44 -07:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Ryan Houdek 44a14e7fd0 OpcodeDispatcher: Cleans up RFLAGS size handling
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.

Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
2023-09-03 01:41:04 -07:00
Mai 2d22176699 Merge pull request #3052 from Sonicadvance1/remove_todo_shlimm
OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
2023-09-03 04:38:52 -04:00
Ryan Houdek c425db7284 OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
This one was easy, barely anything changes behaviour, as seen by
InstCountCI changes.
2023-09-03 00:27:28 -07:00