Commit Graph
7661 Commits
Author SHA1 Message Date
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Mai 750d90939d Merge pull request #3088 from Sonicadvance1/instcountci_missing_secondary_opsize
InstCountCI: Adds missing instructions from Secondary OpSize tables
2023-09-13 22:49:17 -04:00
Mai 31d828390f Merge pull request #3087 from Sonicadvance1/tbl2_implementation
OpcodeDispatcher: Implement shufps with VTBL2 in worst case
2023-09-13 22:48:53 -04:00
Ryan Houdek 2aea401189 InstCountCI: Adds missing instructions from Secondary OpSize tables
I managed to miss a whole section of instructions from the secondary
opsize tables. This resulted in four instructions missing from the
database.

Adds cmppd, pinsrw, pextrw, and shufpd which are all non-optimal
instruction implementations.
2023-09-13 11:48:50 -07:00
Ryan Houdek c008671509 unittests/asm: Add test with inverted sources
To ensure this is tested with non sequential source registers.
2023-09-13 11:31:20 -07:00
Ryan Houdek 5903be156c InstCountCI: Update for shufps tbl opt 2023-09-13 11:31:20 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Ryan Houdek cfe620ab15 Merge pull request #3085 from Sonicadvance1/optimize_shufps
OpcodeDispatcher: Optimize a bunch of shufps variants
2023-09-12 21:53:33 -07:00
Ryan Houdek ea8d63350a InstCountCI: Updates for optimized shufps 2023-09-12 19:58:07 -07:00
Ryan Houdek e37cef8283 unittests: Implement shufps optimization test
Tests all current forms of shufps optimizations.
2023-09-12 19:58:07 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek d5c3036bc2 JITx86: Fixes VREV64 with 32-bit element size.
This has been incorrect since it has been implemented.
Noticed when implementing optimizations.
2023-09-12 19:23:59 -07:00
Mai ebdca02218 Merge pull request #3084 from Sonicadvance1/optimize_bswap
OpcodeDispatcher: Optimize 32-bit bswap
2023-09-12 20:09:00 -04:00
Mai dda5861bdd Merge pull request #3081 from Sonicadvance1/fix_waitpid
Tools: Fixes usage of waitpid in the face of EINTR
2023-09-12 19:35:05 -04:00
Mai f7e652b616 Merge pull request #3083 from Sonicadvance1/optimize_nop_move
OpcodeDispatcher: Optimize NOP vector move
2023-09-12 19:34:36 -04:00
Ryan Houdek 65bc159ff1 InstCountCI: Update for bswap optimization 2023-09-12 16:19:55 -07:00
Ryan Houdek c362d3a9d8 OpcodeDispatcher: Optimize 32-bit bswap
Removes a redundant move, making it optimal now.
2023-09-12 16:19:10 -07:00
Ryan Houdek 8a44be0c30 InstCountCI: Update for NOP vector moves
Adds a couple of instructions that get tested in this code path.
2023-09-12 16:11:43 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Ryan Houdek aa017116b3 Tools: Fixes usage of waitpid in the face of EINTR
waitpid can return early if interrupted due to EINTR.
Loop on this case and try again.
2023-09-12 12:41:43 -07:00
Mai 90f7937146 Merge pull request #3079 from Sonicadvance1/recover_two_temps
Arm64: Recover two unused vector vector temporary registers
2023-09-11 22:06:03 -04:00
Mai 98f148766d Merge pull request #3078 from Sonicadvance1/detect_flagm
HostFeatures: Detect FlagM/2
2023-09-11 20:57:43 -04:00
Ryan Houdek 9c44e295fa InstCountCI: Update for recovering two vector temps
All the changes are RA changes and spilling/filling taking another
instruction.
2023-09-11 16:50:52 -07:00
Ryan Houdek b5a1d323c2 Arm64: Recover two unused vector vector temporary registers
This leaves us with two temporary vectors that the JIT can use.
As of last month we stopped using v2 and v3 as temporaries and these can
now be given back to the JIT.

Ensures that the registers are still sequentially ordered and adds
support for spilling the FPR counts that are aligned by 2 instead of 4.
Adds a couple of instructions to filling and spilling but isn't that big
of an issue.

InstcountCI has some ridiculously large changes just because RA is
starting at a new register number.
2023-09-11 16:48:25 -07:00
Ryan Houdek b453439968 HostFeatures: Detect FlagM/2
Currently unused but at least detect the feature so that our Arm64 JIT
can use it in the future.
2023-09-11 16:41:30 -07:00
Mai 48521a4416 Merge pull request #3075 from Sonicadvance1/optimize_bt_ops
OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
2023-09-11 16:05:33 -04:00
Mai 6fe643d270 Merge pull request #3076 from Sonicadvance1/enable_enhanced_rep_movs
CPUID: Enabled Enhanced REP MOVSB/STOSB
2023-09-11 15:35:56 -04:00
Mai fbc4bda7a6 Merge pull request #3074 from Sonicadvance1/hwcap2_fsgsbase
ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
2023-09-11 15:35:26 -04:00
Mai 6d9b52452e Merge pull request #3072 from Sonicadvance1/crc32_is_fixed_size
IR: Changes crc32 operation to always return a 32-bit result.
2023-09-11 15:34:37 -04:00
Mai 950007c815 Merge pull request #3071 from Sonicadvance1/update_rcl_opsize
OpcodeDispatcher: Update 32/64-bit RCL for operating size
2023-09-11 15:34:06 -04:00
Mai d029394c27 Merge pull request #3070 from Sonicadvance1/update_rcr_opsize
OpcodeDispatcher: Update 32/64-bit RCR for operating size
2023-09-11 15:33:34 -04:00
Mai 879fcdc6fe Merge pull request #3069 from Sonicadvance1/fix_redundant_load_rclse
IR:RCLSE: Partially reenables the RCLSE pass
2023-09-11 15:32:55 -04:00
Ryan Houdek 2f77982b54 CPUID: Enabled Enhanced REP MOVSB/STOSB
Missed with #2490.
This changes behaviour of glibc's memmove slightly, seems to recover a
bit of performance on Half-Life 2's title screen.
2023-09-10 20:49:08 -07:00
Ryan Houdek e3a00fb2fb InstCountCI: Update for BT minor opt 2023-09-10 20:23:08 -07:00
Ryan Houdek 4feb059f51 OpcodeDispatcher: Optimize the case of all flags invalidated
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
2023-09-10 20:16:29 -07:00
Ryan Houdek 3d1bbe505d OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
2023-09-10 20:16:29 -07:00
Ryan Houdek b2a42b6c61 ELFCodeLoader: Expose FSGSBase in getauxval HWCAP2
We have supported this since #163 but we haven't been exposing the
feature in hwcap2.

We have exposed it in CPUID this entire time, just not in hwcap2.
2023-09-10 17:00:21 -07:00
Ryan Houdek 315d1855de IR: Changes crc32 operation to always return a 32-bit result.
CRC32 is always a 32-bit sized operation even with a 64-bit source
value.
This doesn't change any InstCountCI results.
2023-09-09 10:02:30 -07:00
Ryan Houdek 93246878e2 InstCountCI: Update for rcl explicit size change 2023-09-09 09:40:36 -07:00
Ryan Houdek 6c62691af0 OpcodeDispatcher: Update 32/64-bit RCL for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:40:12 -07:00
Ryan Houdek ee5aed51d8 InstCountCI: Update for rcr expliti size change 2023-09-09 09:35:13 -07:00
Ryan Houdek 47f50a7008 OpcodeDispatcher: Update 32/64-bit RCR for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:33:27 -07:00
Ryan Houdek be07254935 Merge pull request #3067 from neobrain/refactor_thunks
Thunks: Minor restructuring and small cleanups
2023-09-07 20:16:17 -07:00
Ryan Houdek 636f8aa4a7 Arm64: Fix undefined behaviour in Push operation
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.

This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.

Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
2023-09-07 17:38:39 -07:00
Ryan Houdek 22ca46a227 Arm64: Fixes SVE V{S,U}MulH
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```

We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```

As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.

InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
2023-09-07 16:47:08 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Ryan Houdek b753b9ffa2 InstCountCI: Updates for RCLSE fix
Adds two `packsswb` tests to ensure redundant sources are getting
optimized as expected.
2023-09-07 15:58:43 -07:00
Ryan Houdek c62b5a3103 IR:RCLSE: Partially reenables the RCLSE pass
This is taking steps to start fixing RCLSE which was started by #2700.
Same situation as that PR, since #2170 when we converted
{Load,Store}Context in to {Load,Store}Register we broke this pass
entirely. It hasn't been doing anything for redundant GPRs and FPRs
since at least November of last year.

Technically it was potentially still optimizing redundant MMX
accesses, but it is so broken that it doesn't matter.

Instead of going all in like #2700 did, tear down the pass and start
again. We are now /only/ optimizing redundant context/register loads.
This fixes an issue that comes up commonly where the same register used
as sources was getting loaded twice, causing redundant moves.

`packsswb xmm0, xmm0` for example was generating a four instruction
sequence instead of three instructions because we weren't eliminating
the redundant load.

Going to take reimplementing all the optimizations that this pass does
in steps. This way we can track any regression in the independent steps
unlike what happened in #2700.

Confirmed that Proton/Sonic Mania still works after this.
2023-09-07 15:50:27 -07:00