Commit Graph
153 Commits
Author SHA1 Message Date
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Alyssa Rosenzweig 8efe2eeef6 OpcodeDispatcher: Defer PF invert
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.

This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:

* Zeroing PF. This now requires writing 1 instead of 0, which may require an
  extra move for the constant. Some of this will go away when we merge PF+AF
  into a single register, which is next up on the list. In that case, the
  two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
  x86 view of PF), that will get absorbed into the or, if we also write AF. If
  we leave AF undefined and need to write a 1, that's a single mov instruction
  and we couldn't do better anyway if not inverted (since we'd still have a mov
  wzr even then). So in view of the future work, this isn't something I'm
  concerned about.

* Float comparisons that put Unordered into PF. These require an extra invert to
  match the new convention. These are already so unnecessarily bloated that I'm
  not convinced I'm making things materially worse here. But we realistically
  need multiple destination support in the IR to fix this particular mess.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 15:58:18 -04:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig ff0b514da8 OpcodeDispatcher: Invalidate PF/AF in more cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig 240260576b OpcodeDispatcher: Stop zeroing so many flags
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:10:29 -04:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00
Alyssa Rosenzweig c88b022d9f OpcodeDispatcher: Add AF accessors
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Mai 2d22176699 Merge pull request #3052 from Sonicadvance1/remove_todo_shlimm
OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
2023-09-03 04:38:52 -04:00
Ryan Houdek c425db7284 OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
This one was easy, barely anything changes behaviour, as seen by
InstCountCI changes.
2023-09-03 00:27:28 -07:00
Ryan Houdek d0595b5f13 OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront 2023-09-03 00:16:26 -07:00
Billy Laws 13b8f95f85 X87: Switch all stack pointer accesses to 32-bit OpSize 2023-09-02 09:17:33 -07:00
Billy Laws cb49373f47 FEXCore: Rework X87 tag word handling
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
2023-09-02 09:17:33 -07:00
Alyssa Rosenzweig dcb09b085b OpcodeDispatcher: Use AddNZCV/SubNZCV
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:41 -04:00
Alyssa Rosenzweig 45d99a2cce OpcodeDispatcher: Fix source sizes for Sub flags
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:39 -04:00
Ryan Houdek 61df7a576a OpcodeDispatcher: Be super defensive when starting a new block
Ensure all cached data is correct.
2023-08-30 20:43:11 -07:00
Ryan Houdek ffaa908475 OpcodeDispatcher: Use the named zero register for each usage
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
2023-08-30 20:43:11 -07:00
Ryan Houdek a2307f28d6 Arm64: Optimize AES operations by caching a zero register
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.

Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.

AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
2023-08-30 19:00:43 -07:00
Ryan Houdek 24215f7ad0 OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
2023-08-30 17:09:34 -07:00
Mai 02b891c0fe Merge pull request #3040 from Sonicadvance1/optimize_aeskeygen
Arm64: Optimize AESKeyGenAssist
2023-08-30 15:58:29 -04:00
Ryan Houdek d8f131fa3d Arm64: Optimize AESKeyGenAssist
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.

I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
2023-08-30 12:15:09 -07:00
Ryan Houdek 9bfd4b650f OpcodeDispatcher: Removes erroneous debug log 2023-08-30 11:12:12 -07:00
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00
Ryan Houdek 8bc84c202c IR: Removes implicit sized xor 2023-08-28 19:51:14 -07:00
Ryan Houdek e9a3848602 Merge pull request #3027 from Sonicadvance1/remove_implicit_andn
IR: Removes implicit sized andn
2023-08-28 19:39:49 -07:00
Ryan Houdek 1699ec9a76 IR: Removes implicit sized andn 2023-08-28 19:16:16 -07:00
Ryan Houdek db6c8852fc IR: Removes implicit sized or 2023-08-28 19:06:05 -07:00
Ryan Houdek 65dc6f3e90 IR: Removes implicit sized lshr 2023-08-28 18:16:56 -07:00
Ryan Houdek 60c4438780 IR: Removes implicit sized lshl 2023-08-28 17:50:41 -07:00
Ryan Houdek b9e4a1423f IR: Removes sext IR helper
You hold no power here IR operation.
2023-08-28 17:03:38 -07:00
Ryan Houdek 48669b7006 IR: Removes implicit sized orlshl/orlshr 2023-08-28 05:04:32 -07:00
Ryan Houdek 5768444ce9 IR: Removes implicit sized abs 2023-08-28 05:04:32 -07:00
Ryan Houdek ce8392d5ae IR: Removes implicit sized ror 2023-08-28 05:02:01 -07:00
Ryan Houdek 386cf36cfd IR: Removes implicit sized sbfe
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.

We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek b95648a4ab IR: Removes implicit sized FindMSB 2023-08-28 05:02:01 -07:00
Ryan Houdek bf18672999 IR: Removes implicit sized FindLSB 2023-08-28 05:02:01 -07:00
Ryan Houdek 514a8223d9 OpcodeDispatcher: Optimizes SSE movmaskps
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.

I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek 8d110738ac IR: Add option to disable vector shift range clamping
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.

Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.

When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Ryan Houdek 9ba46f429e X8764: Ensure frndint uses host rounding mode
This previously used `Round_Nearest` which had a bug on Arm64 that it
actually was always using `Round_Host` aka frinti.
Ever since 393cea2e8ba47a15a3ce31d07a6088a2ff91653c[1] this has been fixed
so that `Round_Nearest` actually uses frintn for neaest.

This instruction actually wants to use the host rounding mode.
Once issue with this is that x87 and SSE have different rounding mode
flags and currently we conflate the two in our JIT. This will need to be
fixed in the future.

In the meantime this restores behaviour that it actually uses the host
rounding mode, which fixes black screen and broken vertices in Grim
Fandango Remastered.

[1] e89321dc60 for scalar.
2023-08-25 16:04:01 -07:00