Commit Graph
190 Commits
Author SHA1 Message Date
Alyssa Rosenzweig 7a06cc9727 IR: Use adcs/sbcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 09:06:46 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Alyssa Rosenzweig c8519b0b87 OpcodeDispatcher: Remove LoadPF
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 68d32ad70d OpcodeDispatcher: Optimize PF in lahf
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.

Closes #3117

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 86063411dc Revert "OpcodeDispatcher: Use plain Lshl for flags"
This logic is unused since 8adfaa9aa ("OpcodeDispatcher: Use SelectCC for x87"),
which addressed the underlying issue.

This reverts commit df3833edbe.
2023-09-24 20:47:50 -04:00
Alyssa Rosenzweig c5fc03dac4 OpcodeDispatcher: Use cset for blsr/etc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig e63871ed2e OpcodeDispatcher: Handle sub in CalculateOF
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig ea8b7633eb OpcodeDispatcher: Optimize OF calc of immediates
If we know the sign of one of the sources, we can do better when calculating OF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 18:16:09 -04:00
Alyssa Rosenzweig b1231c24ef OpcodeDispatcher: Omit AF xor for common constants
The only reason we need to XOR arguments for AF is to get bit 4 correct. But if
the operand in question is known to have bit 4 clear, the XOR will be an
effective no-op and can be skipped. This saves an instruction in a bunch of
common cases, like inc/dec. If we dedicated a register to AF to eliminate the
store, we would not save an instruction from this but would still come out ahead
due to an eor turning into a (zero cycle?) mov that can be handled by the
renamer.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig 699aa85c4b OpcodeDispatcher: Opt PF selection
Fold the and in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 1596e33f58 OpcodeDispatcher: Remove pointless or
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 07d03f1610 OpcodeDispatcher: Don't opencode bfe, badly
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig a8b48dcacd OpcodeDispatcher: Swap some selects
...if it lets us use cset.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig bb87b2a19d OpcodeDispatcher: Use more Orlshl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 09:13:41 -04:00
Alyssa Rosenzweig 19eff62c77 OpcodeDispatcher: Use orlshl for FCW
Potentially easier on the RA (bfi has a tied operand), mostly whatever here.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-21 08:55:25 -04:00
Mai 5fc8699db9 Merge pull request #3130 from Sonicadvance1/optimize_fsw
OpcodeDispatcher: Optimize reconstructing FSW
2023-09-21 08:35:16 -04:00
Ryan Houdek 5664195e49 OpcodeDispatcher: Optimize reconstructing FSW
Minor optimization using Bfi to insert C0, C1, C2, & C3
2023-09-21 02:07:27 -07:00
Ryan Houdek 8e9e87f631 OpcodeDispatcher: Removes non-explicit SelectCC function
Renames the explicit sized one to `SelectCC`
Cleans up a bit of duplicated code.
2023-09-21 01:56:38 -07:00
Ryan Houdek 67680d71a4 Merge pull request #3125 from Sonicadvance1/spdx_fexcore
FEXCore: Adds SPDX identifier
2023-09-19 17:42:07 -07:00
Ryan Houdek 38f1536255 FEXCore/Interface/Core/OpcodeDispatcher: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Alyssa Rosenzweig 8adfaa9aa6 OpcodeDispatcher: Use SelectCC for x87
Better code gen and will benefit from future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:37:54 -04:00
Alyssa Rosenzweig df3833edbe OpcodeDispatcher: Use plain Lshl for flags
If we have PF but no CF this simplifies the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek 759cc0025a OpcodeDispatcher: Add a dirty flag for tracking NZCV status
Cached NZCV reads don't need to be written back at the end of the block.
This will remove one instruction from the end of some blocks.
2023-09-15 10:09:37 -07:00
Alyssa Rosenzweig d29b8bab36 OpcodeDispatcher: Inline constant in PF calculation
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:33:53 -04:00
Alyssa Rosenzweig fc02f38435 IR: Only invert CF for NZCV if needed
If we are going to throw away the updated value of CF anyway there is no point
wasting an instruction to invert CF. Add an IR toggle for that so the arm64 JIT
can make better choices.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-15 12:08:22 -04:00
Ryan Houdek d5782567e8 Merge pull request #3077 from Sonicadvance1/x86_shifted
FEXCore: Implements support for shifted bitwise ops
2023-09-15 08:09:35 -07:00
Ryan Houdek 9866e238d5 Merge pull request #3080 from Sonicadvance1/defer_softfloat
FEXCore: Defer setting x87 softflow rounding mode until use
2023-09-15 08:08:04 -07:00
Ryan Houdek f730339365 OpcodeDispatcher: Reorder vector loads in shifts
This affects codegen due to RA quirks. This ensures that the wide shifts
don't have to generate a movprfx.
2023-09-14 19:34:00 -07:00
Mai 92824f5e4d Merge pull request #3093 from Sonicadvance1/optimize_blendp
OpcodeDispatcher: Optimize blendp{s,d}
2023-09-14 00:33:00 -04:00
Mai 213d3c4e2b Merge pull request #3091 from Sonicadvance1/optimize_pinsr
OpcodeDispatcher: Optimize pins{b,w,d,q}
2023-09-14 00:30:56 -04:00
Mai d4c6749d2a Merge pull request #3090 from Sonicadvance1/optimize_pextr
OpcodeDispatcher: Optimize pextr{b,w}
2023-09-14 00:30:43 -04:00
Ryan Houdek 29f824cf7a OpcodeDispatcher: Optimize blendp{s,d}
Optimal blendps is worst case 2 instructions.
FEX's RA doesn't quite get there since it can't see through multiple
instructions with SRA destinations. That'll be fixed in the future.

Optimal blendps is always one instruction, one is a no-op.
We always hit this.
2023-09-13 20:06:39 -07:00
Ryan Houdek 5e7d793a6a OpcodeDispatcher: Optimize pins{b,w,d,q}
Inserting from a GPR and memory can both be optimized. These are now
optimal

Needs #3088 merged first.
2023-09-13 19:53:05 -07:00
Ryan Houdek 33a2fbb896 OpcodeDispatcher: Optimize pextr{b,w}
Cleans up the code which had special cased some 32-bit optimization
which is unnecessary now that both 8-bit and 16-bit are also optimized.

When FEX does a VExtractToGPR, the result is zero extended to the full
GPR register size. This means we don't need to do a zero extend when
storing to a guest GPR.

Makes pextr{b,w} optimal now.

Needs #3088 merged first.
2023-09-13 19:51:56 -07:00
Ryan Houdek 67914157cb OpcodeDispatcher: Optimize shufpd
This one is very satisfying since there are only four variants and each
one of them converts to a single instruction.

Needs #3088 merged first
2023-09-13 19:50:15 -07:00
Ryan Houdek db5056f275 OpcodeDispatcher: Implement shufps with VTBL2 in worst case
In the case that source registers are sequential then this turns in to a
load of the vector constant (2 instructions) and the single tbl
instruction.

If the registers aren't sequential then the tbl turns in to 2 moves and
then the single tbl, which with zero-cycle rename isn't too bad.

Since this is a worst case option this is significantly better than the
previous implementation doing a bunch of inserts which was always 9
instructions.
We should still strive to implement faster versions without the use of
TBL2 if possible but this makes it less of a concern.
2023-09-13 11:31:20 -07:00
Ryan Houdek 3f1979286f OpcodeDispatcher: Optimize a bunch of shufps variants
Hits a whole bunch of common cases, most of which then emit optimal code
generation.
Two cases that use VInsElement hit the RA quirk where the SRA
destination is dead but RA doesn't see it, so it ends up doing a couple
moves. If RA gets fixed then those two moves will go away.

There are definitely still cases that we could emit more optimal code.
Additionally we could implement a TBL2 IR operation to do a LUT approach
for ones we don't cover.

Problem with implementing a TBL2 ir operation is that we have no way to
ensure registers are sequential so we would need to always do moves
```asm
ldr v2, <LUT Table>
mov v0, v16
mov v1, v18
tbl v16.16b, { v0.16b, v1.16b }, v2.16b
```

Which to be fair isn't terrible, and if we're lucky that the guest uses
sequential registers we can naturally get the more optimal code path.
Ideally our RA could push some operations in to sequential registers but
that's not possible currently.

I'll do a follow-up PR that implements TBL2.
2023-09-12 19:58:07 -07:00
Ryan Houdek 304dba5f20 OpcodeDispatcher: Optimize NOP vector move
Move instruction to itself here is a nop.
Need to be careful about AVX operations which use a different handler
since those might actually zero the upper bits on 128-bit move
2023-09-12 16:10:39 -07:00
Ryan Houdek 76bd81af15 FEXCore: Defer setting x87 softflow rounding mode until use
Currently FEX will always jump out of the JIT any time FCW was getting
written to, ensuring that the softfloat state is setup to rounding at
the time of FCW getting written.
This has the unintended side-effect that even in "x87 reduced precision"
mode we were jumping out of the JIT.
This hit a real world use case of an installer reloading FCW after every
x87 operation and generating a block with 2297 instructions.

Instead when jumping out of the JIT for handling x87 operations, load
FCW and pass it as the first argument of the handler. Setting the
softfloat state at that point.

This helps the installer's hottest block by cutting it down to 1477
instructions. 64.3% of the original size. The code block is still
burning 90% of the CPU time of the installer but the performance is
significantly better while it is doing its decompression.

In order to optimize this installer's block of code more then we will
likely need to optimize out x87 stack usage.
2023-09-12 05:21:06 -07:00
Ryan Houdek 863331b117 FEXCore: Implements support for shifted bitwise ops
This wasn't implemented initially for the interpreter and x86 JIT.

This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.

The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
2023-09-11 13:17:35 -07:00
Ryan Houdek 7b80427de0 OpcodeDispatcher: Remove BLENDV "optimization"
Now that the RCLSE pass finally optimizes redundant loads again this
optimization that lives in the OpcodeDispatcher can be removed.

With InstCountCI reran, the pblendvb results don't change at all, as
expected.
2023-09-07 16:00:56 -07:00
Alyssa Rosenzweig 8efe2eeef6 OpcodeDispatcher: Defer PF invert
Now the calculation of PF is entirely deferred, by inverting our internal
representation of PF. All the (e.g.) logical op needs to do is store the low
8-bits of the result.

This is a bit of a mixed bag. Primary ALU ops all save an instruction, by
skip the XOR. Loading PF takes an extra instruction, that's expected. The tricky
cases are:

* Zeroing PF. This now requires writing 1 instead of 0, which may require an
  extra move for the constant. Some of this will go away when we merge PF+AF
  into a single register, which is next up on the list. In that case, the
  two stores will turn into 1 `or`. So if we need to write a 1 to PF (zeroing
  x86 view of PF), that will get absorbed into the or, if we also write AF. If
  we leave AF undefined and need to write a 1, that's a single mov instruction
  and we couldn't do better anyway if not inverted (since we'd still have a mov
  wzr even then). So in view of the future work, this isn't something I'm
  concerned about.

* Float comparisons that put Unordered into PF. These require an extra invert to
  match the new convention. These are already so unnecessarily bloated that I'm
  not convinced I'm making things materially worse here. But we realistically
  need multiple destination support in the IR to fix this particular mess.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 15:58:18 -04:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig ff0b514da8 OpcodeDispatcher: Invalidate PF/AF in more cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:25:14 -04:00
Alyssa Rosenzweig 240260576b OpcodeDispatcher: Stop zeroing so many flags
Use an explicit invalidate, so we can zero easily enough if we need to for
debugging later but we can save the instrs ordinarily.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 12:10:29 -04:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00