Commit Graph
82 Commits
Author SHA1 Message Date
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek d5c3036bc2 JITx86: Fixes VREV64 with 32-bit element size.
This has been incorrect since it has been implemented.
Noticed when implementing optimizations.
2023-09-12 19:23:59 -07:00
Ryan Houdek 636f8aa4a7 Arm64: Fix undefined behaviour in Push operation
Arm64 store with writeback when source register is the same register as
the address is undefined behaviour.
Depending on hardware details this can do a whole bunch of things.

This situation happens when the x86 code does `push rsp` which is quite
common for applications to do. We would then convert this to a `str x8, [x8, #-8]!`
Which results in undefined behaviour.

Now that redundant loads are optimized this showed up as an issue. Adds
a unit test to ensure we don't hit this again.
2023-09-07 17:38:39 -07:00
Ryan Houdek 22ca46a227 Arm64: Fixes SVE V{S,U}MulH
When the destination overlaps one of the sources we must be careful to
follow a movprfx rule.
```
The destination register must not refer to architectural register state
referenced by any other source operand register of this instruction.
```

We ended up in a situation in the vpmulh{u,}w AVX tests where zm was
overlapping the destination which violated that rule. This also
generated invalid code for this instruction.
```
[INFO] movprfx z6, z4
[INFO] umulh z6.h, p6/m, z6.h, z6.h
```

As seen, we were overwriting one of the sources because the destination
overlapped it. Now instead check if each individual overlap so invalid
code isn't generated.

InstCountCI results aren't affected since this only happens in
situations with multiple instructions.
2023-09-07 16:47:08 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Ryan Houdek ca1c33047c Arm64: Only allocate vixl::Decoder if enabled
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.

Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
2023-09-03 12:53:18 -07:00
Alyssa Rosenzweig 659568ec1b ConstProp: Propagate 0 to first argument of SubNZCV
This allows inlining a constant into the comparison for Neg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:01 -04:00
Alyssa Rosenzweig f688364bf2 IR: Add AddNZCV/SubNZCV op
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.

Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 07:52:12 -04:00
Ryan Houdek a2307f28d6 Arm64: Optimize AES operations by caching a zero register
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.

Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.

AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
2023-08-30 19:00:43 -07:00
Ryan Houdek 1446d4fe12 IR: Adds support for named vector zero
This is useful for caching a zero register vector which we use in
various locations. This will be abused soon.
2023-08-30 18:59:38 -07:00
Ryan Houdek d8f131fa3d Arm64: Optimize AESKeyGenAssist
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.

I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
2023-08-30 12:15:09 -07:00
Ryan Houdek 62a9a075b7 Arm64: Leave a comment that 32-bit division shouldn't leave garbage in upper 64-bits 2023-08-28 05:02:00 -07:00
Ryan Houdek 6f2b3e76ac Merge pull request #3013 from Sonicadvance1/32bit_sra
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek 1d7c280367 Merge pull request #3012 from Sonicadvance1/optimize_movmskps
OpcodeDispatcher: Optimizes SSE movmaskps
2023-08-27 21:29:04 -07:00
Ryan Houdek 514a8223d9 OpcodeDispatcher: Optimizes SSE movmaskps
This now improves the instruction implementation from 17 instructions
down to 5 or 6 depending on if the host supports SVE.

I would say this is now optimal.
2023-08-27 21:07:20 -07:00
Ryan Houdek 8d110738ac IR: Add option to disable vector shift range clamping
The range check and clamping is necessary in the cases of passing x86
shift amounts directly through VUSHL/VSSHR.

Some AVX operations are still using these with range clamping. A future
investigation task should be the check if they can be switched over to
the wide variants that we implemented for the SSE instructions.

When consuming our own controlled data, we don't want the range clamping
to be enabled.
2023-08-27 21:07:20 -07:00
Ryan Houdek e4bb0df486 IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.

This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.

Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek 8f7925d06f Arm64: Simple typo fix 2023-08-26 18:22:50 -07:00
Ryan Houdek a01e69092d Arm64: Ensure Bfe and Sbfe operate at 32-bit or 64-bit op size
For Sbfe at least it ensures the upper bits don't get filled with
garbage.
Bfe it doesn't change behaviour but best to be correct.
2023-08-26 18:22:50 -07:00
Ryan Houdek 2fde2140ef Arm64: Ensure assert is testing correct array 2023-08-26 18:22:50 -07:00
Ryan Houdek 7f63d87295 IR: Adds support for new LoadNamedVectorIndexedConstant IR 2023-08-25 12:59:40 -07:00
Ryan Houdek 189b0da68f JIT/Int: Add support for scalar conversion as well 2023-08-25 03:19:11 -07:00
Ryan Houdek 62156f2152 ARM64JIT: Adds support for scalar cvt 2023-08-25 02:34:30 -07:00
Ryan Houdek 72ce7ddf2d Arm64: Optimize CVT operations for 64-bit variants
Using 128-bit converts for 64-bit versions cuts their throughput in half
on Cortex. Ensure we use the 64-bit version when possible.
2023-08-24 15:45:07 -07:00
Ryan Houdek ba01eac467 IR: Adds support for ARM's FCMA FCADD instruction 2023-08-24 15:00:41 -07:00
Lioncache 42ccc18606 x86_64/MemoryOps: Fix mislabeled IR op messages 2023-08-23 22:54:36 -04:00
Ryan Houdek 05b9651279 IR: Implements new vector multiply returning high bits
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Ryan Houdek 6e4765d48b Merge pull request #2989 from lioncash/ins
Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
2023-08-23 18:09:15 -07:00
Ryan Houdek 172c8f3ba6 Merge pull request #2988 from lioncash/half
Arm64/ConversionOps: Add missing half-precision conversions to scalar functions
2023-08-23 17:56:23 -07:00
Lioncache 203a2b1105 Arm64/ConversionOps: Remove redundant moves in AdvSIMD VInsGPR
If Dst and DestVector alias one another, then we don't need to
move the vector unnecessarily.
2023-08-23 20:50:21 -04:00
Ryan Houdek 4297e13fcf Merge pull request #2986 from lioncash/ext
Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
2023-08-23 17:40:50 -07:00
Lioncache 5ad56ad52e Arm64/ConversionOps: Add missing half-precision operations to Float_FromGPR_S
Provides parity with vector operations.
2023-08-23 20:36:34 -04:00
Lioncache 24e7baf28f Arm64/ConversionOps: Add missing half-precision conversions to Float_FToF
Provides parity with the vector conversion operations.
2023-08-23 20:31:49 -04:00
Lioncache b248ae4c04 Arm64/VectorOps: Remove redundant moves from SVE SQSHL
We don't need to emit a move if the destination and source alias.
2023-08-23 20:12:09 -04:00
Lioncache 95bea864cf Arm64/VectorOps: Remove redundant moves from SVE SRSHR
We don't need to perform a move is the destination aliases
the source vector to be shifted.
2023-08-23 20:10:51 -04:00
Lioncache 47c4507bb6 Arm64/VectorOps: Remove redundant moves from SVE VSQXTUN2
We don't need to perform a move if the destination aliases the lower vector.
2023-08-23 20:10:29 -04:00
Lioncache 5ea0b6db28 Arm64/VectorOps: Remove redundant moves from SVE VSQXTN2
We don't need to perform a move if the destination aliases the
lower vector.
2023-08-23 20:02:28 -04:00
Mai ee10153d14 Merge pull request #2984 from Sonicadvance1/optimize_pack
OpcodeDispatcher: Use new IR ops for pack instructions
2023-08-23 20:02:16 -04:00
Lioncache d0d94adabe Arm64/VectorOps: Remove redundant moves in SVE VExtr when possible
We don't need to do any moves here is the destination aliases the
lower bits.
2023-08-23 19:56:10 -04:00
Ryan Houdek 926b8c2c97 Merge pull request #2985 from lioncash/shift
Arm64/VectorOps: Remove redundant moves from SVE variable/immediate/vector shifts when possible
2023-08-23 16:41:39 -07:00
Lioncache 18ebcdc9de Arm64/VectorOps: Remove redundant moves in VUshrNI2
If the destination and VectorLower alias, then we don't need
to emit a movprfx.
2023-08-23 18:52:32 -04:00
Lioncache f31a9a52e6 Arm64/VectorOps: Remove redundant moves from SVE immediate vector shifts when possible
If the destination and source vector alias one another, then the
operation can largely be done in place.
2023-08-23 18:36:21 -04:00
Lioncache 03504a5f8c Arm64/VectorOps: Remove redundant moves from SVE vector shifts when possible
If the destination and the vector to be shifted alias, then we can
avoid needing to move some data around.
2023-08-23 18:24:58 -04:00
Lioncache d29b4de1ee Arm64/VectorOps: Remove redundant moves from SVE variable vector register shifts when possible
In the event that the destination and the vector to be shifted
alias one another, then we can skip the movprfx, since it's not
necessary.
2023-08-23 18:24:53 -04:00
Ryan Houdek c508570da0 IR: Implements VSQXT{U,}NPair operations
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Lioncache d5e145c4b0 Arm64/VectorOps: Remove redundant moves from SVE BSL when possible
If the destination and true vector alias one another, then we can
perform the operation in place instead of moving data around.
2023-08-23 17:54:10 -04:00
Ryan Houdek 350bca97c6 Merge pull request #2982 from lioncash/imin
Arm64/VectorOps: Remove redundant moves from SVE V{S,U}Min/V{S,U}Max when possible
2023-08-23 14:53:15 -07:00
Ryan Houdek 226405880f Merge pull request #2981 from lioncash/fmin
Arm64/VectorOps: Remove redundant moves from SVE VFMin/VFMax when possible
2023-08-23 14:46:03 -07:00
Lioncache 37a8cb6821 Arm64/VectorOps: Remove redundant moves from SVE VSMax when possible
When the destination and first source alias one another, then we
can perform the operation in place instead of moving data around.
2023-08-23 17:34:52 -04:00
Lioncache fe2c7dbf97 Arm64/VectorOps: Remove redundant moves from SVE VUMax when possible
When the destination and source alias one another, then we
can perform the operation in place without needing to move
data around.
2023-08-23 17:32:17 -04:00