Commit Graph
20 Commits
Author SHA1 Message Date
Ryan Houdek e9d96ce538 IR: Implements support for VTBL2
Skips implementing it for the x86 JIT because that's a bit of a
nightmare to think about.

The ARM64 implementation requires sequential registers which means if
the incoming sources aren't sequential then we need to move the sources
in to the two vector temporaries. This is fine since we have zero-cycle
vector renames and the alternative is slower.
2023-09-13 11:31:20 -07:00
Ryan Houdek 444d4c082d Int: Fixes typo in LoadNamedVectorIndexedConstant
Surprising this didn't break anything before this.
2023-09-13 11:31:20 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Ryan Houdek 1446d4fe12 IR: Adds support for named vector zero
This is useful for caching a zero register vector which we use in
various locations. This will be abused soon.
2023-08-30 18:59:38 -07:00
Ryan Houdek e4bb0df486 IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.

This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.

Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek 7f63d87295 IR: Adds support for new LoadNamedVectorIndexedConstant IR 2023-08-25 12:59:40 -07:00
Ryan Houdek 189b0da68f JIT/Int: Add support for scalar conversion as well 2023-08-25 03:19:11 -07:00
Ryan Houdek ba01eac467 IR: Adds support for ARM's FCMA FCADD instruction 2023-08-24 15:00:41 -07:00
Ryan Houdek 05b9651279 IR: Implements new vector multiply returning high bits
SVE implemented a new instruction that does this explicitly, so we
should support it directly.
2023-08-23 18:38:05 -07:00
Ryan Houdek c508570da0 IR: Implements VSQXT{U,}NPair operations
This takes the two independent VSXT{U}N{2,} operations and merges them
in to a single IR operations.
In some cases this can result in a more optimal implementation since
there is no need for moves inbetween.
2023-08-23 15:13:07 -07:00
Ryan Houdek 6aa2cab41c IR: Implement support for vector store element
Matches ARM64 ST1 semantics
2023-08-22 20:42:24 -07:00
Ryan Houdek 5e57ec94cf IR: Implement support for vector load element
Matches Arm64 LD1 semantics.
2023-08-22 20:15:16 -07:00
Ryan Houdek bbf9cb9d52 IR: Implements new VRev32 and LoadNamedVectorConstant ops
VRev32 matches Arm64 semantics directly.
LoadNamedVectorConstant allows FEX to quickly load "named constants".
This will allow us to have specific hardcoded vector constant values
that we can load with a ldr(State)+ldr(Value) and will be more abused in
the future.
This also allows us to do a very simple optimization in the future where
we can optimize away redundant loads of these loads if they are used
multiple times in the same block. (Not implemented here).
2023-08-22 16:29:06 -07:00
Lioncache f1d020ce95 Interpreter: Tie SSA data elements to supported vector width
Now, if we ever increase our vector sizes, the allocated data elements
will follow suit without needing to remember to handle this as well.
2023-08-21 14:41:00 -04:00
Lioncache a9a7cbce21 Interpreter: Use alias for temporary vector data
Lets us extract the size into one location for easy size
changes in the future if necessary.
2023-08-21 14:12:55 -04:00
Ryan Houdek a3bf952f2b IR: Implements support for wide scalar shifts
This matches x86 vector shift behaviour closely for ps{rl,ra,ll}{w,d,q}
where the vector is shifted by a scalar value that is 64-bits wide.
Anything larger than the element size will set that element to zero.

With SVE we have some new wide element shifts that match this behaviour
exactly (except supports wide shift sources rather than scalar).

This is a significant improvement even on platforms that only support
128-bit SVE.
2023-08-20 19:16:40 -07:00
Ryan Houdek a523858f66 Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek 1fdc4d2c62 IR: Implement support for a push operation
This is a bit of tricky operation where due to our our usage of SSA, the
incoming source isn't guaranteed to end its live-range at this
instruction.

This gives us a behaviour where to be optimal we need to take different
paths depending on if the incoming address register is the same as the
destination node.
Once we have form of RA constraints or non-SSA IR form that can
guarantee this restriction then this will go away.
2023-08-18 14:14:38 -07:00
Ryan Houdek d19e2507e5 FEXCore: Adds telemetry around legacy segment register setting
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.

Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.

It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.

Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.

A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.

[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
   - 3.6 - Restricted Subset of Segmentation
      - `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
        registers; the base for CS, DS, ES, and SS is ignored for 32-bit
        mode, same as 64-bit mode (treated as zero).`
   - 4.2.17 - MOV to Segment Register
      - Will fault if SS is written (Breaking anything that writes to
        SS).
      - Will not fault if CS, DS, ES are written (Thus it sets the
        segment but gets ignored due to 3.6).
2023-08-17 17:00:41 -07:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00