Commit Graph
1159 Commits
Author SHA1 Message Date
Ryan Houdek 082a0baff3 JIT: Implement missing Vector_FToF2 2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig be8ff9ccb9 Merge pull request #3740 from Sonicadvance1/avx_12
AVX128: More various instructions
2024-06-24 09:28:40 -04:00
Ryan Houdek 9c531d97b0 AVX128: Implements the various vector shift instructions
These are very closely related to each other so it makes sense to
implement the roughly three different families in one commit.
2024-06-24 09:20:19 -04:00
Ryan Houdek 6edf4619d4 Merge pull request #3742 from Sonicadvance1/export_avx_reg_helpers
FEXCore: Implement AVX reconstruction helpers
2024-06-24 05:57:44 -07:00
Ryan Houdek 8f769ce5a3 Merge pull request #3743 from alyssarosenzweig/cleanup/literal
X86Tables: add Literal() helper
2024-06-23 13:45:42 -07:00
Ryan Houdek d52a1da501 FEXCore: Implement support for fetching/setting YMM registers
Because we have two views of the YMM registers depending on if the host
supports SVE256 or not, add helper functions to fetch them correctly.

We fetch them in the way that Linux desires them in signal handlers, if
we want to return the converged view directly, that is easy to add
support for. It's unnecessary for now.
2024-06-21 17:13:56 -04:00
Ryan Houdek abdcaa7c86 AVX128: Implement support for vpinsr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek ad122cf463 AVX128: Implement support for vpmovmskb 2024-06-21 15:53:52 -04:00
Ryan Houdek b58a57d225 AVX128: Implement support for vmovmskp{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek 28d679de98 AVX128: Implement support for vpmov{s,z}{b,w,d}{w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek d1dd055e6a AVX128: Implement support for vpextr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 3045578da4 AVX128: Implement vmov{d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 9566dda73e AVX128: Implement support for vcmps{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek a0ced2b685 AVX128: Implement support for vcmpp{s,d} 2024-06-21 15:50:26 -04:00
Ryan Houdek df232f567b AVX128: Implement support for v{add,sub,mul,fmin,fmax,fdiv,sqrt,rsqrt,rcp}s{s,d} 2024-06-21 15:50:05 -04:00
Ryan Houdek 2a6d6a9d13 AVX128: Implement support for v{u,}comis{s,d} 2024-06-21 15:50:05 -04:00
Alyssa Rosenzweig cd03932bd1 OpcodeDispatcher: tweak InsertScalarFCMPOpImpl signature
so AVX128 can reuse it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 15:50:05 -04:00
Ryan Houdek 2e5fa1ef1b Merge pull request #3739 from Sonicadvance1/avx_11
Frontend: Expose AVX W flag
2024-06-21 12:15:14 -07:00
Alyssa Rosenzweig 0c6c4cd532 OpcodeDispatcher: make FCMP more compact
I told Ryan to change this for AVX, but it needs to be changed in the original
to match!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 15:02:10 -04:00
Alyssa Rosenzweig 25f8a87429 OpcodeDispatcher: use Literal() helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:58:49 -04:00
Alyssa Rosenzweig edf1a7970d X86Tables: add Literal() helper
Any time we get the value of Literal, we want to assert that it's actually a
literal. We've been open coding this pattern sporadically throughout the
opcodedispatcher. Let's add an ergonomic helper to fetch the value of literal,
asserting that the value is indeed literal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:46:46 -04:00
Ryan Houdek fac9972bad Merge pull request #3741 from alyssarosenzweig/cleanup/comiss
OpcodeDispatcher: refactor Comiss helper
2024-06-21 11:43:05 -07:00
Alyssa Rosenzweig 9ecb960f3a OpcodeDispatcher: refactor Comiss helper
AVX128 will use this, it's not SSE-specific.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:23:20 -04:00
Ryan Houdek 3d26e23891 Merge pull request #3737 from Sonicadvance1/avx_10
Arm64: Implement support for emulated masked vector loadstores
2024-06-21 11:04:01 -07:00
Ryan Houdek 7bbbd95775 Merge pull request #3736 from Sonicadvance1/avx_9
AVX128: Some pun pickles, moves and conversions
2024-06-21 10:55:19 -07:00
Ryan Houdek bb308899b9 Frontend: Expose AVX W flag
Previously we could always tell the size of the operation depending on
how this effects the operating size of the instruction. Converting
64-bit down to 32-bit as an example.

AVX gather instructions are the first instruction class that can't infer
this information. The element load size is determined by the W flag but
the operating size of 128-bit or 256-bit is determined by other means.

Expose this flag so we can determine this difference. The FMA
instructions are going to need this flag as well.
2024-06-21 10:54:42 -07:00
Ryan Houdek e95c8d703c Arm64: Implement support for emulated masked vector loadstores
In order to support `vmaskmov{ps,pd}` without SVE128 this is required.
It's pretty gnarly but they aren't often used so that's fine from a
compatibility perspective.

Example SVE128 implementation:
```json
    "vmaskmovps ymm0, ymm1, [rax]": {
      "ExpectedInstructionCount": 9,
      "Comment": [
        "Map 2 0b01 0x2c 256-bit"
      ],
      "ExpectedArm64ASM": [
        "ldr q2, [x28, #32]",
        "mrs x20, nzcv",
        "cmplt p0.s, p6/z, z17.s, #0",
        "ld1w {z16.s}, p0/z, [x4]",
        "add x21, x4, #0x10 (16)",
        "cmplt p0.s, p6/z, z2.s, #0",
        "ld1w {z2.s}, p0/z, [x21]",
        "str q2, [x28, #16]",
        "msr nzcv, x20"
      ]
    },
```

Example ASIMD implementation
```json
    "vmaskmovps ymm0, ymm1, [rax]": {
      "ExpectedInstructionCount": 37,
      "Comment": [
        "Map 2 0b01 0x2c 256-bit"
      ],
      "ExpectedArm64ASM": [
        "ldr q2, [x28, #32]",
        "mrs x20, nzcv",
        "movi v0.2d, #0x0",
        "mov x1, x4",
        "mov x0, v17.d[0]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[0], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[1], [x1]",
        "add x1, x1, #0x4 (4)",
        "mov x0, v17.d[1]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[2], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[3], [x1]",
        "mov v16.16b, v0.16b",
        "add x21, x4, #0x10 (16)",
        "movi v0.2d, #0x0",
        "mov x1, x21",
        "mov x0, v2.d[0]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[0], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[1], [x1]",
        "add x1, x1, #0x4 (4)",
        "mov x0, v2.d[1]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[2], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[3], [x1]",
        "mov v2.16b, v0.16b",
        "str q2, [x28, #16]",
        "msr nzcv, x20"
      ]
    },
```

There's a little bit of an improvement where nzcv isn't needed to get
touched on the ASIMD implementation, but I'll leave that for a future
improvement.
2024-06-21 08:21:32 -07:00
Ryan Houdek 903d6a742e CPUBackend: Removes SupportsSaturatingRoundingShifts option
This has always been true ever since we removed the x86 JIT and
Interpreter. This was left over and adding more code for no reason.
2024-06-21 08:11:22 -07:00
Ryan Houdek 424218e327 AVX128: Implement support for vpsign{b,w,d} 2024-06-21 08:11:22 -07:00
Ryan Houdek 17dc03d414 AVX128: Implement support for vpack{s,u}{wb,dw} 2024-06-21 08:11:21 -07:00
Ryan Houdek baf699c6e1 AVX128: Implements support for vandnps and vpandn
This can't use the previous binary operator handler since the register
sources need to be swapped.
2024-06-21 08:11:21 -07:00
Ryan Houdek 1431af1ff5 AVX128: Implements support for vcvt{t,}s{s,d}2si 2024-06-21 08:11:21 -07:00
Ryan Houdek 775a41b903 AVX128: Implement support for vcvtsi2s{s,d} 2024-06-21 08:11:21 -07:00
Ryan Houdek e614340c0c CPUID: Update labeling on some reserved bits
These aren't reserved and I was confused that they were missing.
2024-06-21 05:34:44 -07:00
Ryan Houdek 3c293b9aed Arm64: Loosen restrictions on V{Load,Store}VectorMasked to allow 128-bit operation 2024-06-21 04:26:09 -07:00
Ryan Houdek 283c2861c9 AVX128: Implement suppor for vlddqu 2024-06-21 00:56:36 -07:00
Ryan Houdek 757dc95116 AVX128: Implement support for the punpckh instructions 2024-06-21 00:56:32 -07:00
Ryan Houdek 6192250b8a AVX128: Implement support for the punpckl instructions 2024-06-21 00:56:28 -07:00
Ryan Houdek f489135b1d Merge pull request #3734 from Sonicadvance1/avx_8
AVX128: Move moves!
2024-06-21 00:53:41 -07:00
Ryan Houdek 3f232e631e Merge pull request #3730 from Sonicadvance1/avx_4
Vector: Helper refactorings
2024-06-21 00:31:14 -07:00
Ryan Houdek 6e3643c3ef Merge pull request #3714 from pmatos/FSTstiTagSet
Set tag properly in X87 FST(reg)
2024-06-21 00:27:24 -07:00
Ryan Houdek c28824f94d AVX128: Implements support for vbroadcast* 2024-06-20 09:43:10 -07:00
Ryan Houdek 664d766b45 AVX128: Implement support for vmovshdup 2024-06-20 09:43:10 -07:00
Ryan Houdek fce694ed92 AVX128: Implement support for vmovsldup 2024-06-20 09:43:10 -07:00
Ryan Houdek 96aafb4f07 AVX128: Implement support for vmovddup
This instruction is a little weird.
When accessing memory, the 128-bit operating size of the instruction
only loads 64-bits.
Meanwhile the 256-bit operating size of the instruction fetches a full
256-bits.

Theoretically the hardware could get away with two 64-bit loads or a
wacky 24-byte load, but it looks like to simplify hardware they just
spec'd it that the 256-bit version will always load the full range.
2024-06-20 09:43:10 -07:00
Ryan Houdek dbaf95a8f3 AVX128: Implement support for vmovhps/d 2024-06-20 06:53:21 -07:00
Ryan Houdek e67df96ad9 AVX128: Implement support for movlps/d 2024-06-20 06:53:17 -07:00
Ryan Houdek 56de94578d AVX128: Implement support for vmovq 2024-06-20 06:53:13 -07:00
Ryan Houdek 06fc2f5ef0 AVX128: Implement support for non-temporal moves. 2024-06-20 06:53:09 -07:00
Ryan Houdek b3ba315cbd AVX128: Implements unary/binary lambda helper 2024-06-20 06:53:05 -07:00