Commit Graph
232 Commits
Author SHA1 Message Date
Ryan Houdek 9c531d97b0 AVX128: Implements the various vector shift instructions
These are very closely related to each other so it makes sense to
implement the roughly three different families in one commit.
2024-06-24 09:20:19 -04:00
Ryan Houdek abdcaa7c86 AVX128: Implement support for vpinsr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek ad122cf463 AVX128: Implement support for vpmovmskb 2024-06-21 15:53:52 -04:00
Ryan Houdek b58a57d225 AVX128: Implement support for vmovmskp{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek 28d679de98 AVX128: Implement support for vpmov{s,z}{b,w,d}{w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek d1dd055e6a AVX128: Implement support for vpextr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 3045578da4 AVX128: Implement vmov{d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 9566dda73e AVX128: Implement support for vcmps{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek a0ced2b685 AVX128: Implement support for vcmpp{s,d} 2024-06-21 15:50:26 -04:00
Ryan Houdek df232f567b AVX128: Implement support for v{add,sub,mul,fmin,fmax,fdiv,sqrt,rsqrt,rcp}s{s,d} 2024-06-21 15:50:05 -04:00
Ryan Houdek 2a6d6a9d13 AVX128: Implement support for v{u,}comis{s,d} 2024-06-21 15:50:05 -04:00
Alyssa Rosenzweig cd03932bd1 OpcodeDispatcher: tweak InsertScalarFCMPOpImpl signature
so AVX128 can reuse it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 15:50:05 -04:00
Ryan Houdek fac9972bad Merge pull request #3741 from alyssarosenzweig/cleanup/comiss
OpcodeDispatcher: refactor Comiss helper
2024-06-21 11:43:05 -07:00
Alyssa Rosenzweig 9ecb960f3a OpcodeDispatcher: refactor Comiss helper
AVX128 will use this, it's not SSE-specific.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:23:20 -04:00
Ryan Houdek 424218e327 AVX128: Implement support for vpsign{b,w,d} 2024-06-21 08:11:22 -07:00
Ryan Houdek 17dc03d414 AVX128: Implement support for vpack{s,u}{wb,dw} 2024-06-21 08:11:21 -07:00
Ryan Houdek baf699c6e1 AVX128: Implements support for vandnps and vpandn
This can't use the previous binary operator handler since the register
sources need to be swapped.
2024-06-21 08:11:21 -07:00
Ryan Houdek 1431af1ff5 AVX128: Implements support for vcvt{t,}s{s,d}2si 2024-06-21 08:11:21 -07:00
Ryan Houdek 775a41b903 AVX128: Implement support for vcvtsi2s{s,d} 2024-06-21 08:11:21 -07:00
Ryan Houdek 283c2861c9 AVX128: Implement suppor for vlddqu 2024-06-21 00:56:36 -07:00
Ryan Houdek 757dc95116 AVX128: Implement support for the punpckh instructions 2024-06-21 00:56:32 -07:00
Ryan Houdek 6192250b8a AVX128: Implement support for the punpckl instructions 2024-06-21 00:56:28 -07:00
Ryan Houdek f489135b1d Merge pull request #3734 from Sonicadvance1/avx_8
AVX128: Move moves!
2024-06-21 00:53:41 -07:00
Ryan Houdek c28824f94d AVX128: Implements support for vbroadcast* 2024-06-20 09:43:10 -07:00
Ryan Houdek 664d766b45 AVX128: Implement support for vmovshdup 2024-06-20 09:43:10 -07:00
Ryan Houdek fce694ed92 AVX128: Implement support for vmovsldup 2024-06-20 09:43:10 -07:00
Ryan Houdek 96aafb4f07 AVX128: Implement support for vmovddup
This instruction is a little weird.
When accessing memory, the 128-bit operating size of the instruction
only loads 64-bits.
Meanwhile the 256-bit operating size of the instruction fetches a full
256-bits.

Theoretically the hardware could get away with two 64-bit loads or a
wacky 24-byte load, but it looks like to simplify hardware they just
spec'd it that the 256-bit version will always load the full range.
2024-06-20 09:43:10 -07:00
Ryan Houdek dbaf95a8f3 AVX128: Implement support for vmovhps/d 2024-06-20 06:53:21 -07:00
Ryan Houdek e67df96ad9 AVX128: Implement support for movlps/d 2024-06-20 06:53:17 -07:00
Ryan Houdek 56de94578d AVX128: Implement support for vmovq 2024-06-20 06:53:13 -07:00
Ryan Houdek 06fc2f5ef0 AVX128: Implement support for non-temporal moves. 2024-06-20 06:53:09 -07:00
Ryan Houdek b3ba315cbd AVX128: Implements unary/binary lambda helper 2024-06-20 06:53:05 -07:00
Ryan Houdek e5a531e683 Vector: Refactor MPSADBWOpImpl so AVX128 can use it. 2024-06-20 06:43:57 -07:00
Ryan Houdek e2de57bd04 Vector: Refactor PSADBWOpImpl so AVX128 can use it. 2024-06-20 06:43:57 -07:00
Ryan Houdek 4eebca93e3 Vector: Refactor PSHUFBOpImpl. This will be reused for AVX128 2024-06-20 06:33:27 -07:00
Ryan Houdek 3919ec9692 Vector: Expose VBLENDOpImpl in the OpcodeDispatcher. It will be reused by AVX128 2024-06-20 06:33:21 -07:00
Ryan Houdek 02aeb0ac1a Vector: Restructure PMADDWDOpImpl. It's going to get reused for AVX128 2024-06-20 06:33:15 -07:00
Ryan Houdek 206544ad09 Vector: Reconfigure PMADDUBSWOpImpl, it's going to get reused for AVX128 2024-06-20 06:33:08 -07:00
Ryan Houdek 3854cd2b2f Vector: Restruture SHUFOpImpl. AVX128 is going to reuse it. 2024-06-20 06:32:58 -07:00
Ryan Houdek acbd920c9a OpcodeDispatcher: Adds initial groundwork for decomposed AVX operations
Only installs the tables if SVE256 isn't supported yet AVX is explicitly
enabled with HostFeatures, to protect accidental enablement early.

- Only implements 85 instructions starting out
- Basic vector moves
- Basic vector unary operations
- Basic vector binary operations
- VZeroUpper/VZeroAll

The bulk of the implementation is currently the handling for loading and
storing the halves of the registers from the context or from memory.

This means the load/store helpers must always return a pair unless only
requesting the bottom half of the register, which occurs with 128-bit
AVX operations. The store side then needing to consume the named zero
register if it occurs since those cases will zero the upper bits.

This implementation approach has a few benefits.
- I can pound this out extremely quickly
- SSE implementations are unaffected and don't need to deal with the
  insert behaviour of SVE256.
- We still keep the SVE256 implementation for the inevitable future when
  hardware vendors actually do implement it (Give it 8 years or
  something).
- We can actually unit test this path in CI once it is complete.
- We can partially optimize some paths with SVE128 (Gathers) and support
  a full ASIMD path if necessary.

One downside is that I can't enable this in CI yet because it can't pass
all unittests. but that's a non-issue since it is going to be in heavy
flux as I'm hammering out the implementation. It'll get switched on at
the end when it's passing all 1265 AVX unittests. Currently at 1001 on
this.
2024-06-20 08:44:14 -04:00
Alyssa Rosenzweig db0bdd48e5 Merge pull request #3729 from alyssarosenzweig/refactor/address-modes
OpcodeDispatcher: Refactor address modes
2024-06-20 08:18:33 -04:00
Alyssa Rosenzweig ec03831a21 OpcodeDispatcher: plumb A.NonTSO deeper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 9ca821316a OpcodeDispatcher: factor out DecodeAddress
this is the common guts of the load/store routines.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 025a060337 OpcodeDispatcher: extract IsNonTSOReg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 371d6f0730 OpcodeDispatcher: extract IsOperandMem
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig af8cfb79e5 OpcodeDispatcher: refactor zero vector loads
AVX128 is going to slam this, so make it more ergonomic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 11:44:46 -04:00
Ryan Houdek 1ce27a5e6b FEXCore: Disentangle the SVE256 feature from AVX
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.

While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.

Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.

This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
2024-06-17 17:20:32 -07:00
Alyssa Rosenzweig 4bd84eb523 OpcodeDispatcher: extract PF/AF invalidate helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig e2073dcd30 OpcodeDispatcher: extract safe Thunk
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig fd72669c7e OpcodeDispatcher: extract safe Break
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00