Commit Graph
23 Commits
Author SHA1 Message Date
Ryan Houdek 4c409ea47d Merge pull request #2938 from lioncash/fcmp
OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
2023-08-19 19:27:24 -07:00
Ryan Houdek ac53913c37 Merge pull request #2937 from lioncash/scalarunary
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
2023-08-19 19:22:10 -07:00
Ryan Houdek 2224c23c79 Merge pull request #2934 from lioncash/scalarfp
OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
2023-08-19 19:20:20 -07:00
Ryan Houdek 6ce380d7a9 Merge pull request #2936 from lioncash/scalaralu
OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorScalarALUOp
2023-08-19 19:09:06 -07:00
Lioncache 5152854b98 OpcodeDispatcher: Remove extraneous moves in {V}CVTSD2SS/{V}CVTSS2SD
Since all we're going to be doing is an insert as the final operation,
in the cases where our source is a vector, we can specify the size of
the vector rather than the size of the element to avoid doing unnecessary
zero-extending.
2023-08-19 22:07:25 -04:00
Ryan Houdek 8dade7eea1 Merge pull request #2935 from lioncash/scalarfp2
OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
2023-08-19 19:05:52 -07:00
Lioncache 343b00818d OpcodeDispatcher: Eliminate unnecessary moves in {AVX}VFCMPOp
We dealing with scalar vector sources, we don't need to zero-extend
the vector, and we can just use it as is.
2023-08-19 21:20:03 -04:00
Lioncache 09addb217a OpcodeDispatcher: Remove unnecessary moves in {AVX}VectorUnaryOp
When dealing with source vectors, we can use the vector length
rather than using a smaller size and zero extending the register,
especially since the resulting value is just inserted into another
vector.
2023-08-19 20:09:19 -04:00
Lioncache 4d1f002dea OpcodeDispatcher: Remove unnecessary moves in AVXVectorScalarALUOp
Same thing as the SSE variant, but for AVX.
2023-08-19 19:22:42 -04:00
Lioncache 1158ad7b2a OpcodeDispatcher: Remove unnecessary moves in VectorScalarALUOp
We can explicitly specify the vector width when working with a
vector source, so that we don't do any unnecessary zero-extending
on the element.
2023-08-19 19:15:13 -04:00
Lioncache 6907fdca6b OpcodeDispatcher: Remove redundant moves from {V}CVTSD2SI/{V}CVTSS2SI
We can specify the full vector length when dealing with a source vector
to avoid zero-extending the vector unnecessarily. When dealing with a
memory operand, however, we only want to load the exact source size.
2023-08-19 18:46:08 -04:00
Lioncache c31329609f OpcodeDispatcher: Unify handling code for MOVSD and MOVSS
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
2023-08-19 17:44:45 -04:00
Lioncache 1fe8470933 OpcodeDispatcher: Remove extraneous moves from VMOVSS/VMOVSD xmm to mem case
Like the changes made to the xmm to xmm case, since we're going to be storing
a 64-bit value, we don't directly need to zero-extend the vector on a load.
2023-08-19 17:35:52 -04:00
Lioncache 99b5aaa426 OpcodeDispatcher: Remove extraneous moves in VMOVSS/VMOVSD register case
In the event that we have a full length vector, we can just load and move
from it, which gets rid of a little bit of mov noise. Since all we intend
to do is perform an insert from one vector into another, we don't need the
zero-extending behavior that an 64-bit vector load would do.
2023-08-19 17:35:12 -04:00
Lioncache bbed4d73ed OpcodeDispatcher: Improve VPERMQ/VPERMPD broadcast cases
For a bunch of cases that act as broadcasts (where all
indices in the imm8 specify the same element), we
can use VDupElement here rather than iterating through.
2023-08-19 01:08:30 -04:00
Ryan Houdek f09d9af3db Merge pull request #2922 from lioncash/psrld
OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
2023-08-17 17:05:35 -07:00
Lioncache 9e54ec2724 OpcodeDispatcher: Improve {V}PSRLDQ shift by 0
While it would be bizarre if this actually occurred frequently
in practice, we can still tune it so there's no subpar assembly
output in the cases it actually does happen.
2023-08-17 19:33:09 -04:00
Ryan Houdek 461ca6fe7c Merge pull request #2921 from lioncash/shift
OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
2023-08-17 16:03:16 -07:00
Lioncache 5a1f32c339 OpcodeDispatcher: Remove unnecessary conditionals in {V}PSLLIOp
PSLLIImpl already checks for and handles a shift value of zero.
2023-08-17 18:47:50 -04:00
Lioncache 01515cea2c OpcodeDispatcher: Improve VMOVDDUP output
We can make use of TRN1 here to collapse a bunch of these moves.
2023-08-17 18:30:44 -04:00
Lioncache 764c844225 OpcodeDispatcher: Improve output of {V}MOVSHDUP 2023-08-17 17:01:50 -04:00
Lioncache 31719aac6a OpcodeDispatcher: Improve output of {V}MOVSLDUP 2023-08-17 17:01:46 -04:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00