Commit Graph
88 Commits
Author SHA1 Message Date
Alyssa Rosenzweig 57978accc1 OpcodeDispatcher: Add heuristic to prefer rmif for InsertNZCV
Decent instcountci win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig 157f95b08f IR: Extend TestNZ to two sources
Unlock the power of AND.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig ef544fecf2 OpcodeDispatcher: Use predicated neg for x87 fild
Saves an instruction on non-CSSC platforms by deleting a redundant cmp.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig 5471367db1 OpcodeDispatcher: Optimize branches
For native cases. Big perf win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig b8265b1067 OpcodeDispatcher: Rework SelectCC for branches
Separate out the NZCV bits from the more complex stuff so we can specially
optimize the branches.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig 0d70c6a0d0 OpcodeDispatcher: Use cset for getting nzcv flags
Usually better in practice... some rotates are slightly regressed by this but
they were already terrible.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig bfa069c4d5 OpcodeDispatcher: Dirty NZCV on new blocks
This worked only by accident before.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig 61bdf64e15 OpcodeDispatcher: Cleanup PF select
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig b187a853e7 OpcodeDispatcher: Use NZCVSelect for SelectCC
Massively better codegen.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-10 10:46:58 -04:00
Alyssa Rosenzweig 279afd88bb OpcodeDispatcher: Generalize rmif trick
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 10:05:51 -04:00
Alyssa Rosenzweig c0a6d82025 OpcodeDispatcher: Use rmif for NZCV inserts
Optimizes piles of s/w flag generation on flagm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 3a03e1c93c OpcodeDispatcher: rework InsertNZCV
in prep for rmif.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 5336129b58 IR: Optimize sub/sbb
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig b6f6c84790 IR: Optimize tests
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 04e4993d9b OpcodeDispatcher: Add a kludge to save NZCV less
Some opcodes only clobber NZCV under certain circumstances, we don't yet have
a good way of encoding that. In the mean time this hot fixes some would-be
instcountci regressions.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig c1dbc28aa2 OpcodeDispatcher: Implement SaveNZCV
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig b3055523b4 IR: Switch to dedicated NZCV load/store
Semantics differ markedly from the non-NZCV flags, splitting this out makes it a
lot easier to do things correctly imho. Gets the dest/src size correct
(important for spilling), as well as makes our existing opt passes skip this
which is needed for correctness at the moment anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig 9b81a83894 OpcodeDispatcher: Add DeriveOp helper
The "create op with wrong opcode, then change the opcode" pattern is REALLY
dangerous. This does not address that. But when we start doing NZCV trickery, it
will get /more/ dangerous, and so it's time to add a helper and make the
convenient thing the safe(r) thing. This helper correctly saves NZCV /before/
the instruction like the real builders would. It also provides a spot for future
safety asserts if someone is motivated.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-07 12:05:00 -04:00
Ryan Houdek 5103f2d92b Merge pull request #3247 from alyssarosenzweig/refactor/nzcv-prereq
Preparatory patches for nzcv
2023-11-01 14:07:14 -07:00
Alyssa Rosenzweig 5522c6db9c OpcodeDispatcher: Use jump wrappers
Mostly automated replacement + renaming for build fixing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Alyssa Rosenzweig 367e1658ad OpcodeDispatcher: Add jump wrappers
These should always be used in the dispatcher rather than the raw jumps they
translate to, as they ensure that flags are flushed. Eliminates a class of bugs
that will become a lot easier to hit with the new nzcv work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-11-01 15:44:36 -04:00
Mai 9612b2fe4b Merge pull request #3240 from Sonicadvance1/optimize_palignr_zero
OpcodeDispatcher: Optimize palignr with zero immediate
2023-11-01 05:55:52 +01:00
Mai 77d92872bc Merge pull request #3212 from Sonicadvance1/dpp_opt
OpcodeDispatcher: Optimize 128-bit DPPS and DPPD
2023-11-01 04:01:05 +01:00
Ryan Houdek 13cd8b33a2 OpcodeDispatcher: Optimize palignr with zero immediate
These turns in to moves
2023-10-27 15:09:31 -07:00
Alyssa Rosenzweig d87155e4ee IR: Add infrastructure for modelling flag clobbers
Lots of instructions clobber NZCV inadvertently but are not intended to write to
the host flags from the IR point-of-view. As an example, Abs logically has no
side effects but physically clobbers NZCV due to its cmp/csneg impl on non-CSSC
hw. Add infrastructure to model this in the IR so we can deal with it when we
start using NZCV for things.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-10-23 10:21:47 -04:00
Ryan Houdek 887200e571 OpcodeDispatcher: Optimize 128-bit DPPS and DPPD
These instructions aren't super amazing due to the fact that they have
both a source mask and a destination duplication mask.

Setup a case where we can generate more optimal code in /most/ cases.

There are a few that still fall down a "bad" path for the result
broadcast but in most cases they are optimal. Still to be seen what
games typically use the broadcast mask as.

AVX in its infinite wisdom expanded DPPS to 256-bit, while leaving DPPD
to only support 128-bit still. This leaves the original implementation
alone for 256-bit DPPS since I don't want to break it.

This is another instruction that gets a free optimization when
SVE-128bit is supported!
2023-10-20 18:02:27 +02:00
Lioncache 2b67f87054 OpcodeDispatcher: Handle SSE vector moves into themselves a little better
Obviously, it's silly to do this, but we should still be generating
optimal code for this case (which is none at all).
2023-10-18 14:58:57 +02:00
Lioncache 47a0f14537 OpcodeDispatcher: Remove unnecessary 128-bit truncating moves from StoreResult
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.

This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.

Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
2023-10-17 11:07:04 +02:00
Lioncache 2304cfc530 OpcodeDispatcher: Remove prefixing from MemoryAccessType enum
Since this is an enum class, we don't need to add a prefix.
2023-10-16 03:10:33 +02:00
Lioncache 1a39de4509 OpcodeDispatcher: Put extra LoadSource options in a struct
Allows for easier expansion without needing to expand the function definitons.

Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
2023-10-15 21:18:00 +02:00
Ryan Houdek 3bff42e6a7 OpcodeDispatcher: Wire up support for the new scalar insert operations 2023-10-10 03:44:58 -07:00
Ryan Houdek 6403290019 FEXCore: Renames raw FLAGS location names to signify they can't be used directly
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.

SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.

These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.

When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.

If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.

Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
2023-10-08 11:51:11 -07:00
Alyssa Rosenzweig 92211bf8c6 OpcodeDispatcher: Add AllowUpperGarbage option
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:08:20 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Alyssa Rosenzweig c8519b0b87 OpcodeDispatcher: Remove LoadPF
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig e63871ed2e OpcodeDispatcher: Handle sub in CalculateOF
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig 223a6562ff IR: Support <32-bit TestNZ
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:08:26 -04:00
Alyssa Rosenzweig 699aa85c4b OpcodeDispatcher: Opt PF selection
Fold the and in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2d65a3677b OpcodeDispatcher: Optimize NZCV selects
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Ryan Houdek 8e9e87f631 OpcodeDispatcher: Removes non-explicit SelectCC function
Renames the explicit sized one to `SelectCC`
Cleans up a bit of duplicated code.
2023-09-21 01:56:38 -07:00
Ryan Houdek e4613477b1 FEXCore/Interface/Core: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek 6dbbd9ecfc OpcodeDispatcher: Duplicate SelectCC but with Explicit result size
This is a temporary measure as we are moving Select operations over to
explicit sizes. Once we remove all uses of SelectCC then it will get
removed.
2023-09-15 10:09:37 -07:00
Ryan Houdek 759cc0025a OpcodeDispatcher: Add a dirty flag for tracking NZCV status
Cached NZCV reads don't need to be written back at the end of the block.
This will remove one instruction from the end of some blocks.
2023-09-15 10:09:37 -07:00
Ryan Houdek 863331b117 FEXCore: Implements support for shifted bitwise ops
This wasn't implemented initially for the interpreter and x86 JIT.

This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.

The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
2023-09-11 13:17:35 -07:00
Ryan Houdek 4feb059f51 OpcodeDispatcher: Optimize the case of all flags invalidated
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
2023-09-10 20:16:29 -07:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig 2a44acb144 OpcodeDispatcher: Defer AF extract
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:37:35 -04:00
Alyssa Rosenzweig e5883fe892 OpcodeDispatcher: Extract CalculateAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:13 -04:00