Not supposed to touch flags at all, so don't! instead of making a terrible mess
of csels. a lot less instructions, and probably faster because the branch should
be predicted correctly in practice in hot loops.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Semantics differ markedly from the non-NZCV flags, splitting this out makes it a
lot easier to do things correctly imho. Gets the dest/src size correct
(important for spilling), as well as makes our existing opt passes skip this
which is needed for correctness at the moment anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Replace every instance of the Op overwrite pattern, and ban that anti-pattern
from the codebase in the future. This will prevent piles of NZCV related
regressions.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Requires #3238 to be merged first since this uses the tbx IR operation.
Worst case is now a three instruction sequence of ldr+ldr+tbx.
Some operations are special-cased, which definitely doesn't cover all
possible cases we could use without tbx, but as a worst case improvement
this is a significant improvement.
A bunch of blendps swizzles weren't optimal. This optimizes all swizzles
to be optimal.
Two instructions can be more optimal without a tbx but the rest required
tbx to be optimal since they don't match ARM's swizzle mechanics.
Instead of using VInsElement in pi2fw and pf2iw, just use uzp1 to ensure
we don't unintentionally add to RA pressure.
Additionally we can generate the constant needed for pmulhrw directly
using the movi instruction. Converts two instructions in to one.
Under FEX's current constraints this makes all 3DNow! instructions
optimal.
These instructions aren't super amazing due to the fact that they have
both a source mask and a destination duplication mask.
Setup a case where we can generate more optimal code in /most/ cases.
There are a few that still fall down a "bad" path for the result
broadcast but in most cases they are optimal. Still to be seen what
games typically use the broadcast mask as.
AVX in its infinite wisdom expanded DPPS to 256-bit, while leaving DPPD
to only support 128-bit still. This leaves the original implementation
alone for 256-bit DPPS since I don't want to break it.
This is another instruction that gets a free optimization when
SVE-128bit is supported!
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.
This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.
Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
Allows for easier expansion without needing to expand the function definitons.
Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.
SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.
These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.
When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.
If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.
Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
Suggested by Alyssa. Adding an IR operation can be a little tedious
since you need to add the definition to JIT.cpp for the dispatch switch,
JITClass.h for the function declared, and then actually defining the
implementation in the correct file.
Instead support the common case where an IR operation just gets
dispatched through to the regular handler. This lets the developer just
put the function definition in to the json and the relevent cpp file and
it just gets picked up.
Some minor things:
- Needs to support dynamic dispatch for {Load,Store}Register and
{Load,Store}Mem
- This is just a bool in the json
- It needs to not output JIT dispatch for some IR operations
- SSE4.2 string instructions and x87 operations
- These go down the "Unhandled" path
- Needs to support a Dispatcher function override
- This is just for handling NoOp IR operations that get used for
other reasons.
- Finally removes VSMul and VUMul, consolidating to VMul
- Unlike V{U,S}Mull, signed or unsigned doesn't change behaviour here
- Fixed a couple random handler names not matching the IR operation
name.
Using the cached zero value is less efficient than loading it in to the
register for all these cases.
Lets us use rename hardware more efficiently and removes a dependency
chain on a single register.
Original:
```
movi v2.2d, #0x0
mov z16.d, p7/m, z2.d
<... 16 more times>
mov z31.d, p7/m, z2.d
```
Result:
```
movi v16.2d, #0x0
<... 16 more times>
movi v31.2d, #0x0
```
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.
Deletes a pile of code both in FEX and the output :-)
ADC/SBC could probably get similar treatment later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.
Closes#3117
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The only reason we need to XOR arguments for AF is to get bit 4 correct. But if
the operand in question is known to have bit 4 clear, the XOR will be an
effective no-op and can be skipped. This saves an instruction in a bunch of
common cases, like inc/dec. If we dedicated a register to AF to eliminate the
store, we would not save an instruction from this but would still come out ahead
due to an eor turning into a (zero cycle?) mov that can be handled by the
renamer.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>