Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.
SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.
These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.
When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.
If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.
Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.
Deletes a pile of code both in FEX and the output :-)
ADC/SBC could probably get similar treatment later.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This wasn't implemented initially for the interpreter and x86 JIT.
This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.
The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AF is calculated as:
((Src1 ^ Src2) ^ Res)[4]
Due to the extract, this is equivalent to
((Src1 ^ Src2) ^ (Res ^ 1))[4]
We already store (Res ^ 1) as the PF byte. So, it suffices to store
AF Byte = Src1 ^ Src2
and then we can recover the flag value
AF = (AF Byte ^ PF Byte)[4]
This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:
* Both PF and AF written together, the coupling is correct.
* PF written but AF invalidated, irrelevant.
* Both invalidated, irrelevant.
None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.
Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.
This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.