Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
Originally this was going to use setf8/setf16, but it looks like the approach of
shift-and-test turns out to be faster. As a bonus this is a nice delete-the-code
win :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This wasn't implemented initially for the interpreter and x86 JIT.
This meant we are maintaining two codepaths. Implement these operations
in the interpreter and x86 JIT so we no longer need to do that.
The emitted code in the x86 JIT is hot garbage, but it's only necessary
for correctness testing, not performance testing there.
When flags are invalidated but we're going to insert a new flag we end
up in a situation where we loaded the prior value from memory, claimed
unknown cache status (they were all invalid!), and then did an insert.
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
AF is calculated as:
((Src1 ^ Src2) ^ Res)[4]
Due to the extract, this is equivalent to
((Src1 ^ Src2) ^ (Res ^ 1))[4]
We already store (Res ^ 1) as the PF byte. So, it suffices to store
AF Byte = Src1 ^ Src2
and then we can recover the flag value
AF = (AF Byte ^ PF Byte)[4]
This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:
* Both PF and AF written together, the coupling is correct.
* PF written but AF invalidated, irrelevant.
* Both invalidated, irrelevant.
None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The AF calculation is a Bfe of an XOR result. We can't defer the XOR (since it
combines multiple inputs into one), but we can & should defer the Bfe. Since AF
is written much more often than it is read, this should come out ahead.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
For now these are trivial to let us refactor without functional changes. Later
in this series, they will be made nontrivial to let us defer AF calculation.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.
Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.
This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.
Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.
Needs #2993 merged first.
Zero-extension will already occur if necessary upon storing.
Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.
Back to back bfi is actually less optimal due to dependency tracking.
With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.
Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.
It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.
Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.
A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.
[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
- 3.6 - Restricted Subset of Segmentation
- `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
registers; the base for CS, DS, ES, and SS is ignored for 32-bit
mode, same as 64-bit mode (treated as zero).`
- 4.2.17 - MOV to Segment Register
- Will fault if SS is written (Breaking anything that writes to
SS).
- Will not fault if CS, DS, ES are written (Thus it sets the
segment but gets ignored due to 3.6).