Fixed offset x86 code doesn't quite solve the issue, so adjust this
heuristic just to get instcounci to stop flaking.
This code is going to heavily change soon anyway so +50 doesn't change
much.
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.
Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.
No functional change with this, just moving headers around.
x86 has a few prefetch instructions.
- prefetch - One of two classic 3DNow! instructions
- Prefetch in to L1 data cache
- prefetchw - One of two classic 3DNow! instructions
- Implies prefetch in to L1 data cache
- Prefetch cacheline with intent to write and exclusive ownership
- prefetchnta
- Prefetch non-temporal data in respect to /all/ cache levels
- Assumes inclusive caches?
- prefetch{t0,t1,t2}
- Prefetch data with respect to each cache level
- T0 = L1 and higher
- T1 = L2 and higher
- T2 = L3 and higher
**Some silly duplicates**
- prefetchwt1
- Duplicate of prefetchw but explicitly L1 data cache
- prefetch_exclusive
- Duplicate of prefetch
God Of War 2018 uses prefetchw as a hint for exclusive ownership of the
cacheline in some very aggressive spin-loops. Let's implement the
operations to help it along.
Folds reg+const memory address into addressing mode,
if the constant is within 16Kb.
Update instcountci files.
Add test 32Bit_ASM/FEX_bugs/SubAddrBug.asm
The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!
Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This usually happens on backwards memcpy where we know the direction of
the copy because the code will typically do as follows:
```
std
rep movsb
cld
```
This is because the direction flag is part of the ABI and needs to be
set back to the forward direction if it was modified.
This typically doesn't get picked up on forward copies because we won't
have visibility of a cld instruction in the block.
This optimization allows us to only emit half of the code for the memcpy
if it is a compile time constant.
There's definitely some future task that could assume forward direction
if unknown and recompile the code if the assumption has failed, but not
doing that here.
Unused since TestNZ+NZCVSelect accomplishes the same and good riddance. Might
bring them back later for tbz/tbnz, but certainly not in this Selectful form.
(I added them when I thought we were going to RA the flags. With the more
effective static approach, we don't need this for that.)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
via the 0 reg. We really need a more generalized approach to taking advantage of
wzr, but this optimizes the special case I care about for seta (saving a move to
make the impl optimal).
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This will replace Select soon, as it lets us take advantage of
NZCV-generating instructions and it doesn't clobber NZCV.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
When SubShift (LSL) occurs with both sources constant then optimize away
the calculation.
Additionally if add is found to have one immediate constant where the
inverse of the constant fits in to ImmAddSub range, then invert the
constant and change it in to a sub.
This optimizes the cases when direction flag is known upfront in an
instruction.
Lots of instructions clobber NZCV inadvertently but are not intended to write to
the host flags from the IR point-of-view. As an example, Abs logically has no
side effects but physically clobbers NZCV due to its cmp/csneg impl on non-CSSC
hw. Add infrastructure to model this in the IR so we can deal with it when we
start using NZCV for things.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
With the previous RCLSE pass optimization that fixes store->load
forwarding, this pass started optimizing harder.
This hit a bug with this vmov removal that previously didn't get hit.
In particular this would eliminate vmov IR operations even if they were
zero extending a vector.
Since we have dramatically cleaned up the amount of vmov IR operations
we are generating, remove this optimization entirely. In the games I
tested, the only game that hit this "optimization" was Ender Lilies and
it started generating broken code for the single block of instructions
that did.
Adds a unit test for this case just in-case it comes back in the future
for some reason.
Fixes an issue where Ender Lilies would flash the screen to black every
time an enemy hit the player character.
The range was slightly incorrect which mostly wouldn't have caused
issues.
The lowest byte would have just generated slightly less optimal code.
The upper byte could have generated broken code, which our CI couldn't
catch since TSO instructions only get enabled when multiple threads are
in-flight.
Easy enough to fix.
Add new synthetic condition codes that do an AND as their relational operator,
testing the result. This is 1 IR op for things like
(A & B) == 0 ? C : D
This can translate to
tst A, B
csel A, B, eq
In the future, if A is the NZCV register and B is a supported immediate, eg
(NZCV & 0x80000000) == 0 ? C : D
this will be able to translate to a single instruction with the appropriate
condition
csel A, B, pl
but that needs RA support.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is blocking performance improvements. This backend is almost
unilaterally unused except for when I'm testing if games run on Radeon
video drivers.
Hopefully AmpereOne and Orin/Grace can fulfill this role when they
launch next year.
FEX has a problem with large blocks that uses a ton of constants spread
throughout the block. Once a block gets large enough with enough
constants that have large live ranges, FEX slows down to unusable speeds
due to the register allocator spending more time calculating node
interferences than anything else in the program.
This adds a little heuristic to ensure that constants aren't reused if
the previous value is past a certain distance threshold. This threshold
works well enough that XeSS's pedantic initialization code doesn't have
issues now. See https://github.com/FEX-Emu/FEX/issues/2688 for more
information about that.
FEX itself should work to remove bad constant usages to make this pass
less necessary anyway. In most cases we are materializing duplicated 0,
1, and masks which could be done without a constant entirely.
Maybe once we've improve that enough we could remove this constant
pooling entirely.
To note, this doesn't fix the issue that XeSS causes our register
allocator, this is purely a heuristic workaround.