We can perform less moves by checking for scenarios where aliasing
occurs. Since addition is commutative (usually, general-case anyway),
order of inputs doesn't strictly matter here.
In the event no source vectors alias the destination,
we can just move the first source vector into it and
then perform the divide without needing to move afterword.
We can avoid needing to use movprfx here by moving
directly into the destination when possible and just
doing the UMAX directly.
Also expands the unsigned max tests to test values with
the sign bit set to ensure all behavior is caught.
Since SMAX performs a comparison and returns the max value regardless
of how the operands are provided, we can check for when the second
input aliases the destination.
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.
This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.
Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
Allows for easier expansion without needing to expand the function definitons.
Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
Notably this bugfix version also introduces support for formatting
std::atomic types and std::atomic_flag.
Also, of course keeps our tracked external up to date.
A couple of games were hitting these. Not sure how they were missed in
PR #3159 but adds the missing one.
Small rearrangement to make this easier as well. Hopefully thunk stuff
lands sooner rather than later to automate this for Vulkan.
Maybe `-isystem` instead of `-I` needs to be used unlike what #2076,
might depend on what is installed on the host system.
Some simple tests to showcase instructions that we can optimize.
- Back to back pushes could be optimized
- Back to back scalar vector operations can be optimized
- Show with AFP that back to back scalar is already optimal
- Also ensures we don't break this stage.
We don't need zero in the upper bits for a push.
Makes a couple variants optimal.
Adds missing tests to the 32-bit file, since only 32-bit can push a
32-bit register.
These IR operations are required to support AFP's NEP mode which does
vector insert in to the destination register. Additionally it gives us
tracking information to allow optimizing out redundant inserts on
devices that don't support AFP natively.
In order to match x86 semantics we need to support binary and unary
scalar operations that do a final insert in to a vector. With optional
zeroing of the top 128-bits for AVX variants.
A tricky thing is that in binary operations this means that the
destination and first source have an intrinsically linked property
depending on if it is SSE or AVX.
SSE example:
- addss xmm0, xmm1
- xmm0 is both the destination and the first source.
- This means xmm0[31:0] = xmm0[31:0] + xmm1[31:0]
- Bits [127:32] are UNMODIFIED.
FEX's JIT jumps through some hoops so that if the destination register
equals the first source register, then it hits the optimal path the
AFP.NEP will insert in to the result. AVX throws a small wrench in to
this due to changed behaviour
AVX example:
- vaddss xmm0, xmm1, xmm2
- xmm0 is ONLY the destination, xmm1 and xmm2 are the sources
- This operation copies the bits above the scalar result from the
first source (xmm1).
- Additionally this will zero bits above the original 128-bit xmm
register.
- xmm0[31:0] = xmm1[31:0] + xmm2[31:0]
- xmm0[127:32] = xmm1[127:32]
- ymm0[255:127] = 0
This causes these instructions to support a fairly large table depending
on if the instruction is an SSE or AVX instruction, plus if the host CPU
supports AFP or not.
So while fairly complex, it's handling all the edge cases and gives us
optimization opportunities as we move forward. Currently on non-AFP
supporting devices this has a minor benefit that these IR operations
remove one temporary register, lowering the Register Allocation
overhead.
In the coming weeks I am likely to introduce an optimization pass that
removes redundant inserts because FEX currently does /really/ badly with
scalar code loops.
Needs #3184 merged first.
When FEX is in the JIT we need to make sure to enable NEP and AH and
then disable when leaving.
Explicitly disabled when the vixl simulator is used since even
attempting to set the bits will cause it to fault out. Ensures
InstCountCI keeps working.
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.
eg:
```asm
; Can be optimized to a single stp
push eax
push ebx
; Can remove half of the copy since we know the direction
cld
rep movsb
; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```
This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.
There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.
Example in the json
```json
"push ax, bx": {
"ExpectedInstructionCount": 4,
"Optimal": "No",
"Comment": "0x50",
"x86Insts": [
"push ax",
"push bx"
],
"ExpectedArm64ASM": [
"uxth w20, w4",
"strh w20, [x8, #-2]!",
"uxth w20, w7",
"strh w20, [x8, #-2]!"
]
}
```
This adds all the missing atomic tests in to their own tests files.
This includes all of them except a few choice ones that are in their
original files.
- BTC, BTR, BTS are in their Secondary/SecondaryGroup files
- CMPXCHG, CMPXCHG8B, CMPXCHG16B are in their Secondary/SecondaryGroup
files
- These always imply lock semantics even without the prefix.
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.
SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.
These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.
When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.
If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.
Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.