The usual tricks, also requires introducing a bare adc op to optimize adcs to,
but we wanted that anyway!
Also support a zero source, so we can calculate "foo + CF" in one instruction to
optimize the "lock adc" cases.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is used for instcountci to ensure instruction counts don't change
when a compiler supports this feature or not. Always runtime disable
when running in instcountci.
CMake option from #3394 can still be useful so leaving that in place.
- use better algorithm that is O(# set bits) instead of O(# total bits)
- eliminate spilling by careful management of our temporaries
- fix nzcv clobber bug (whoops)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The ARM64EC SRA layout will use x0-3 for x86_64 registers, as such any
arguments passed to C ABI functions need to proxy their arguments
through the temporaries and move as appropriate.
Primary goal for this is to ensure that the delinker doesn't need to
allocate any memory. This delinker can end up getting hit heavily with
JIT code so we don't want it to be allocating memory.
The delinker step of the JIT was using std::function with capture
lambdas that required memory allocation when unnecessary.
Because the compiler can't see through our std::function usage it could
never decompose these by itself.
By passing the Thread's frame and record to the function as arguments
then we can have the signature be a raw function pointer.
This fixes an area of concern from:
https://github.com/FEX-Emu/FEX/blob/main/docs/ProgrammingConcerns.md#stdfunction-and-lambdas
If the Dst register is allocated as VectorIndices or VectorTable,
using Dst as an operand to perform the tbx operation will result in an error.
For example:
%131(FPR0) i128 = LoadNamedVectorIndexedConstant u8:Tmp:RegisterSize, #0x6, #0xaa0
%132(FPR0) i128 = VTBX1 u8:Tmp:RegisterSize, %129(FPRFixed6) i32v4, %126(FPRFixed10) i16v8, %131(FPR0) i128
Since the tbx instruction's destination register is also the original operand,
this is consistent with the semantics of VTBX1. Therefore,
directly using VectorSrcDst as the destination operand for the tbx instruction is safe.
We can safely call virtual functions through the JIT with a little bit
of work.
FEX's JIT has quite a few steps before it gets to a syscall handler.
Before this commit:
JIT->static HandleSyscall->SyscallHandler::HandleSyscall->SyscallHandler
After this commit:
JIT->SyscallHandler::HandleSyscall->SyscallHandler
A bit hard to notice this when this interface can spin at 67-million
calls per second though.
This was a temporary header to help with when this header was migrated
to our public API headers.
It's temporary nature is no longer necessary, just get rid of it.
In some situations TestNZ is generated with a constant that is using a
constant that can't fit inside of the tst instruction.
This was found in libGLX with virgl, crashing invalid instruction
generation and crashing steamwebhelper
This usually happens on backwards memcpy where we know the direction of
the copy because the code will typically do as follows:
```
std
rep movsb
cld
```
This is because the direction flag is part of the ABI and needs to be
set back to the forward direction if it was modified.
This typically doesn't get picked up on forward copies because we won't
have visibility of a cld instruction in the block.
This optimization allows us to only emit half of the code for the memcpy
if it is a compile time constant.
There's definitely some future task that could assume forward direction
if unknown and recompile the code if the assumption has failed, but not
doing that here.
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).
To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).
The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>