In the case of modrm + immediate then the immediate would end up
overwriting Src1 due to the the order of the decoding.
Changes Src1 and Src2 to an array and use a variable to index the array.
Causes a bit of code churn but fixes instruction decoding and allows
easier expansion in the future for instructions that have more sources
like AVX
Register class for loadstores must be declared upfront.
This was causing a pain point when we were trying to load <16byte in to
FPRs, which was requiring a GPR<->FPR dance.
These are whole vector shift instructions and they shift by bytes rather
than bits.
You only get an immediate offset for the instruction.
Implements two new IR ops to account for these instructions
This also doesn't match AArch64 behaviour 100%, requires two
instructions to emulate rather than the single one on the x86-side
Simple linear scan forward of instruction decoding isn't viable to do.
We need to break up the decoding at the block boundaries.
Otherwise our decoding gets in to the weeds and decodes trash.
This changes over CondJump to have a defined True and False branch path.
Additionally needing to change over everything to handle this.
This now means that a false branch condition will no longer assume
fallthrough past this instruction.
Currently the x86-64 JIT generates an additional jmp instruction on
fallthrough that isn't optimized away.
This is working towards improvements in our IR branching model and
improving IRValidation in doing so.
This is working towards the model that IRBlocks can only have 0-2
successors.
Ret/Exit - 0 successors
Jump - 2 successors
CondJump - 2 successors
TBD:
Syscall/CPUID - Breaking block?
Adds an EVEX table and throws some ops in to an unimplemented function.
This is necessary for multiblock in the future where it will see
unsupported instructions but not actually execute them.
I had to change how blocks are represented to make it easier to parse
This required a fairly substantial refactor that makes it so blocks are
represented differently and we can walk them sequentially.
This will make future analysis easier to deal with.
Had to rewrite the passes and core's parsing of the IR afterwards.
Moved RA in to a optimization pass to be shared between the JIT backends
This works because x86-64 and AArch64 RA can be identical.
Still doesn't support PHI nodes or spilling correctly, this is the first
step in the process of getting there.
Argumentless IR emitter functions were prone to generating invalid code.
Remove them from the python emitter and change the branch instructions
that were using them to a new version instead.
Adds NumUse tracking as well.