With multiblock enabled, host code generated from guest code with a
lower address may be placed after host code generated from guest code
with a higher address in a multiblock. As each guest RIP reconstruction
entry is always relative to the one before it the offset needs to be
signed to allow this.
This avoids both the generation of multiblocks that cover massive spans
of guest code, which causes issues for both context reconstruction
overflowing the RIP offset and attempting to decode branch targets
in unmapped memory regions.
Once support for querying mappings from the FEX frontend is in place this
limit could be increased if necessary, but this seems fine for now.
'add [rax], al' is almost never seen in actual code so the assumption
can be made that we are most likely trying to explore garbage code and
that this will never be hit. If it is then code will be generated at
that point (where Entrypoint == true).
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.
Discussed in #4252.
Currently FinalInstruction causes only to the currently decoding block
to be terminated, but that is not enough as both MaxInst and
DefaultDecodedBufferSize are global limits that apply across all blocks
within a multiblock.
It's faster to load the f80 sign mask from our named vector constants
than synthesizing the values. Changes a 4 instruction sequence to
synthesize to be 1 load.
Most of this table ignores REX.W, but two encodings change behaviour
based on REX.W. These two encodings are PEXTRD/PEXTRQ and PINSRD/PINSRQ.
For every other instruction encoding, they will ignore REX.W, but FEX
was requiring that they didn't have REX.W encoding. I had special cased
this in the past by adding PALIGNR, but that didn't handle any of the
other instructions.
We can't just handle REX.W in the OpcodeDispatcher and remove the two
special cased instructions because these vector operations also interact
with instruction prefix 0x66 which changes the operating size to 16bit
with regular instructions.
So instead just generate all listings of instructions with REX.W being
zero and one and install handlers in all cases.
No need to extract the subregisters out before operating on them since
the long division and long remainder IR operations correctly zero/sign
extend the incoming sources as necessary. Saves a couple of
instructions.
guest instruction
Single instruction blocks need to be treated specially when inline SMC
is detected, the frontend only needs to reprotect RWX and invalidate
caches then continue execution as side effects from the SMC shouldn't be
seen until the instruction executes.
Frontends need to detect this in order to handle SMC within the current
block (inline SMC) differently to regular SMC which can just reprotect
and continue.
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.