Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.
Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
Telemetry value address generation was forcing an indirection at all
times which was unnecessary. These values live in the BSS, zero
initialized at process start and is unnecessary.
Instead change the wrapper defines to directly operate on the enum
passed in which saves an indirection on all of these telemetry
operations (except for the ones in the JIT which are required to be PIC
compliant).
This also fixes an annoying warning about
`FEXCORE_TELEMETRY_STATIC_INIT` causing initialization and destruction
order being unspecified, so two wins.
These three options are mutually exclusive with each other and could
potentially result in invalid encodings of the table on accident.
Change over to a 2-bit bitfield to encode if the operand that consumes
the VEX option is none, destination, 1st src, or 2nd src.
This ensures the table can't ever be incorrectly encoded.
With the prior approach, backwards jumps into existing blocks would
explore the overlapping part rather than splitting the block, generating
needless code and wasting time decoding. Similarly, the current block
wouldn't be split when it is extended to overlap with a pending jump target.
Solve this by tracking blocks in a sorted vector and splitting existing blocks
on jumps when appropriate, in order to avoid any possibility of overlapping
blocks, which would break the lookup, misaligned and zero instruction blocks
are disallowed.
This avoids both the generation of multiblocks that cover massive spans
of guest code, which causes issues for both context reconstruction
overflowing the RIP offset and attempting to decode branch targets
in unmapped memory regions.
Once support for querying mappings from the FEX frontend is in place this
limit could be increased if necessary, but this seems fine for now.
'add [rax], al' is almost never seen in actual code so the assumption
can be made that we are most likely trying to explore garbage code and
that this will never be hit. If it is then code will be generated at
that point (where Entrypoint == true).
Currently FinalInstruction causes only to the currently decoding block
to be terminated, but that is not enough as both MaxInst and
DefaultDecodedBufferSize are global limits that apply across all blocks
within a multiblock.
A source of overhead with multiblock is hitting instructions through a
conditional branch that can never be executed. Usually AVX512
instructions in glibc. This causes us to emit partial blocks for a ton
of targets that will never get executed.
Instead, when we have multiblock enabled, if a block hits an instruction
encoding we don't support, then remove all the decoded instructions from
the block and early terminate it if it isn't the entry block. This
resolves the issue of emitting a bunch of IR and code for blocks never
executed.
If the block of code has an invalid instruction in the entry block for
decoding then it'll still emit code up to the invalid instruction and
raise a SIGILL. This has the potential for generating some additional
blocks of code if a game is abusing SIGILL, but since that's unlikely
it's a good trade-off.
Also removes a few log instructions that don't really provide anything
anymore and just show up as confusing messages when multiblock is
enabled.
This is no longer necessary and it also no longer provides us any useful
information. Since we expose the AVX CPUID flag, basically everything
uses VEX encoding now, so it is basically always set.
In regular SIB land the index register encoding of 0b100 encodes to "no
register", this feature lets you get SIB encodings without an index
register for flexibility.
In VSIB encoding this isn't expected behaviour and instead there are no
encodings where an index register is missing. Allowing you to encode all
sixteen registers as an index register.
This was causing an abort in `AVX128_LoadVSIB` because the index turned
in to an invalid register.
Working instruction:
`vgatherdps ymm2, dword [eax+ymm5*4], ymm7`
Broken instruction:
`vgatherdps ymm0, dword [eax+ymm4*4], ymm7`
This fixes a crash in libfmod where it is using gathers in the wild.
Fixing a crash in Ender Lilies.
Previously we could always tell the size of the operation depending on
how this effects the operating size of the instruction. Converting
64-bit down to 32-bit as an example.
AVX gather instructions are the first instruction class that can't infer
this information. The element load size is determined by the W flag but
the operating size of 128-bit or 256-bit is determined by other means.
Expose this flag so we can determine this difference. The FMA
instructions are going to need this flag as well.
FEXCore includes was including an FHU header which would result in
compilation failure for external projects trying to link to libFEXCore.
Moves it over to fix this, it was the only FHU usage in FEXCore/include
NFC
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.
eg:
```asm
; Can be optimized to a single stp
push eax
push ebx
; Can remove half of the copy since we know the direction
cld
rep movsb
; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```
This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.
There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.
Example in the json
```json
"push ax, bx": {
"ExpectedInstructionCount": 4,
"Optimal": "No",
"Comment": "0x50",
"x86Insts": [
"push ax",
"push bx"
],
"ExpectedArm64ASM": [
"uxth w20, w4",
"strh w20, [x8, #-2]!",
"uxth w20, w7",
"strh w20, [x8, #-2]!"
]
}
```
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>