Now that NX is tracked in the frontend, we need to ensure that adjusted
RIP pages are tracked correctly. Keep around both instruction stream
pointers, validate the the original RIP is executable, and read from the
adjusted RIP as appropriate.
Fixes#5544
When the LOCK prefix is on an instruction that doesn't support LOCK then
it raises a SIGILL. Make sure to pass that up.
Additionally if the instruction does support lock prefix, has a lock
prefix, but the destination is not memory then that is also invalid.
Otherwise we hit an assert in FEXCore backend with code discovery
hitting things that look like AVX.
Fixes a crash in Uplay.
Also adds a test to just ensure that the instruction faults out and is
captured instead of crashing in FEX itself.
- Do compiler/architecture checks EARLY, don't waste time doing random
configuration stuff if the user can't even compile in the first place
- MSVC is unsupported, I assume? So add a check to disallow. There's
literally no MSVC or MSC_VER checks anywhere, so...
- Rather than using the MSVC architecture definitions, use our own
`ARCHITECTURE_arm64` et al. Hijacking existing "standard" definitions
is a very bad idea. Also makes it more readable in CMake
- Change the x86 host check to `x86|amd64`. Some systems still refer to
themselves as x86 despite being 64-bit for... reasons, and I saw one a
very long time ago that referred to it as amd64. This should
basically never come up, nor is it really relevant given that FEX is
for arm64... but it kinda annoyed me so whatever.
TODOs:
- Should we check `CMAKE_SIZEOF_VOID_P (equal) 64`? I don't think anyone
is even trying to compile this thing on armv7 or older, but might as
well? maybe?
- What's the status of *BSD, Solaris, macOS? Technically macOS does
support Wine, not sure about the others.
Signed-off-by: crueter <crueter@eden-emu.dev>
When relocations are loaded, all immediates read from memory are checked
against the relocation map and transformed into an appropriately
sign-extended entrypoint-relative variant of the specific operands
addressing mode. As almost every case of an unhandled relocation will
lead to a later, likely harder to debug, crash at runtime just bail out
early if any such cases are encountered. Note that while this
handles/detects all cases of relocated immediates, if relocations were
applied to instructions themselves (occurs in some malware variants)
these would be missed without any errors reported.
When compiling code at runtime there is no harm to including jumps to
different sections within a multiblock, when enforcing as such would
introduce a lookup cost for every decode invocation (or some caching).
However when compiling offline as each cache blob is tied to a specific
library these boundaries should be enforced.
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.
Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.
This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.
Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.
Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.
Right now just change the log so we can determine if we need to support
this edge case.
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.
If we packed this would get to a further 62 bytes but probably not worth
it.
Unity games crash with TSO disabled due to the SPSC GfxDevice
ThreadedStreamBuffer read/write pointer updates missing acq/rel
semantics. Rather than attempting to use heuristics to match cases like
this, which could be overzealous and hit more accesses than necessary, just
target the problem directly as this is consist across 32/64 bit Unity versions
for at least the past 10 years. Gate this behind the existing Unity mono
hacks to avoid false-positives in non-Unity games.
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
A wild use case of union over a variant because we don't want to
increase the encoding size from 128-bit to 256-bit (because of padding).
The type of operation is encoded with the table operation type, so a
variant is unnecessary and we get to keep the 128-bit encoding.
This allows us to have "recursive" x86 table descriptions. But in
reality this is going to only be one layer deep. As this will allow the
Frontend decoder to select instruction encodings based on arch bitness
once the tables are generated correctly.
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.
Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
Now executable page tracking is implemented, this limit is technically
unnecessary and effectively never hit in normal code. Keep a reasonable
limit however to avoid accidentally inlining tail calls and exploring
dead branches in obfuscated code.
Only the last prefix byte is retained when multiple are set. We were
accidentally generating a mask.
Additionally with 64-bit code, the legacy segment prefixes don't
overwrite if FS or GS have been set. So no weird behaviour where FS/GS
is set, a legacy prefix is used for padding, and then it "ignores" a bad
prefix by ignoring only the latest one.
This has the Frontend and OpcodeDispatcher select their operating mode
depending on the incoming code segment long-mode flag.
Adds some asserts since currently it is unexpected if the configuration
changes at runtime.
This is fairly straightforward for an initial setup but isn't fully
fleshed out.
Right now FEX's x86 tables aren't setup in a way to support choosing a
different instruction decoding depending on runtime operating mode
change, so that would break in interesting ways.
Primarily this just gets FEX setup to start piping the operating mode
through from the frontend to the backend. This is a long term task, so
it is going to take a long time to iron out all the issues.
Since CodePages is now a member of the guest to host map, which could
be replaced when JITing ARM code, any additions to it must be moved after that.
Additionally there is no benefit marking code pages for invalidation at all if
they are never added to the cache as in the single-step case.
This does technically prolong the window of an existing race where guest code
modifications could be missed, however this is unlikely to cause issues and didn't
prior.
Avoids an additional layer of indirection for callbacks. Passing them
around deep into instruction decoding logic doesn't provide much benefit
seeing as there will always be one frontend object per thread.