When relocations are loaded, all immediates read from memory are checked
against the relocation map and transformed into an appropriately
sign-extended entrypoint-relative variant of the specific operands
addressing mode. As almost every case of an unhandled relocation will
lead to a later, likely harder to debug, crash at runtime just bail out
early if any such cases are encountered. Note that while this
handles/detects all cases of relocated immediates, if relocations were
applied to instructions themselves (occurs in some malware variants)
these would be missed without any errors reported.
When compiling code at runtime there is no harm to including jumps to
different sections within a multiblock, when enforcing as such would
introduce a lookup cost for every decode invocation (or some caching).
However when compiling offline as each cache blob is tied to a specific
library these boundaries should be enforced.
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
When loading code caches, these constants get patched up for the new guest
address. The new value may be larger than the original, so the padding bytes
ensure the maximum of 16 bytes of encoding space is always available.
Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.
Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
have their entries invalidated for the given range, these codebuffers
are tracked internally within FEXCore. It is at this point that delinking
callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
naively invalidating the L1/L2 caches for every invalidated block for
every thread, threads now track on their own what specific entries
have been potentially fetched into their L1/L2 caches. This is
aided by GuestToHostMap now tracking the pages each block touches. (an
inverse CodePages so to speak).
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.
Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.
This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.
Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.
Enables memcpy optimization of 80bit floats on reduced precision.
Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.