Saves power and responds faster. Pass in the atomic to `WaitPred` with
the predicate checking if the buffer has been flushed yet. Same
behaviour as previous code but more efficient on our hardware.
Default to enabled because this is the config we expect by default.
In the future will get some benchmarking in various games like
Assassin's Creed, and Call of Duty.
A few games were generating "Can't handle adddress size".
I implemented 0x67 prefix handling for CMPSOp and SCASOP and improved
the error messages for the remainder. This will implement the address
modifier on 64bit systems, and keep issuing an error on 32bits.
Shared code buffer support introduced the concept of having a single
GuestToHostMaps shared across many threads. In the common case all
threads will share one however if e.g. a resize recently occured and
specific thread is yet to compile any code with the new codebuffer it
will still use the old GuestToHostMap. The current invalidation
approach handles this by repeatedly calling erase for every single
thread's GuestToHostMap, even if it is repeated. An accumulator is used
to ensure when two threads share a map, the L1/L2 cache entries in the
second thread will still be invalidated even if the the iteration for
the first thread removed them from the map.
Unfortunately this is incredibly slow in cases with many threads, as
a significant number of redundant map lookups and L1/L2 cache erasures
on threads that never even observed a given block can occur. Solve this
by introducing a two-pass model:
- First, all active codebuffers (and their associated GuestToHostMaps)
have their entries invalidated for the given range, these codebuffers
are tracked internally within FEXCore. It is at this point that delinking
callbacks are ran.
- Second, each thread will have its caches invalidated. But rather than
naively invalidating the L1/L2 caches for every invalidated block for
every thread, threads now track on their own what specific entries
have been potentially fetched into their L1/L2 caches. This is
aided by GuestToHostMap now tracking the pages each block touches. (an
inverse CodePages so to speak).
We currently rely on the frontend to keep track of threads and then
iterate over all threads to perform per-codebuffer operations. However
as codebuffers are shared between many threads (the common case is a
single code buffer across all) this ends up being inefficient. Introduce
a list of codebuffers to solve that (new codebuffers are very rare, so a
vector is plenty fine here for erasing invalid weak refs).
Adds it to the VDSO handling, it's not necessarily a VDSO function but
it behaves as such as it is in every single process. This means we get
to reuse the mapped page for every process when thunks are built,
shaving a page out of 32-bit processes.
Also, fixes a bug in guest VDSO symbol loading where clang sticks all
symbols in to `.dynsym` where gcc sticks them in to `.symtab`. Search
both. This effectively meant the couple of guest VDSO symbols were
always failing to get found, causing us to allocate yet another page on
32-bit. So effectively three pages stolen.
This also means we can remove the Linux specific X86HelperGen stuff from
FEXCore, only passing a single "VDSO" function pointer to the backend
for the dispatcher. Once again moving the Linux stuff to the frontend is
good.
Fixes an assert about about untracked noexec code `NoExec
instruction in entry block: FFFFE000` whenever thunk callbacks were
used.