Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.
The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.
Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.
I didn't do that exercise since that can be followed up in a subsequent
PR.
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
This isn't actually used anywhere in the IR header, so we can
move it to where it's actually used.
Now the IR interface header doesn't have anything related to the
independent passes in it.
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.
This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.
Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
The dispatcher/block linker will handle this, but if the instruction
following a POPF flag doesn't otherwise trigger one of those the
interrupt would be missed.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
Hashmaps are super expensive and there's no reason not to use a vector - we
already have compact block IDs so we don't benefit from the sparseness. Huge win
for very little effort.
Spotted when profiling FEX. CondJump() in the JIT was almost 4% of our time (?!)
and all because of map slowness. Easy fix.
Difference at 95.0% confidence
-0.0196494 +/- 0.00194956
-3.92827% +/- 0.389753%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.
To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset
If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.
Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.
(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.
(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).
(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).
(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).
(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).
(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.
Been a while since I last looked at this, added a new instcountci file
to show the improvement.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.