The dispatcher/block linker will handle this, but if the instruction
following a POPF flag doesn't otherwise trigger one of those the
interrupt would be missed.
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
These x87 f64 reduced precision operations don't use FCW so we don't
need to load it from the context. So just remove loading it. This falls
within noise while benchmarking.
Hashmaps are super expensive and there's no reason not to use a vector - we
already have compact block IDs so we don't benefit from the sparseness. Huge win
for very little effort.
Spotted when profiling FEX. CondJump() in the JIT was almost 4% of our time (?!)
and all because of map slowness. Easy fix.
Difference at 95.0% confidence
-0.0196494 +/- 0.00194956
-3.92827% +/- 0.389753%
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
The constraints introduced by shared code buffers make supporting
calls with the previous layout impossible. The main additional constraint
imposed by call-ret that if a host location is ever pushed onto the
call-ret stack, then it must forever be a valid jump target. While
this is reasonable in the: unlinked, direct linked, unlinked,
direct linked case; it's almost impossible to achieve in the: unlinked,
indirect linked, unlinked, direct linked case while ensuring
all backpatching cases are valid with the current approach.
To solve this introduce an additional layer of indirection, jump thunks,
these are emitted at the end of a multiblock and are used to handle the
two cases of calling the initial linker, and calling an indirect linked
block. Initially at the ExitFunction location a branch/call to a unique
jump thunk will be emitted, which will have the code layout:
00: b 0x8
04: br TMP1
08: ldr TMP1, <Shared exit linker>
0c: blr TMP1
10: HostCode
18: GuestRIP
20: CallerOffset
If a direct link can be performed, then the initial branch/call to the
jump thunk can be linked/unlinked to point to the jump thunk in a
single 32-bit atomic operation. For an indirect link, the HostCode
member is updated with a 64 bit atomic operation, and then a 32 bit
atomic operation is used to replace the branch at 00 with a load of
HostCode. Indirect unlinks are done by placing back the b 0x8 at 00.
Safety:
(1)
Sequential link (e.g. one waiting to lock, one locked and linking):
Linking is idempotent, would just rewrite the same data atomically.
(2)
Simultaneous link or simultaneous delink:
Impossible due to LookupCache locking.
(3)
Simultaneous link and execute:
(3.1)
Direct link: Either the direct link is observed at the thunk
callsite, or it is not observed and the linker is entered - this is
then just (1).
(3.2)
Indirect link: Either the branch at 00 in the thunk is observed
to be replaced with an ldr, in which case the modified HostCode
must be observed due to the cache flush. Alternatively the branch
replacement isn't observed and it's just (1).
(4)
Simultaneous unlink and execute:
(4.1)
Direct link: Either the jump to the jump thunk is seen, which must
be in its base unlinked state with the branch at 00 as that would
be inserted by any previous indirect unlink. In such a case the
linker would just be entered, giving (5). Alternatively the modified
jump isn't seen and it calls the original host code (which is fine).
(4.2)
Indirect link: If an ldr is seen at 00, then the rest of that sequence
will function fine as HostCode is left untouched. If a branch is seen
at 00, then it will just call the linker giving (5).
(5)
Sequential unlink then link:
Unlinking restores the callsite and jump thunk to their original
contents (aside from a modified HostCode). Linking then works as
usual.
This is made slightly awkward by the many potential orderings of blocks
and desire to support both fallthrough jumps and calls without additional
branches.
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
Now that the PF flag isn't using popcount, this is a win across the
board if the hardware supports it.
Been a while since I last looked at this, added a new instcountci file
to show the improvement.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
If a multiblock contains a call instruction, we know at the point
of compilation that the instruction after that call will likely be
jumped to at some point. Avoid redundant recompilation by tracking
such cases and including an entrypoint for that instruction in the
multiblock aswell.
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.
```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second
64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801
After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659
80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719
After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749
Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```
Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.
Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
this will eliminate an annoying special case in post-RA opts.
No difference proven at 95.0% confidence
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
the modifying thread. Old versions of the CodeBuffer persist as read-only
data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
decrease the reference count and eventually trigger deallocation of the
old version
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>