We already have an equivalent header within FEXCore that's header only,
so we can adapt it to conform for both cases, allowing for removal of
one of them.
Moves the IR include into one of the more specific headers, which avoids dumping the IR header into any core bits that use the interface.
Also uncovered a missing header guard.
We definitely don't want the boolean null test to be able to be implicitly converted
(e.g. to an int or whatever else).
For example a non-explicit bool operator allows for silly things like:
NonMovableUniquePtr<...> ptr;
// ...
auto k = 5 + ptr;
to build without issue, which we should really force the user to be explicit about if it's *really* a desired behavior.
We can reduce includes, such as logging by specifying a concrete size for logging levels,
allowing the enum to be forward declared. We can also move FillHeader into the cpp file,
allowing the syscalls header to be removed.
Since all of the information comes from the Dispatcher, we can have the dispatcher
provide that information. This way we can also eliminate a bunch of now-redundant
public interface members and simplify the config setup within InitCoreImpl().
Conveniently, this also allows making all members of the dispatcher non-public.
The name itself is already qualified with SignalDelegator, so this can reasonably be outside the class itself.
This also allows for forward declarations of the config struct
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
Taking this very slowly because this is very fickle code. The frontend
needs to manage GDT and LDT, but before we get there, we need to
actually add support for LDT in the backend. Split the segments to two
arrays so the JIT can actually update their cached values correctly.
Still treats GDT and LDT as mirrors like how the JIT previously did (By
it ignoring the selector's TI bit).
Previously we were only storing the 32-bit base address which isn't
actually how segment descriptors work.
In reality segment descriptors are 64-bit descriptors that are laid out
in a particular layout depending on the 4-bit type value. In reality we
only care about code and data segment layouts since the rest are
bonkers.
Describe these descriptors correctly and setup a default code descriptor
for the operating mode that FEX is starting in.
This option is free and only enabled if the config option is set. Enable
it always at build time so that users can pick it up without enabling
the full gpuviz/tracy paths.
This can't be handled fully within FEXCore due to the frontend-specific
handling of guard pages. Frontends can populate this at init time and
are expected to handle setting the CPUState field and register as approriate.
Prevents invalidations being missed under the following circumstances:
Thread A JITs block A into the global codebuffer, adding the guest to host
mapping to its CodePages, thread A is then killed.
Thread B then performs SMC on block A. An exception will be triggered but
as CodePages was stored per-thread, and thread A is now killed when all
threads are iterated over by the frontend to perform invalidations it
will be missed.
The accumulator is introduced to handle the case where multiple threads
have the same code entry in their local caches but share the same codebuffer.
Consider a thread C in the above example that also has block A in its cache,
without an accumulator, when invalidating thread B the entrypoint of A is erased
from the shared guest to host map. So when C is invalidated, the local cache entry
for A is not removed since it was removed from CodePages when invalidating B.
Part of waitpkg is the TPAUSE instruction. This instruction gives an
RDTSC deadline to go in to a low power sleep mode with the CPU.
Semantically we can't implement umonitor and umwait with ARM's exclusive
monitor implementation, but a nop implementation is sane. Just need to
make sure to clear the pre-req flags.
This lowers power consumption of UE5 games since their job handler now
goes to a tpause based implementation instead of a `pause` spinloop
implementation.
Buffers are tied to the lifetime of their owned flag, and as that
is a member of PoolBufferWithTimedRetirement we must always unclaim here.
Avoids the need to manually remember this quirk (which was forgot for the
temporary compilation buffer in JIT.cpp) at every use-site.
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.
```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second
64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801
After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659
80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719
After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749
Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```
Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.
Disabled in the simulator because we can't easily support pairs of
vector registers being returned.