In most cases we aren't using JIT symbols but still creating the perf
map file.
Early check if we should generate the file or not, this way we stop
polluting the /tmp folder.
I misread the implementation details of this instruction when
implementing.
The pseudocode says `ST(0) = ST(0) ∗ 2^rndint(ST(1))` so I understood
the instruction to use the current rounding mode of the host to extract
the integer portion of `ST(1)`.
The actual implementation is in the details of the statement `the
integer portion of the floating- point value in ST(1).`
This behaves like round towards zero/truncate, additional hardware
testing and documentation reading confirms this.
Fixes#1584
This isn't correct and breaks games.
This makes the FREM and REM1 implementation the same.
While not 100% correct, it is still better than before.
New issues will be created to handle the differences in the future.
Fixes#1374.
Also fixes most of the HL2 issues, just not the seam issue.
Brings along a bunch of enhancements and ensures we always build against
the latest version.
Also fixes up a few issues that arose due to changes in fmt
Migrates lingering instances of the old logger over to fmt where
applicable. This allows removing some of the old defines and functions.
The only remaining usages of the printf-based variant of the logger is
in Tests/LinuxSyscalls/Syscalls.cpp for the strace handling.
This lets us have JITsymbols grouped by library.
Useful for determining where to thunk.
Sadly perf doesn't have an option to deduplicate regions by name, so
some external tooling is necessary to make it look nice.
This is useful for testing the ops that are emulated using different
precision.
Haven't found anything that changes behaviour but useful to keep around
This isn't quite a 100% clean sweep of IWYU.
There are some false positives where clang fails.
Additionally there are still a few missed in the frontend side of things
that I didn't get to
This information is only ever going to be offline. Will be useful for multiple reasons.
1) Searching for split lock usage in applications, which can be a programming bug.
a) This isn't visible on AMD systems and on Intel is a fairly new linux feature
2) Having more information about when an application breaks.
3) Useful for some minor profiling for devs looking for statistical data
C++ no-op functions can't optimize out the predicate arguments in all cases.
This was causing a problem where zero cost assertions weren't actually zero cost.
The only way to resolve this is to actually use macros sadly enough.
This will give a fairly hefty performance uplift with anything operating on IR.
Prevents a class of sneaky logic bugs from slipping through into the
codebase.
This also resolves a case of such a bug within the Decorder's ReadData()
where all 3 byte reads would be performed as if they were a 4 byte read.
Instead of having the configuration being loaded and stored in to a
frontend system. First moves the backing store of the configuration in
to FEXCore.
Each layer that is constructed then loads its particular configuration.
After the layers are loaded, then a meta layer is constructed that
merges the layers flat.
This means we will no longer hit the problem where a configuration is
stored in the "main" configuration file, then environment and arguments
passed manage to overwrite it.
The order of the layers going from inner most layer to outer most, with
outermost overwriting previous layer configurations is as follows:
Main < Global application < Local application < Arguments < Environment
One step that needs to be changed in the future is that FEXCore can then
just load its configuration from the layers directly since the data is
in FEXCore now. This will be reserved for a future change so we are less
disruptive.
This implements support for 130 of the 132 x87 ops as interpreter
fallbacks.
The two missing ops are the BCD load and BCD store instructions and can
be implemented another time.
This allows us to no longer have to rely on a libstdc++ modification to
handle their usage of long double. This instead works entirely through
using `long double` directly in the interpreter. Which will either use
real x87 on an x86 host. Or in the case of ARM devices, fall down
glibc's long double soft float path.
This doesn't necessarily need to be quick right now. It's more important
to have the compatibility improvement from this.