Follow-up to #2327.
Split off from #2176 and improved.
32-bit signals are a bit more complex than 64-bit due to behaviour
changing depending on if `rt_sigaction` and `sigaction` syscall is used
and if `SA_SIGINFO` is passed in to the flags.
With `SA_SIGINFO` used, both turn in to an `RT` frame, which is encoded
differently than without `SA_SIGINFO`.
Additionally 32-bit signals support both regular Linux stack ABI and
`regparm(3)` ABI.
Without `SA_SIGINFO` then `siginfo_t` is removed from the signal handler
arguments, but most of the rest still remains.
Also two of the arguments to the signal handler are forced to be nullptr
with `regparm(3)`.
Split off from #2243 to remove each member individually.
Shaves 8-bits off of each IR op.
No need to cart around this data when it is constant for each operation.
Especially since most optimization passes don't need the data anyway.
Needed to add a new `GetRAArgs` to get the number of SSA arguments that
get RA versus `GetArgs` which returns all SSA arguments the IR operation
owns. This is what was causing #2243 to fail CI since it needs to know
the difference in some places.
Split off from #2243 to remove each member individually.
IR ops are hardcoded by operation to have a destination or not.
No need to have each operation have a boolean for determining if the
operation has a destination or not.
The number of places things need to know if the operation has
destination or not is better served by using a lookup instead.
This isn't really a performance issue, more just something that is gross
looking at while looking at unit tests.
Break /usually/ isn't abused heavily by games (Except Denuvo) so not
really a perf concern regardless.
This will be useful for keying specific executables to steamids.
This is sadly required because a bunch of games end up naming themselves
"game.exe" so we can't safely enable thunks for all things shipping a
generic name.
Segment registers are indexed significantly more than they are changed.
Pay the cost of indexing during the set and store rather than the per
register index.
Should be a fairly significant performance improvement for 32-bit
applications. At least on hardware that doesn't have a data dependent
prefetcher.
Breaks Steam atm and isn't clean.
This creates a generic interface that FEXCore can use for timeline
profiling. This allows us to create a generic interface which the
backend details are hidden so we can support multiple timeline profile
APIs.
The only API supported right now is ftrace/gpuvis. Which is extremely
lightweight of an interface with minimal overhead.
We must be careful here since in most cases will will have dozens of
FEX instances running at any given time. So a timeline profiler like
Microprofiler can have major issues since that only ever expects a
single process at a time.
Not enabled by default but just needs the `ENABLE_FEXCORE_PROFILER`
cmake option set to enable.
In the event that the host doesn't support the requirements for running
AVX-enabled applications (SVE2 with at least 256-bit wide vectors) but
still wanted to run regular SSE-enabled applications, they would be
taking a performance hit due to a load/store pessimization (necessary in
order for 256-bit loads/stores to work)
However, we can add an alternate view into the xmm data that would allow
those hosts to use the previous optimization, while still supporting
AVX-capable hosts.
While we were getting the application name for the application layer, we
were failing to store the filename for telemetry.
Save the filename we get for application layers and store it for the
telemetry file.
Otherwise these were just alway ending up as wine or wine-preloader.
glibc lazily initializes the SETXID signal handler until first thread
creation.
Once we create our first pthread, steal it back from GLIBC after the
fact.
Fixes SOMA again.
x86 has six instructions that will fault on us that we mostly handled. A
few of these weren't being handled correctly.
One problem is that RIP needs to synchronize differently depending on
which fault instruction it is. Some instructions fault at the
instruction RIP, some at the instruction afterwards.
Additionally some of the metadata generated around the signal delegation
wasn't correct.
With behaviour of all of these instructions changed, it will now be
easier to switch over to non-faulting guest synchronous signals off of
these instructions. This doesn't go far enough to change that behaviour
yet.
I noticed this when looking at Elden Ring's weird faulting behaviour
with ud2 and `int 0x2d`. This made me investigate since I had a
suspicion that we weren't handling both cases correctly.
With this change, Elden Ring is now stabilized and works under FEX.
This comes in the form of a new define and a new function.
Assuming assert doesn't quite cover 100% of our asserting logging use
cases so we need to break this out.
If the compiler can't see through side effects at compile time for the
predicate then it'll throw a warning.
This gives us a nice behaviour. In the case that asserts are enabled, we
fall down the assert checking path. So if an assert fails to predicate
we will do an assert log like normal.
The change comes when we build release without asserts, we use the
`__builtin_assume` definition to allow the compiler to optimize around
assumptions. Since these aren't runtime asserts in release build, this
gives us some small optimizations in various locations.
Clang will very specifically do additional optimizations when you
provide it `__builtin_assume` directives that can really be worth it.
Allows for a smaller struct size.
This movement of the struct members also requires us to modify the
DeadContextStorePass to take the new locations into account.
On the plus side, we get to remove all of the padding bits in the
CPUState struct, so we can remove the handling for them.
pressure-vessel overrides our rootfs when it does a pivot_root.
Since we are still communicating to the FEXServer we were pulling the
configured rootfs.
Instead check if we are in pressure vessel and avoid doing that.
This fixes FEX running under pressure-vessel.
(There may be some implications to this down the road with code caching
but let's worry about that later)