In order to implement the SSE4.2 string instructions in a reasonable
manner, we can make use of a fallback implementation for the time
being.
This implementation just returns the intermediate result and leaves it
up to the function making use of it to derive the final result from said
intermediate result. This is fine, considering we have the immediate
control byte that tells us exactly what is desired as far as output
formats go.
Given that the result of this IR op will never take up more than
16-bits, we store the flags we need to set in the upper 16 bits of the
result to avoid needing to implement multiple return values in the JIT.
Also, since the IR op just returns the intermediate result, this can be
used to implement all of the explicit string instructions with a single IR op.
The implementation is pretty heavily documented to help make heads or
tails of these monster instructions.
Will be used to implement the load variants of VMASKMOVP{D, S} and
VPMASKMOV{D, Q}
Particularly useful, since with SVE this behavior can be collapsed into
two instructions (CMPGT followed by the relevant LD1 load instruction)
Causes Portal to go from 85FPS to 120FPS on Lenovo X13s.
When running in 32-bit mode we were wasting 8 GPRs and 8 FPRs by still
allocating the top 8 registers of each even though 32-bit can't use
them.
Reallocate them to be register allocated registers when running 32-bit
applications which reduces spills and lowers the cost of
spilling/filling SRA registers.
Quite a significant speed boost for a little bit of work.
This operation directly matches what the x86 STOS instruction does
without supporting its faulting behaviour.
STOS faulting behaviour is that RCX and RDI get updated to the last word
written. Which is something that FEX hasn't ever supported.
This has been a long time coming. The C interface has been a thorn in
our side for no reason for a long time.
The purpose of this step is to remove the C interface without changing
behaviour as much as possible. This means that with this commit there
are still some bad practices but the remaining issues will be solved
with followup PRs.
Primarily, we still have a `DestroyContext(CTX)` static function which calls
the Context implementation's `DestroyContext` and does a raw C++ delete.
Follow up PR will remove that, but I didn't want to touch it yet since
it'll require checking to ensure the unique_ptr changes play nice with
our allocator hooking. Which this is already a huge PR without trying to
change behaviour.
This will allow us to support operating on 256-bit vectors.
Currently only sets up the bits and pieces on the x86-64 side, since
facilities for testing the 256-bit operations on ARM isn't set up yet.
Follow-up to #2327.
Split off from #2176 and improved.
32-bit signals are a bit more complex than 64-bit due to behaviour
changing depending on if `rt_sigaction` and `sigaction` syscall is used
and if `SA_SIGINFO` is passed in to the flags.
With `SA_SIGINFO` used, both turn in to an `RT` frame, which is encoded
differently than without `SA_SIGINFO`.
Additionally 32-bit signals support both regular Linux stack ABI and
`regparm(3)` ABI.
Without `SA_SIGINFO` then `siginfo_t` is removed from the signal handler
arguments, but most of the rest still remains.
Also two of the arguments to the signal handler are forced to be nullptr
with `regparm(3)`.
Split off from #2243 to remove each member individually.
Shaves 8-bits off of each IR op.
No need to cart around this data when it is constant for each operation.
Especially since most optimization passes don't need the data anyway.
Needed to add a new `GetRAArgs` to get the number of SSA arguments that
get RA versus `GetArgs` which returns all SSA arguments the IR operation
owns. This is what was causing #2243 to fail CI since it needs to know
the difference in some places.
In large blocks we can be generating a ton of inline constants. But in
most cases these end up being 0, 1, or (1 << N).
Add these to a map and reuse if possible. Makes some IR blocks
significantly smaller for later optimization passes.
Split off from #2243 to remove each member individually.
IR ops are hardcoded by operation to have a destination or not.
No need to have each operation have a boolean for determining if the
operation has a destination or not.
The number of places things need to know if the operation has
destination or not is better served by using a lookup instead.