Commit Graph
501 Commits
Author SHA1 Message Date
Lioncache 30cb1aaaed IR: Add VPCMPESTRX fallback
In order to implement the SSE4.2 string instructions in a reasonable
manner, we can make use of a fallback implementation for the time
being.

This implementation just returns the intermediate result and leaves it
up to the function making use of it to derive the final result from said
intermediate result. This is fine, considering we have the immediate
control byte that tells us exactly what is desired as far as output
formats go.

Given that the result of this IR op will never take up more than
16-bits, we store the flags we need to set in the upper 16 bits of the
result to avoid needing to implement multiple return values in the JIT.

Also, since the IR op just returns the intermediate result, this can be
used to implement all of the explicit string instructions with a single IR op.

The implementation is pretty heavily documented to help make heads or
tails of these monster instructions.
2023-04-17 21:39:32 -04:00
Ryan Houdek 51afcb7143 RA: Use FindFirstSetBit helper 2023-04-15 18:41:35 -07:00
Ryan Houdek e98a46aa5f Review comments 2023-04-07 17:01:53 -07:00
Ryan Houdek 4d70f4fc4e Remove some unused headers now. 2023-04-07 17:01:52 -07:00
Ryan Houdek 4ab822aebb IRParser: Convert to fextl 2023-04-07 17:01:52 -07:00
Ryan Houdek 2d18156e15 AOT: Convert fstream to fextl and raw files 2023-04-07 17:01:52 -07:00
Ryan Houdek 001a086d85 Convert remaining fmt::format to fextl 2023-04-07 17:01:51 -07:00
Ryan Houdek 546a1edb55 CodeReview 2023-04-01 09:27:01 -07:00
Ryan Houdek 97daec3dba Review comments 2023-03-31 06:03:06 -07:00
Ryan Houdek 1eac7e7105 AOTIR: Convert to FHU to remove glibc 2023-03-30 16:28:34 -07:00
Ryan Houdek 1eb36b8b31 Convert a ton of things over to fextl 2023-03-30 16:28:33 -07:00
Ryan Houdek 465ecd9b19 Mark code regions that require glibc memory allocations.
This ensures that when we enable glibc fault testing these sections
won't break CI.
2023-03-30 16:28:33 -07:00
Lioncache 5abf9de8a5 IR: Add VStoreVectorMasked IR op
Will be used to implement the store variants of VPMASKMOV and
VMASKMOVP{D, S}
2023-03-29 14:03:20 -04:00
Lioncache eb8626c1f7 IR: Add VLoadVectorMasked IR op
Will be used to implement the load variants of VMASKMOVP{D, S} and
VPMASKMOV{D, Q}

Particularly useful, since with SVE this behavior can be collapsed into
two instructions (CMPGT followed by the relevant LD1 load instruction)
2023-03-28 01:57:25 -04:00
Ryan Houdek 7022b3b825 Review c_str() changes 2023-03-23 12:45:14 -07:00
Ryan Houdek 0ab2a550b1 FEXCore: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek e5867b89ed FEXCore: Convert sstream to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek 3f98eff6e5 AOT: Convert string to fextl 2023-03-16 03:21:46 -07:00
Ryan Houdek a649e6aedf IRParser: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek f77f243ae6 PassManager: Convert string to fextl 2023-03-16 03:21:45 -07:00
Ryan Houdek 606242472a Convert the rest of map to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek 8941b8a312 Convert the rest of vector to fextl 2023-03-16 03:03:08 -07:00
Ryan Houdek 95a4994200 Cache: Convert over to fextl
Requires #2542 merged first. First commit is cherry-picked from that PR.
2023-03-15 12:22:04 -07:00
Ryan Houdek 91330d6b86 IR: Convert deque to fextl 2023-03-15 12:22:04 -07:00
Ryan Houdek 4bde9f796a IR: Convert set to fextl 2023-03-15 12:22:03 -07:00
Ryan Houdek 1c0a0c1fe5 IR: Convert unordered_map to fextl 2023-03-15 12:14:10 -07:00
Ryan Houdek bac5eff296 RAPass: Convert unordered_set to fextl 2023-03-15 12:10:32 -07:00
Ryan Houdek af91430007 IR: Changes vector to fextl 2023-03-15 12:04:26 -07:00
Ryan Houdek 042b511126 IR/Passes: Changes vector to fextl 2023-03-14 11:53:35 -07:00
Ryan Houdek 64c45ed70a Arm64: Reclaim SRA registers on 32-bit
Causes Portal to go from 85FPS to 120FPS on Lenovo X13s.

When running in 32-bit mode we were wasting 8 GPRs and 8 FPRs by still
allocating the top 8 registers of each even though 32-bit can't use
them.

Reallocate them to be register allocated registers when running 32-bit
applications which reduces spills and lowers the cost of
spilling/filling SRA registers.

Quite a significant speed boost for a little bit of work.
2023-03-14 07:15:05 -07:00
Ryan Houdek 81e52ca19f IR: Implements a memcpy operation
This matches the x86 REP MOVS instruction behaviour.
2023-03-06 16:38:43 -08:00
Ryan Houdek d6f50bf7b0 IR: Implement support for MemSet operation
This operation directly matches what the x86 STOS instruction does
without supporting its faulting behaviour.

STOS faulting behaviour is that RCX and RDI get updated to the last word
written. Which is something that FEX hasn't ever supported.
2023-03-03 09:16:28 -08:00
Lioncache 8beae0fce4 OpcodeDispatcher: Add VTrn/VTrn2 IR opcodes
Provides a convenient way to propogate indices at given intervals in
vectors. This makes permutation instructions a little less annoying to
implement.
2023-02-20 16:44:14 -05:00
Ryan Houdek 35af4bd42a FEXCore: Removes C wrapper interface
This has been a long time coming. The C interface has been a thorn in
our side for no reason for a long time.

The purpose of this step is to remove the C interface without changing
behaviour as much as possible. This means that with this commit there
are still some bad practices but the remaining issues will be solved
with followup PRs.

Primarily, we still have a `DestroyContext(CTX)` static function which calls
the Context implementation's `DestroyContext` and does a raw C++ delete.

Follow up PR will remove that, but I didn't want to touch it yet since
it'll require checking to ensure the unique_ptr changes play nice with
our allocator hooking. Which this is already a huge PR without trying to
change behaviour.
2023-02-19 11:59:11 -08:00
Lioncache 4bb7f49c2a IR: Add VDupFromGPR
Allows broadcasting constants into vectors from GPRs. Resolves the only
remaining TODOs within our vector ops.
2023-02-13 16:52:47 -05:00
Mai 143ef57141 Merge pull request #2345 from Sonicadvance1/user_sigreturn
Support user supplied signal restorer.
2023-02-08 20:11:47 -05:00
Lioncache 0218c966bd IR: Allow specifying register size for PCLMUL
This will allow us to support 256-bit vector operation in the future.
2023-02-08 16:35:20 -05:00
Lioncache ec5bc9cf3e IR: Allow specifying register sizes for AES enc/dec ops
This will allow us to support operating on 256-bit vectors.

Currently only sets up the bits and pieces on the x86-64 side, since
facilities for testing the 256-bit operations on ARM isn't set up yet.
2023-02-08 16:25:19 -05:00
Lioncache acbfee55b4 IR: Allow provising register size for VBSL
Necessary, since this will now be used with both 256-bit and 128-bit
registers, rather than just 128-bit.
2023-02-06 23:04:26 -05:00
Ryan Houdek ba5ad72ca2 JIT: Adds a JIT data header and tail.
This will be used to store various bits of data about the code going
forward.

Currently unused but that will change as we move forward.
2023-02-06 14:06:37 -08:00
Ryan Houdek abc596c634 IR: Removes SignalReturn op
This will no longer be used as we are swithing over to using the Linux
system call directly.
2023-02-04 10:35:06 -08:00
Ryan Houdek fa1193f14c Merge pull request #2344 from Sonicadvance1/siginfo_32
FEXCore: Fixup 32-bit signal handling
2023-01-31 20:26:36 -08:00
Mai 7be2e1ad34 Merge pull request #2330 from Sonicadvance1/implement_flushes
OpDispatcher: Adds support for CLWB and CLFLUSHOPT
2023-01-31 04:01:26 +00:00
Ryan Houdek d75e1f996f FEXCore: Fixup 32-bit signal handling
Follow-up to #2327.

Split off from #2176 and improved.

32-bit signals are a bit more complex than 64-bit due to behaviour
changing depending on if `rt_sigaction` and `sigaction` syscall is used
and if `SA_SIGINFO` is passed in to the flags.

With `SA_SIGINFO` used, both turn in to an `RT` frame, which is encoded
differently than without `SA_SIGINFO`.
Additionally 32-bit signals support both regular Linux stack ABI and
`regparm(3)` ABI.

Without `SA_SIGINFO` then `siginfo_t` is removed from the signal handler
arguments, but most of the rest still remains.
Also two of the arguments to the signal handler are forced to be nullptr
with `regparm(3)`.
2023-01-30 13:30:15 -08:00
Ryan Houdek 14fe95bd14 IR: Removes NumArgs member from IR ops
Split off from #2243 to remove each member individually.

Shaves 8-bits off of each IR op.
No need to cart around this data when it is constant for each operation.
Especially since most optimization passes don't need the data anyway.

Needed to add a new `GetRAArgs` to get the number of SSA arguments that
get RA versus `GetArgs` which returns all SSA arguments the IR operation
owns. This is what was causing #2243 to fail CI since it needs to know
the difference in some places.
2023-01-30 11:53:05 -08:00
Mai f8e762fcfb Merge pull request #2319 from Sonicadvance1/remove_has_dest
IR: Remove HasDest member
2023-01-30 16:25:09 +00:00
Ryan Houdek c6d46801ad ConstProp: Pool inline constants
In large blocks we can be generating a ton of inline constants. But in
most cases these end up being 0, 1, or (1 << N).
Add these to a map and reuse if possible. Makes some IR blocks
significantly smaller for later optimization passes.
2023-01-24 12:58:29 -08:00
Ryan Houdek 4582c8d380 IR: Adds support for CacheLineClean and non-serializing clear
These will be used in the next commit.
2023-01-18 17:56:21 -08:00
Ryan Houdek 5c98db5f47 IR: Remove HasDest member
Split off from #2243 to remove each member individually.

IR ops are hardcoded by operation to have a destination or not.
No need to have each operation have a boolean for determining if the
operation has a destination or not.

The number of places things need to know if the operation has
destination or not is better served by using a lookup instead.
2023-01-08 17:57:47 -08:00
Ryan Houdek c9622f6fd4 unittests/IR: Update tests for new IR semantics 2022-11-22 23:06:18 -08:00