Ryan Houdek
ef4c4f6e9b
FEXCore: Disable vixl linking if vixl disasm or simulator is disabled
...
This was mostly there, just needed to remove some extraneous headers and
only insert vixl in to the library list if the options were enabled.
2024-08-16 07:29:41 -07:00
Ryan Houdek
933c65d805
Merge pull request #3956 from alyssarosenzweig/opt/pop-return
...
small optimizations for returns
2024-08-15 03:30:23 -07:00
Ryan Houdek
df0ecad15b
Merge pull request #3933 from bylaws/arm64-suspend
...
Support cooperative suspend on ARM64EC
2024-08-15 01:22:28 -07:00
Alyssa Rosenzweig
2dc92c122e
BranchOps: micro-optimize ExitFunction
...
this should be slightly faster on Firestorm and no worse on recent cortex
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 16:59:34 -04:00
Ryan Houdek
aa5d2ff31c
Merge pull request #3951 from alyssarosenzweig/opt/pops
...
Add a hack for multiple destinations & make good use of it
2024-08-14 12:03:00 -07:00
Alyssa Rosenzweig
200c6c054f
IR: introduce POP operation
...
rmw on a source, kind of terrible.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
bc0927b7b1
JIT: optimize moves for cmpxchg
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
64a45c0d29
IR: remove pairs
...
They're now unused. And won't be missed.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
cab02be637
IR: remove unused pair create/extract
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
74f341bc0e
IR: remove pair from cpuid
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
feaa1af1a8
IR: remove pair from XGetBV
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:37:06 -04:00
Alyssa Rosenzweig
a9c26cbf71
IR: add coalescing heuristics for pair replacements
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
b8cac9f7d5
JIT: avoid some moves with caspal
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
13974df204
IR: drop CASPair pair result
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
fa6fe9bf06
IR: drop cmppairz pairs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
746be0824e
IR: drop CASPair source pairs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
93120cabbb
IR: drop pair from memcpy
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Alyssa Rosenzweig
04ae05f4ce
IR: add hack for multiple destinations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-14 09:17:23 -04:00
Billy Laws
890e5e1f0f
FEXCore: Support disabling host cacheline clean/clear operations
2024-08-13 13:25:34 +00:00
Alyssa Rosenzweig
d9c779289c
OpcodeDispatcher: simplify RDRAND
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 15:18:00 -04:00
Alyssa Rosenzweig
5631ff4fd5
OpcodeDispatcher: optimize variable shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
3b8cd44ca4
IR: add NZCVSelectIncrement
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Billy Laws
a469047e7a
FEXCore: Support cooperative suspend on ARM64EC
...
ARM64EC doesn't have an explicit callback for suspend, instead this task
is delegated to the kernel setting a bool at a JIT-defined memory
location, as such the faulting based approach as was done for WOW64
cannot be used.
2024-08-09 14:34:17 +00:00
Ryan Houdek
a4d5302369
Arm64: Adds Int helpers
...
One more vixl step removed.
2024-08-03 21:40:28 -07:00
Ryan Houdek
6ff3c90af3
CodeEmitter: Removes vestigial vixl usage
...
- IsImmLogical already existed in our CodeEmitter. We just forgot to
allow nullptr arguments and to use it.
- Adds an equivalent IsImmAddSub helper and uses it
This gets us closer to removing vixl's global initializers from FEXCore.
2024-08-03 21:04:56 -07:00
Ryan Houdek
7816b150d0
FEXCore: Removes CPUBackendFeatures
...
We were only ever hardcoding true for TBL2 and Flags now. Get rid of it.
2024-07-24 17:19:30 -07:00
Ryan Houdek
3c5b59d985
AVX128: Implement support for scalar FMA with AFP
...
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.
Fixes #3793
2024-07-22 12:58:19 -07:00
Alyssa Rosenzweig
610caf8529
ConstProp: treat StoreContext as zeroable
...
todo: FPR equivalent.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-21 15:49:09 -04:00
Alyssa Rosenzweig
d20b46e46f
IR: drop LoadFlag/StoreFlag ops
...
pointless, we can just load/store the context now.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-21 15:49:09 -04:00
Ryan Houdek
95b15d788b
Arm64: Fix filling static registers
...
Some locations could end up with SRA registers that only spilled one
register.
Allow passing in temporaries from the call site.
Fixes rpid and syscalls asserting.
2024-07-20 15:57:01 -07:00
Ryan Houdek
b78da2e5ad
Arm64: Implements support for DAZ using AFP.FIZ
...
When AFP is supported then we can actually support DAZ. This might also
fix the audio corruption in Animal Well but I can't test it until Steam
is running on Oryon. Requires a bit of plumbing for MXCSR which we were
hacking around before but now we actually want to store the value.
Fixes #3856
2024-07-20 15:34:54 -07:00
Alyssa Rosenzweig
0c3a8d0bc8
IR: remove GetHostFlag
...
it doesn't get host flags, it's just an extra Bfe used in x87. pointless and
confusing!
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-16 14:44:34 -04:00
Ryan Houdek
d79b7fcc49
Merge pull request #3808 from alyssarosenzweig/rclse/3
...
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Mai
e25918d846
Merge pull request #3858 from Sonicadvance1/implement_nt_load
...
Implement support for SSE4.1/AVX NT loads
2024-07-11 14:22:41 -04:00
Alyssa Rosenzweig
3a334c4585
Reapply "IR: drop RCLSE"
...
This reverts commit 78aee4d96e .
2024-07-11 13:21:14 -04:00
Mai
b282620a48
Merge pull request #3857 from Sonicadvance1/sve_bitperm
...
Arm64: Implement support for SVE bitperm
2024-07-11 05:05:41 -04:00
Ryan Houdek
e24b01b6cb
Arm64: Implement support for SVE bitperm
2024-07-11 01:46:35 -07:00
Tony Wasserka
f19fe3b6f3
Fix warning about an expression with side effects being passed to __builtin_assume
...
LOGMAN_THROW_AA_FMT has no benefit over LOGMAN_THROW_A_FMT here, so just use
the latter.
2024-07-11 09:54:31 +02:00
Tony Wasserka
8d2b15665d
Fix unused-variable warnings
2024-07-11 09:54:30 +02:00
Ryan Houdek
4c21aa2604
Arm64: Implement support for NT Loads with ASIMD fallback
2024-07-10 23:06:46 -07:00
Mai
af6a0be832
Merge pull request #3842 from Sonicadvance1/fix_f64_to_i32
...
VCVT{T,}PD2DQ fixes and optimization
2024-07-09 03:49:31 -04:00
Ryan Houdek
d3d76aa8ce
IR: Adds new F64 -> I32 operation that changes behaviour depending on SVE
...
SVE added the ability to do F64 -> I32 conversions directly without an
fcvtn inbetween. So maybe sure to support them.
2024-07-09 00:38:47 -07:00
Ryan Houdek
3bea08da5f
Merge pull request #3843 from Sonicadvance1/remove_half_moves_fma3
...
Arm64: Remove one move if possible in FMA operations
2024-07-09 00:25:07 -07:00
Ryan Houdek
c5a0ae7b34
IR: Adds new QPS gather load variant!
2024-07-08 17:19:18 -07:00
Ryan Houdek
4bd207ebf3
Arm64: Moves 128Bit gather ASIMD emulation to its own helper
...
It is going to get reused.
2024-07-08 17:19:18 -07:00
Ryan Houdek
62cec7b6b2
Arm64: Remove one move if possible in FMA operations
...
If the destination isn't any of the incoming sources then we can avoid
one of the moves at the end. This half works around the problem proposed
in #3794 , but doesn't solve the entire problem.
To solve the other half of the moving problem means we need to solve the
SRA allocation problem for this temporary register with addsub/subadd, so it gets allocated
for both the FMA operation and the XOR operation.
2024-07-08 04:44:40 -07:00
Ryan Houdek
c168ee6940
Arm64: Implements VSSHLL{,2} IR ops
2024-07-06 18:32:35 -07:00
Alyssa Rosenzweig
5a3c0eb83c
OpcodeDispatcher: fix shl with 8/16-bit variable
...
the special case here lines up with the special case of using a larger shift for
a smaller result, so we can just grab CF from the larger result.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 18:38:12 -04:00
Alyssa Rosenzweig
1b552a6f62
JIT: fix ShiftFlags masking
...
we don't update flags for a nonzero shift that masks to zero.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 09:57:42 -04:00
Ryan Houdek
38a823cc54
Arm64: Fixes long signed divide
...
The two halves are provided as two uint64_t values that shouldn't be
sign extended between them. Treat them as uint64_t until combined in to
a single int128_t. Fixes long signed divide.
2024-07-04 16:42:23 -07:00