Commit Graph
484 Commits
Author SHA1 Message Date
Ryan Houdek 81434cd233 Merge pull request #4377 from bylaws/sidt
Implement SIDT/LSL
2025-03-04 17:10:32 -08:00
Ryan Houdek b7f58e68c5 Merge pull request #4363 from pmatos/ReciprocalsFix
Improve reciprocal estimate and tests
2025-03-01 00:58:45 -08:00
Billy Laws cdae654fe4 FEXCore: Somewhat implement LSL
Emulate by always returning failure, this deviates from both Linux
and Windows but shouldn't be depended on by anything.
2025-02-27 23:45:14 +00:00
Billy Laws a506c84bc7 FEXCore: Implement SIDT 2025-02-27 23:45:14 +00:00
Paulo Matos b5241e0f60 Improve reciprocal estimate and tests
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.

* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.

Fixes #4319.
2025-02-27 15:07:14 +01:00
Ryan Houdek e718fc35f8 OpcodeDispatcher: Reuse PSHUFD shuffle mask for sha data shuffling
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.

OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).

Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
2025-02-24 12:07:48 -08:00
Paulo Matos 5666a352d4 NFC: Code cleanup
Removing unused declarations.
Cleaning up unused headers and empty lines.
Avoiding static analysis warnings on `const auto` defaulting to int.
2025-01-22 10:22:25 +01:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Paulo Matos cbda688e29 Revert "Cache predicate register generation from pattern"
This reverts commit 72a4063651.

Caused #4264
2025-01-10 12:52:11 +01:00
Ryan Houdek b2d579a268 OpcodeDispatcher: Assert on invalid size to LoadRegCachePair
Coverity scan
2025-01-03 11:06:30 -08:00
Ryan Houdek eb1050092f OpcodeDispatcher: Assert on invalid size to SelectPairAddressMode
Coverity scan
2025-01-03 11:05:38 -08:00
Billy Laws efd6e95059 OpcodeDispatcher: Match x86 overflow behaviour for F2I conversions
ARM behaviour here is to saturate on overflow or NaN inputs, whereas
X86 returns a sentinel value of 2^(bitsize-1), explicitly emulate this.
2024-12-30 00:42:55 +00:00
Billy Laws 9bdb1f4306 OpcodeDispatcher: Make narrowing implicit for F64->I32 conversions
This is always used, removing it avoids needing to handle unused codepaths.
2024-12-30 00:36:17 +00:00
Billy Laws ae4b7135d5 OpcodeDispatcher: Share AVX F2I/I2F code for 256-bit SVE 2024-12-30 00:29:31 +00:00
Billy Laws 981c3009ee FEXCore: Emulate EFLAGS.TF
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
2024-12-10 15:20:47 +00:00
Paulo Matos 72a4063651 Cache predicate register generation from pattern 2024-12-06 10:15:38 +01:00
Ryan Houdek ce9a860335 OpcodeDispatcher: Also remove unused CacheIndexToSize 2024-12-01 05:38:11 -08:00
Ryan Houdek 0123946ed1 FEXCore: Removes GetGPRSize and convert all uses to GetGPROpSize
Only a few remaining uses left, easy enough to convert. This finally
switches the final few uses over.

NFC
2024-12-01 05:38:09 -08:00
Alyssa Rosenzweig d19473160d OpcodeDispatcher: drop InvalidateDeferredFlags
it is now useless.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 09:02:01 -05:00
Alyssa Rosenzweig aeb2c98cbf OpcodeDispatcher: drop PossiblySetNZCV
this is a pain to track and, it turns out, buys us virtually nothing on flagm
systems. rip it out.

this fixes a bug with failing to set in all the right places.

on non-flagm systems there's a slight instcountci impact, but that is mostly
mitigated by the earlier patches in the series. so overall a wash there but
worth it for making the codebase easier
to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Alyssa Rosenzweig 2350ae5a07 OpcodeDispatcher: optimize BTC on !flagm
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Alyssa Rosenzweig f4ce6fb621 OpcodeDispatcher: optimize AAS/AAD flag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
LC 704841f004 Merge pull request #4144 from Sonicadvance1/iropsize_addrsize
OpcodeDispatcher: Convert address size helpers to use OpSize
2024-10-28 22:12:40 -04:00
Ryan Houdek f74f276d64 OpcodeDispatcher: Convert address size helpers to use OpSize
NFC
2024-10-28 19:02:34 -07:00
Ryan Houdek 4b10cbdafd OpcodeDispatcher: Convert flags helpers over to OpSize
NFC

Plus the tertiary bits that require changing to support it.
2024-10-28 18:56:42 -07:00
Ryan Houdek 5ed82fa0f6 OpcodeDispatcher: Convert {Load,Store}GPRRegister to OpSize
Trivial but quite a few places pass in a raw integer

NFC
2024-10-28 16:35:23 -07:00
Ryan Houdek 2dd0a82059 IR: Change F80CVTTo to use IR::OpSize 2024-10-28 02:23:47 -07:00
Ryan Houdek 84767c8b20 IR: Change PushStack to use IR::OpSize 2024-10-28 02:05:35 -07:00
Ryan Houdek eccfb53bd5 IR: Change StoreStackMemory to use IR::OpSize 2024-10-28 02:01:47 -07:00
Ryan Houdek 44f9df062e IR: Change Vector_FToI to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 38c58706da IR: Change Vector_FToF to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 03bc962564 IR: Change Vector_FToS to use IR::OpSize 2024-10-28 01:28:56 -07:00
Ryan Houdek 34050431ff IR: Change VDupFromGPR to use IR::OpSize 2024-10-28 01:20:06 -07:00
Ryan Houdek 02ebe06496 IR: Change VInsElement to use IR::OpSize 2024-10-28 01:05:34 -07:00
Ryan Houdek 34d5e70e6b IR: Change VUShlSWide to use IR::OpSize 2024-10-28 00:59:33 -07:00
Ryan Houdek 2d8bd7b59d IR: Change VSShrSWide to use IR::OpSize 2024-10-28 00:57:02 -07:00
Ryan Houdek 3a6f5e638b IR: Change VUShrSWide to use IR::OpSize 2024-10-28 00:54:45 -07:00
Ryan Houdek 7ec21b7121 IR: Change VFAddP to use IR::OpSize 2024-10-28 00:41:18 -07:00
Ryan Houdek 33b3814642 IR: Change VTrn2 to use IR::OpSize 2024-10-28 00:37:44 -07:00
Ryan Houdek 85431a8132 IR: Change VTrn to use IR::OpSize 2024-10-28 00:35:12 -07:00
Ryan Houdek b2ae829731 IR: Change VUnZip to use IR::OpSize 2024-10-28 00:31:23 -07:00
Ryan Houdek 9c2292289b IR: Change VZip to use IR::OpSize 2024-10-28 00:26:34 -07:00
Ryan Houdek ac55e468a7 IR: Change VAddP to use IR::OpSize 2024-10-28 00:20:07 -07:00
Ryan Houdek 0dfd5dd96f IR: Change VXor to use IR::OpSize 2024-10-28 00:15:33 -07:00
Ryan Houdek fc04b9113e IR: Change VAnd to use IR::OpSize 2024-10-28 00:11:44 -07:00
Ryan Houdek 68c038085a IR: Change VSub to use IR::OpSize 2024-10-28 00:10:21 -07:00
Ryan Houdek 176f5a2860 IR: Change VAdd to use IR::OpSize 2024-10-28 00:07:30 -07:00
Ryan Houdek 0ba501636e IR: Change VSRSHR to use IR::OpSize 2024-10-27 23:33:49 -07:00
Ryan Houdek ceca9fff17 IR: Change VSQXTUNPair to use IR::OpSize 2024-10-27 23:30:02 -07:00