Commit Graph
717 Commits
Author SHA1 Message Date
Ryan Houdek b0b41d00ee Various: More static analysis warnings cleanup
NFC
2025-03-12 17:27:41 -07:00
Ryan Houdek 2ae02ded74 AVX128: Remove old comment from VEXTRACT{F,I}128
This comment isn't relevant anymore as unused half of ymm loads will get
DCE'd
2025-03-07 12:00:35 -08:00
Ryan Houdek 79ab76b42e AVX128: Optimize vpalignr
When the shift size is exactly 16bytes, then it turns in to a move.

If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
2025-03-06 17:10:22 -08:00
Ryan Houdek 93ed346fd2 JIT: Optimize SVE offset VL loadstores
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.

Fixes #3791.
2025-03-05 12:01:50 -08:00
Ryan Houdek 2bf87ff40f Various: Removes warnings about uninitialized variables
NFC. These wouldn't even occur in practice.
2025-03-04 19:59:26 -08:00
Ryan Houdek 81434cd233 Merge pull request #4377 from bylaws/sidt
Implement SIDT/LSL
2025-03-04 17:10:32 -08:00
Ryan Houdek b7f58e68c5 Merge pull request #4363 from pmatos/ReciprocalsFix
Improve reciprocal estimate and tests
2025-03-01 00:58:45 -08:00
Billy Laws 6ba2accbdc Frontend: Mark INVLPG as permission-restricted 2025-02-28 16:22:47 +00:00
Billy Laws cdae654fe4 FEXCore: Somewhat implement LSL
Emulate by always returning failure, this deviates from both Linux
and Windows but shouldn't be depended on by anything.
2025-02-27 23:45:14 +00:00
Billy Laws a506c84bc7 FEXCore: Implement SIDT 2025-02-27 23:45:14 +00:00
Paulo Matos b5241e0f60 Improve reciprocal estimate and tests
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.

* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.

Fixes #4319.
2025-02-27 15:07:14 +01:00
Ryan Houdek e718fc35f8 OpcodeDispatcher: Reuse PSHUFD shuffle mask for sha data shuffling
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.

OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).

Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
2025-02-24 12:07:48 -08:00
Ryan Houdek de6931b1f5 OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
```diff
     "sha1rnds4 xmm0, xmm1, 10b": {
-      "ExpectedInstructionCount": 55,
+      "ExpectedInstructionCount": 10,
```

So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
2025-02-22 04:57:58 -08:00
LC beef9eee0a Merge pull request #4368 from Sonicadvance1/more_pshufd
OpcodeDispatcher: Implements a few more pshufd masks
2025-02-19 16:50:23 -05:00
Ryan Houdek 50b5971ee5 OpcodeDispatcher: Implements a few more pshufd masks
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.

We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
2025-02-19 12:46:54 -08:00
Ryan Houdek d10853b775 OpcodeDispatcher: Implement support for SHA1MSG2 using SHA instructions
Only saves a handful of instructions, but still an improvement.

```
   "sha1msg2 xmm0, xmm1": {
     -      "ExpectedInstructionCount": 11,
     +      "ExpectedInstructionCount": 7,
```
2025-02-19 11:33:52 -08:00
Ryan Houdek afa8b3a5c9 OpcodeDispatcher: Implement SHA256MSG2 using new SHA256 operation 2025-02-18 18:03:31 -08:00
Paulo Matos 5666a352d4 NFC: Code cleanup
Removing unused declarations.
Cleaning up unused headers and empty lines.
Avoiding static analysis warnings on `const auto` defaulting to int.
2025-01-22 10:22:25 +01:00
Tony Wasserka da58e6a597 Fix warnings about unused variables 2025-01-21 12:07:33 +01:00
Tony Wasserka 26685143be Update code formatting for logging macros 2025-01-21 12:01:33 +01:00
Tony Wasserka e54b9237c6 Drop use of assume-asserting logging macros 2025-01-21 12:01:33 +01:00
Paulo Matos 2d53867668 x87 fst/fld optimization for different addrmodes
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.

Discussed in #4252.
2025-01-14 16:20:33 +01:00
Ryan Houdek b3794f5541 OpcodeDispatcher: FEX_UNREACHABLE in programming error case
Coverity scan
2025-01-03 11:05:38 -08:00
Ryan Houdek 12dc16780f OpcodeDispatcher: Fixes FEX's H0F3A table handling of REX.W
Most of this table ignores REX.W, but two encodings change behaviour
based on REX.W. These two encodings are PEXTRD/PEXTRQ and PINSRD/PINSRQ.

For every other instruction encoding, they will ignore REX.W, but FEX
was requiring that they didn't have REX.W encoding. I had special cased
this in the past by adding PALIGNR, but that didn't handle any of the
other instructions.

We can't just handle REX.W in the OpcodeDispatcher and remove the two
special cased instructions because these vector operations also interact
with instruction prefix 0x66 which changes the operating size to 16bit
with regular instructions.

So instead just generate all listings of instructions with REX.W being
zero and one and install handlers in all cases.
2025-01-01 08:22:19 -08:00
Billy Laws efd6e95059 OpcodeDispatcher: Match x86 overflow behaviour for F2I conversions
ARM behaviour here is to saturate on overflow or NaN inputs, whereas
X86 returns a sentinel value of 2^(bitsize-1), explicitly emulate this.
2024-12-30 00:42:55 +00:00
Billy Laws 9bdb1f4306 OpcodeDispatcher: Make narrowing implicit for F64->I32 conversions
This is always used, removing it avoids needing to handle unused codepaths.
2024-12-30 00:36:17 +00:00
Billy Laws ae4b7135d5 OpcodeDispatcher: Share AVX F2I/I2F code for 256-bit SVE 2024-12-30 00:29:31 +00:00
Alyssa Rosenzweig 77415538f7 OpcodeDispatcher: use 64-bit XOR for AF calc
we don't need masking and the masking gets in the way of constprop.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-12-13 10:44:56 -05:00
Billy Laws 981c3009ee FEXCore: Emulate EFLAGS.TF
When set - either via POPF or a thread context operation - the trap flag
raises a single step exception after the execution of each instruction.
As e.g. a JUMP instruction with TF set will raise an exception at the
jump target. Handle this on the FEX side by storing both the flag itself
(in bit 0) and a 'block exceptions' flag (in bit 1, inverted). Each
generated block when TF is set is then forced to a single instruction
with logic to raise the exception at the start. Initially after setting
TF exceptions are blocked, then at the start of the block they are
unblocked so that after the instruction executes an exception is raised
at the start of the next block.
2024-12-10 15:20:47 +00:00
Ryan Houdek 0123946ed1 FEXCore: Removes GetGPRSize and convert all uses to GetGPROpSize
Only a few remaining uses left, easy enough to convert. This finally
switches the final few uses over.

NFC
2024-12-01 05:38:09 -08:00
LC 2e7fc60dbf Merge pull request #4169 from Sonicadvance1/x87_loadstore_tests
unittests/ASM: Fixes x87 80-bit loads on the edge of page boundaries.
2024-11-30 10:43:10 -05:00
Alyssa Rosenzweig d19473160d OpcodeDispatcher: drop InvalidateDeferredFlags
it is now useless.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 09:02:01 -05:00
Alyssa Rosenzweig aeb2c98cbf OpcodeDispatcher: drop PossiblySetNZCV
this is a pain to track and, it turns out, buys us virtually nothing on flagm
systems. rip it out.

this fixes a bug with failing to set in all the right places.

on non-flagm systems there's a slight instcountci impact, but that is mostly
mitigated by the earlier patches in the series. so overall a wash there but
worth it for making the codebase easier
to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-11-26 08:59:19 -05:00
Ryan Houdek 1465df874b OpcodeDispatcher: Fixes 80-bit loads
Ensures reads don't go past the end of the page boundary.
SVE masked loads can make this more effective but `VLoadVectorMasked`
isn't setup to be efficient for this case yet.
2024-11-24 00:36:50 -08:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
Ryan Houdek 82f936cb6d IR: Converts base IR operations to store OpSize sizes
NFC

Finally converts the IR operations themselves to store the OpSize for
the IR operation size and element sizes.

This also finally, FINALLY, converts that remaining `_Constant` helper
to stop using a size field that is specified in bits rather than bytes
like all the other IR op handlers. That thing was so confusing and now
it's gone.
2024-10-28 21:26:59 -07:00
Ryan Houdek 4b10cbdafd OpcodeDispatcher: Convert flags helpers over to OpSize
NFC

Plus the tertiary bits that require changing to support it.
2024-10-28 18:56:42 -07:00
Ryan Houdek 2dd0a82059 IR: Change F80CVTTo to use IR::OpSize 2024-10-28 02:23:47 -07:00
Ryan Houdek 84767c8b20 IR: Change PushStack to use IR::OpSize 2024-10-28 02:05:35 -07:00
Ryan Houdek eccfb53bd5 IR: Change StoreStackMemory to use IR::OpSize 2024-10-28 02:01:47 -07:00
Ryan Houdek f4e930262f IR: Change PCLMUL to use IR::OpSize 2024-10-28 01:50:24 -07:00
Ryan Houdek 44f9df062e IR: Change Vector_FToI to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 38c58706da IR: Change Vector_FToF to use IR::OpSize 2024-10-28 01:50:23 -07:00
Ryan Houdek 03bc962564 IR: Change Vector_FToS to use IR::OpSize 2024-10-28 01:28:56 -07:00
Ryan Houdek 4544e7c1af IR: Change Float_FromGPR_S to use IR::OpSize 2024-10-28 01:21:48 -07:00
Ryan Houdek 34050431ff IR: Change VDupFromGPR to use IR::OpSize 2024-10-28 01:20:06 -07:00
Ryan Houdek 65439956bf IR: Change VCastFromGPR to use IR::OpSize 2024-10-28 01:18:21 -07:00
Ryan Houdek bd5159c7d5 IR: Change VFMLA to use IR::OpSize 2024-10-28 01:15:01 -07:00
Ryan Houdek 0fa095cec5 IR: Change VFCMPEQ to use IR::OpSize 2024-10-28 01:10:07 -07:00
Ryan Houdek 9e3c50ca2c IR: Change VExtr to use IR::OpSize 2024-10-28 01:08:35 -07:00