Billy Laws
0f6ceaac05
OpcodeDispatcher: Load only the element size from memory for VFMAScalarImpl
2025-07-09 00:52:59 +01:00
Billy Laws
635816c07f
OpcodeDispatcher: Force ElementSize loads for UCOMISxOp
2025-07-09 00:52:59 +01:00
Billy Laws
2fac5c23ff
OpcodeDispatcher: Always use 32-bit load/store for {LD,ST}MXCSR
2025-07-09 00:52:59 +01:00
Billy Laws
c2e2d1e92f
OpcodeDispatcher: Only read at most ElementSize in CVTFPR_To_GPR
2025-07-09 00:50:47 +01:00
Alyssa Rosenzweig
77bb01d384
OpcodeDispatcher: don't generate pointless Xor for AF
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig
589906e6e6
OpcodeDispatcher: generalize AF on constants optimization
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig
f0fcf6d9e6
OpcodeDispatcher: optimize SVE vmovmskpd
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig
1591ced5a7
OpcodeDispatcher: use ArithRef for ADC/SBB flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig
ad7844d7da
OpcodeDispatcher: do not generate useless VMov
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-06-02 14:15:51 -04:00
Ryan Houdek
cdea8d7f74
Merge pull request #4593 from alyssarosenzweig/ici/zeroing-sub-regs
...
InstructionCountCI: add more cases for mov 0/~0
2025-05-29 12:26:15 -07:00
Alyssa Rosenzweig
ece817c691
OpcodeDispatcher: optimize logical flags
...
seems to be strictly better.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-05-28 13:34:50 -04:00
Alyssa Rosenzweig
3fbc8204b7
OpcodeDispatcher: clean up logical flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-05-28 13:28:31 -04:00
Alyssa Rosenzweig
b8dd5d95b0
OpcodeDispatcher: optimize X87FTWTag
...
using bit twiddling tricks :-)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-05-28 13:13:16 -04:00
Billy Laws
416267a238
OpcodeDispatcher: Safely clobber NZCV in FCOMIF64
...
Also fix a small typo that broke the !flagm2 path.
2025-04-16 13:06:37 +01:00
Ryan Houdek
ec1c7797f4
OpcodeDispatcher: Implement support for sha256rnds2 using ARM instructions
...
The big one.
2025-04-02 12:51:03 -07:00
Ryan Houdek
1b18bfaff5
Merge pull request #4473 from bylaws/win32-f
...
Windows: Small fixups
2025-04-01 10:53:28 -07:00
Billy Laws
02d3a319f9
OpTables: Disable thunk opcodes on win32
...
They are of no use here, and are quite frequent in never-taken blocks in Denuvo games
so treating them as invalid avoids wasting some time.
2025-03-31 23:57:58 +01:00
Alyssa Rosenzweig
c4f7b27459
OpcodeDispatcher: accelerate F->I conversions with FRINTTS
...
this should significantly help perf on supported platforms.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2025-03-31 16:09:54 -04:00
Ryan Houdek
21233430aa
Merge pull request #4446 from lioncash/cond
...
OpcodeDispatcher: Remove unnecessary 128-bit check in VPGATHER()
2025-03-27 20:30:49 -07:00
Lioncache
40b1c32008
OpcodeDispatcher: Remove unnecessary 128-bit check in VPGATHER()
...
This is already guaranteed to be true, since it's checked in the outer if,
so this can just be a regular else statement.
2025-03-27 23:08:59 -04:00
Ryan Houdek
8f9d818368
OpcodeDispathcer: static analysis warnings
2025-03-27 04:26:40 -07:00
Paulo Matos
b18f148575
Abstract AddressMode into its own header
...
Also refactor usage of utility functions into x87 Stack Optimization Pass.
2025-03-20 11:41:26 +01:00
Paulo Matos
e18a661b50
Refactor OP_STORESTACKMEM case in x87 Stack Opt Pass
2025-03-18 18:34:06 +01:00
Paulo Matos
16e5777816
Do not attempt direct encoding of 32bit addressing modes
...
Fixes #4393
2025-03-18 15:54:54 +01:00
Paulo Matos
64020e8828
Simplify by merging FSTF64 with FST
2025-03-18 15:54:54 +01:00
Ryan Houdek
929b111648
Merge pull request #4405 from Sonicadvance1/static_analysis_again
...
Various: More static analysis warnings cleanup
2025-03-17 13:08:11 -07:00
Ryan Houdek
b901b42417
OpcodeDispatcher/X87: Optimize GPR moves using new IR operation
...
This saves one instruction per FILD and FRSTOR
2025-03-12 22:22:03 -07:00
Ryan Houdek
b0b41d00ee
Various: More static analysis warnings cleanup
...
NFC
2025-03-12 17:27:41 -07:00
Ryan Houdek
2ae02ded74
AVX128: Remove old comment from VEXTRACT{F,I}128
...
This comment isn't relevant anymore as unused half of ymm loads will get
DCE'd
2025-03-07 12:00:35 -08:00
Ryan Houdek
79ab76b42e
AVX128: Optimize vpalignr
...
When the shift size is exactly 16bytes, then it turns in to a move.
If the shift size is above 16-bytes then synthesize the zero register in
the OpcodeDispatcher, so the backend doesn't synthesize and not cache.
2025-03-06 17:10:22 -08:00
Ryan Houdek
93ed346fd2
JIT: Optimize SVE offset VL loadstores
...
This was only wired up for 256-bit SVE and wasn't ever hit for 128-bit
SVE. Ensure it works with 128-bit SVE, so mulvl needs to know when
128-bit is used. Then wire it up for vmaskmovps/pd. This saves one
instruction per operation.
Fixes #3791 .
2025-03-05 12:01:50 -08:00
Ryan Houdek
2bf87ff40f
Various: Removes warnings about uninitialized variables
...
NFC. These wouldn't even occur in practice.
2025-03-04 19:59:26 -08:00
Ryan Houdek
81434cd233
Merge pull request #4377 from bylaws/sidt
...
Implement SIDT/LSL
2025-03-04 17:10:32 -08:00
Ryan Houdek
b7f58e68c5
Merge pull request #4363 from pmatos/ReciprocalsFix
...
Improve reciprocal estimate and tests
2025-03-01 00:58:45 -08:00
Billy Laws
6ba2accbdc
Frontend: Mark INVLPG as permission-restricted
2025-02-28 16:22:47 +00:00
Billy Laws
cdae654fe4
FEXCore: Somewhat implement LSL
...
Emulate by always returning failure, this deviates from both Linux
and Windows but shouldn't be depended on by anything.
2025-02-27 23:45:14 +00:00
Billy Laws
a506c84bc7
FEXCore: Implement SIDT
2025-02-27 23:45:14 +00:00
Paulo Matos
b5241e0f60
Improve reciprocal estimate and tests
...
3DNow Reciprocal estimations did not have enough accuracy. Tests were enabled
to check for accurate values of reciprocals.
* where needed, reciprocal accuracy was increased.
* 3DNow sqrt reciprocal fixed for negative values.
* New helper VFCopySign IR op added.
Fixes #4319 .
2025-02-27 15:07:14 +01:00
Ryan Houdek
e718fc35f8
OpcodeDispatcher: Reuse PSHUFD shuffle mask for sha data shuffling
...
We already have this mask generated, and because sha instructions
typically don't exist in a vacuum it is actually beneficial to cache the
mask and use a single tbl instruction per shuffle.
OpenSSL has 12 sha1 instructions in their hot loop as an example, so
this would be a fairly good reduction in that loop. Sadly we don't have
it in instcountci, instead having their sha256 hotloop instead (Which
currently doesn't have sha256rnds2 optimized).
Even in a vacuum this is technically 1 instruction savings for each
instruction which is nice.
2025-02-24 12:07:48 -08:00
Ryan Houdek
de6931b1f5
OpcodeDispatcher: Emulate SHA1RNDS4 with ARM sha extensions
...
```diff
"sha1rnds4 xmm0, xmm1, 10b": {
- "ExpectedInstructionCount": 55,
+ "ExpectedInstructionCount": 10,
```
So I spent a few hours glaring at this instruction. Then spent a few
more glaring in to the sunset and then found the optimization.
2025-02-22 04:57:58 -08:00
LC
beef9eee0a
Merge pull request #4368 from Sonicadvance1/more_pshufd
...
OpcodeDispatcher: Implements a few more pshufd masks
2025-02-19 16:50:23 -05:00
Ryan Houdek
50b5971ee5
OpcodeDispatcher: Implements a few more pshufd masks
...
Saw these while scanning around. Funnily it makes it look like libnss is
worse off because there are multiple instructions using the same table
lookup to swizzle. So one instruction turns in to two.
We don't have a way to choose one path or the other, so it's usually
better to go the route that the instruction in a vacuum is improved, so
on average it is also improved.
2025-02-19 12:46:54 -08:00
Ryan Houdek
d10853b775
OpcodeDispatcher: Implement support for SHA1MSG2 using SHA instructions
...
Only saves a handful of instructions, but still an improvement.
```
"sha1msg2 xmm0, xmm1": {
- "ExpectedInstructionCount": 11,
+ "ExpectedInstructionCount": 7,
```
2025-02-19 11:33:52 -08:00
Ryan Houdek
afa8b3a5c9
OpcodeDispatcher: Implement SHA256MSG2 using new SHA256 operation
2025-02-18 18:03:31 -08:00
Paulo Matos
5666a352d4
NFC: Code cleanup
...
Removing unused declarations.
Cleaning up unused headers and empty lines.
Avoiding static analysis warnings on `const auto` defaulting to int.
2025-01-22 10:22:25 +01:00
Tony Wasserka
da58e6a597
Fix warnings about unused variables
2025-01-21 12:07:33 +01:00
Tony Wasserka
26685143be
Update code formatting for logging macros
2025-01-21 12:01:33 +01:00
Tony Wasserka
e54b9237c6
Drop use of assume-asserting logging macros
2025-01-21 12:01:33 +01:00
Paulo Matos
2d53867668
x87 fst/fld optimization for different addrmodes
...
Includes tests and instcountci files and tests.
When the x87 optimizations were implement, we missed
optimizing different addressing modes. This commit addresses this issue.
Discussed in #4252 .
2025-01-14 16:20:33 +01:00
Ryan Houdek
b3794f5541
OpcodeDispatcher: FEX_UNREACHABLE in programming error case
...
Coverity scan
2025-01-03 11:05:38 -08:00