Ryan Houdek
efd5c51110
IR: Change LoadNamedVectorIndexedConstant to use IR::OpSize
2024-10-27 22:08:27 -07:00
Ryan Houdek
886db4ffca
IR: Change LoadNamedVectorConstant to use IR::OpSize
2024-10-27 22:05:42 -07:00
Ryan Houdek
e6f6ee2bcd
IR: Change VectorImm to use IR::OpSize
2024-10-27 21:58:34 -07:00
Ryan Houdek
5566b4455b
IR: Change VFCMPScalarInsert to use IR::OpSize
2024-10-27 18:35:03 -07:00
Ryan Houdek
fed2c13521
IR: Change VFToIScalarInsert to use IR::OpSize
2024-10-27 18:32:11 -07:00
Ryan Houdek
5626f4e50a
IR: Change VSToFVectorInsert to use IR::OpSize
2024-10-27 18:27:39 -07:00
Ryan Houdek
efbc42dac3
IR: Change VFToFScalarInsert to use IR::OpSize
2024-10-27 18:25:13 -07:00
Ryan Houdek
d6f726fc23
IR: Change VFSqrtScalarInsert to use IR::OpSize
2024-10-27 18:20:06 -07:00
Ryan Houdek
081907e168
IR: Change VFAddScalarInsert to use IR::OpSize
2024-10-27 18:14:37 -07:00
Ryan Houdek
1a115a8ce6
IR: Change FCmp to use IR::OpSize
2024-10-27 18:06:34 -07:00
Ryan Houdek
1bde30a196
IR: Change Float_ToGPR_S to use IR::OpSize
2024-10-27 18:02:45 -07:00
Ryan Houdek
764aacaa8f
IR: Change VExtractToGPR to use IR::OpSize
2024-10-27 17:58:53 -07:00
Ryan Houdek
f1a42869d5
IR: Change CondJump to use IR::OpSize
2024-10-27 17:50:51 -07:00
Ryan Houdek
b31ce13f68
IR: Change Pop to use IR::OpSize
2024-10-27 17:41:30 -07:00
Ryan Houdek
260d3b0b4e
IR: Change Push to use IR::OpSize
2024-10-27 17:39:21 -07:00
Ryan Houdek
c8c7ffbf05
IR: Change VBroadcastFromMem to use IR::OpSize
2024-10-27 17:35:56 -07:00
Ryan Houdek
4b03185b77
IR: Change VStoreVectorElement to use IR::OpSize
2024-10-27 17:32:25 -07:00
Ryan Houdek
52ec572db3
IR: Change VLoadVectorElement to use IR::OpSize
2024-10-27 17:27:42 -07:00
Ryan Houdek
3f6cdc2e03
IR: Change VLoadVectorMasked to use IR::OpSize
2024-10-27 17:20:18 -07:00
Ryan Houdek
014917301a
IR: Change StoreMem to use IR::OpSize
2024-10-27 17:14:18 -07:00
Ryan Houdek
07f8a4eadd
IR: Change LoadMem to use IR::OpSize
2024-10-27 16:33:14 -07:00
Ryan Houdek
2f9b0de742
IR: Change StoreContext to use IR::OpSize
2024-10-27 15:46:32 -07:00
Ryan Houdek
40fd4bbb66
IR: Change LoadContext to use IR::OpSize
2024-10-27 15:42:07 -07:00
Ryan Houdek
c045e14837
IR: Change StoreRegister to use IR::OpSize
2024-10-27 15:37:05 -07:00
Ryan Houdek
8f4113d859
IR: Change LoadRegister to use IR::OpSize
2024-10-27 15:34:46 -07:00
Ryan Houdek
4cfc2ac1a4
IR: Change AllocateFPR to use IR::OpSize
2024-10-27 15:30:11 -07:00
Paulo Matos
5f6c0d2245
X87 code simplification
...
Merges some of the code from reduced precision into the main path
since they are practically the same.
2024-10-24 18:15:49 +02:00
Ryan Houdek
caaacb6c15
Merge pull request #4127 from alyssarosenzweig/opt/masking
...
Optimize bsf, bsr, register cmpxchg, pcmpistri
2024-10-23 07:28:29 -07:00
Alyssa Rosenzweig
9c605e7333
OpcodeDispatcher: optimize bsf/bsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-10-22 13:31:49 -04:00
Ryan Houdek
fe8f5c745d
OpcodeDispatcher/AVX128: Use unified PSHUF{L,H}W implementation
...
Allows the AVX128 implementation to use the same implementation as the
128-bit SSE implementation, because it works per 128-bit lane just like
SSE. Is a minor optimization.
2024-10-21 09:24:38 -07:00
Ryan Houdek
d876224358
OpcodeDispatcher: Unify MMX and SSE PSHUF{L,H}W implementations
...
No functional change.
2024-10-21 09:11:05 -07:00
Paulo Matos
10ec6b63b6
Fix FXTRACT for 0.0 and -0.0
...
Fixes fxtract by returning the correct values for 0.0 and -0.0. We moved the split of fxtract into _sig and _exp, to the opcode dispatcher, to ease some comparisons.
Also removed the IR node F80XTRACTStack which is not needed anymore.
2024-10-17 09:05:10 +02:00
Paulo Matos
0d53f2b45c
Implements explicit state switch between X87 and MMX
...
Fixes #3850
2024-10-15 17:58:53 +02:00
Ryan Houdek
c984bdb42a
OpcodeDispatcher: Remove previous template instantiantions and use Bind
2024-09-16 18:52:54 -07:00
Ryan Houdek
37a70e2ec6
FEXCore: Convert Base tables over to constexpr
...
Only doing the single table for review purposes. Once reviewed I will
hammer out the remaining tables.
Similar to #3320 , most of the OpcodeDispatcher tables can be consteval
and made to be a compile time constant. This just requires shuffling the
code slightly. The idea is to get almost all of the table setup out of
the `InstallOpcodeHandlers` function and instead only install the
handlers that change based on 32-bit or 64-bit, just like the x86 tables
we also did.
2024-09-13 11:39:09 -07:00
Alyssa Rosenzweig
d17f33a922
OpcodeDispatcher: don't emit fake 0 for condjump
...
not needed and getting in the way
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig
50a3ca0d6d
IR: introduce dedicated PF/AF instructions
...
this makes reasoning about them a little easier, e.g. for flags. about 1% win
in nodejs.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig
e9ab514962
IR: push parity evaluation down
...
so we can optimize it globally
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig
4eb0948451
IR: push down AXFLAG lowering
...
so we can get the new axflag optimizations on billy's x13s.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig
42f2851575
Merge pull request #4009 from alyssarosenzweig/opt/axflag
...
Optimize AXFLAG-less systems
2024-08-27 08:01:05 -04:00
Alyssa Rosenzweig
8d6b454455
OpcodeDispatcher: optimize AXFLAG emulation
...
this should help on x13s which has flagm but not flagm2.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-25 18:38:03 -04:00
Alyssa Rosenzweig
205ec3e14d
OpcodeDispatcher: refactor axflag
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-25 18:35:27 -04:00
Ryan Houdek
8bf4a124c8
AVX128: Fixes 256-bit float compares
...
Just like the bug in #4006 , we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.
PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek
1d00ad6030
OpcodeDispatcher: Convert VectorVariableBlend to Bind handler
2024-08-22 15:09:44 -07:00
Ryan Houdek
23a076c313
OpcodeDispatcher: Convert packed vector shifts to Bind handler
2024-08-22 15:07:09 -07:00
Ryan Houdek
57eacab654
OpcodeDispatcher: Convert VBROADCASTOp to Bind handler
2024-08-22 14:59:25 -07:00
Ryan Houdek
b4093a8888
OpcodeDispatcher: Convert packed HSub to Bind handler
2024-08-22 14:57:40 -07:00
Ryan Houdek
7c5a9b5d6a
OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler
2024-08-22 14:55:34 -07:00
Ryan Houdek
12b3c82d83
OpcodeDispatcher: Convert PExtr to Bind handler
2024-08-22 14:54:26 -07:00
Ryan Houdek
66520bce0a
OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler
2024-08-22 14:52:41 -07:00