Alyssa Rosenzweig
068599b1ec
OpcodeDispatcher: use size-appropriate alu in bt*
...
saves zero-extending move for 32-bit ops.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
7216415bfc
OpcodeDispatcher: reorder flag calcs in bt*
...
saves big moves.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
3020626506
OpcodeDispatcher: unify bt/btc/bts/btr impls
...
they're all copypastes of each other, unify into one general "bit test & perform
action" template. this means most of the wins from the previous commits now
apply for bt* without more copypaste.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 14:22:21 -04:00
Alyssa Rosenzweig
0a79fa8d5d
OpcodeDispatcher: remove masking for 32/64-bit bt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:40:18 -04:00
Alyssa Rosenzweig
f8380b9adb
OpcodeDispatcher: use smaller shifts for BT
...
if the shift is < N, and we grab bit 0 after, we only need to consider <=N
bits of the source. this lets us use 32-bit lsr for 32-bit bt, which will
reduce masking in the next commit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:40:18 -04:00
Alyssa Rosenzweig
d898028bc3
OpcodeDispatcher: optimize BT with constant
...
use the rmif properly.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:35:23 -04:00
Alyssa Rosenzweig
14ba64a22d
OpcodeDispatcher: use rmif masking for bt
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:20:22 -04:00
Alyssa Rosenzweig
6716077cb6
OpcodeDispatcher: don't mask bt source
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-29 13:16:55 -04:00
Ryan Houdek
a7caf83022
FEXCore: Fixes imul returning garbage data
...
When a 32-bit imul was being executed it had a chance of returning
garbage data in the upper 32-bits of the 64-bit result.
While this didn't typically cause problems, this gets exacerbated from
32-bit applications executing multiplies for address calculations.
A combination of commits 7146691360 and
d01b457727 exposed this problem where
previously there would be multiple moves between the calculation and
data use which would have zero'd the upper bits for us previously.
Now that we are no longer doing that, we need to make sure the opcode
dispatcher doesn't generate broken code instead.
Fixes Dungeon Defenders, which hasn't worked since FEX-2308.
Adds an ASM test that ensures we don't break it again.
2023-11-24 14:04:05 -08:00
Alyssa Rosenzweig
f60608a9c0
OpcodeDispatcher: allow garbage for 32bit inc/dec
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
153d871be2
OpcodeDispatcher: allow garbage for 32-bit cmp
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
149f3e6f6d
OpcodeDispatcher: rm flagsOp unused since select rework
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
56841f0e50
OpcodeDispatcher: avoid moves with 64bit imul
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-14 08:40:41 -04:00
Alyssa Rosenzweig
723146050b
OpcodeDispatcher: allow garbage with multiplies
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-14 08:40:41 -04:00
Alyssa Rosenzweig
89b00c89aa
OpcodeDispatcher: optimize rcl 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
d38917b5f0
OpcodeDispatcher: rm pointless constant
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
910e0242c1
OpcodeDispatcher: optimize rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
651b7bb75d
OpcodeDispatcher: optimize RCL
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
c9f13ae1dd
OpcodeDispatcher: optimize RCR the usual ways
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
862e575100
OpcodeDispatcher: avoid some ubfx for rcr with flagm
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
0f25a960ee
OpcodeDispatcher: remove bfe for small shl imm
...
We allow the garbage in flags calculation, it's ignored.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
282ed3e309
OpcodeDispatcher: optimize bsf/bsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
238e52f74a
OpcodeDispatcher: Don't mask 32-bit bzhi either
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:36:46 -04:00
Alyssa Rosenzweig
b2a9785959
OpcodeDispatcher: optimize bzhi
...
Trickery to save an instruction :')
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
224a1f19a3
OpcodeDispatcher: improve bzhi flag gen
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
472d143021
OpcodeDispatcher: Use TestNZ directly for SelectBit
...
Lets us drop the bitwise select variants, this was the last use.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
5471367db1
OpcodeDispatcher: Optimize branches
...
For native cases. Big perf win.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
b8265b1067
OpcodeDispatcher: Rework SelectCC for branches
...
Separate out the NZCV bits from the more complex stuff so we can specially
optimize the branches.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
ec14a65e23
OpcodeDispatcher: optimize LOOP invert
...
need to reorder since the select clobbers nzcv. before/after diff on the loopne
unit test:
> 4308: [INFO] cset w20, ne
40c41
< 4308: [INFO] mrs x20, nzcv
---
> 4308: [INFO] mrs x21, nzcv
42,49c43,48
< 4308: [INFO] cset x21, ne
< 4308: [INFO] ubfx w22, w20, #30 , #1
< 4308: [INFO] eor x22, x22, #0x1
< 4308: [INFO] and x21, x21, x22
< 4308: [INFO] msr nzcv, x20
< 4308: [INFO] cbnz x21, #+0x8 (addr 0xfffed66f8094)
< 4308: [INFO] b #+0x1c (addr 0xfffed66f80ac)
< 4308: [INFO] ldr x0, pc+8 (addr 0xfffed66f809c)
---
> 4308: [INFO] cset x22, ne
> 4308: [INFO] and x20, x22, x20
> 4308: [INFO] msr nzcv, x21
> 4308: [INFO] cbnz x20, #+0x8 (addr 0xfffec94e8090)
> 4308: [INFO] b #+0x1c (addr 0xfffec94e80a8)
> 4308: [INFO] ldr x0, pc+8 (addr 0xfffec94e8098)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig
61bdf64e15
OpcodeDispatcher: Cleanup PF select
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig
482b35c283
OpcodeDispatcher: Optimize JA/JNA selects
...
Chain two csel instructions together, which should be optimal.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 10:46:59 -04:00
Alyssa Rosenzweig
b74d886017
OpcodeDispatcher: Cleanup selectcc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 10:46:59 -04:00
Alyssa Rosenzweig
1667abad7e
OpcodeDispatcher: Don't fold flags in SelectCC
...
Not beneficial in the new approach.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 10:46:59 -04:00
Alyssa Rosenzweig
b187a853e7
OpcodeDispatcher: Use NZCVSelect for SelectCC
...
Massively better codegen.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 10:46:58 -04:00
Alyssa Rosenzweig
3767f3633d
Merge pull request #3263 from alyssarosenzweig/opt/not-garbage
...
OpcodeDispatcher: Make "not" not garbage
2023-11-09 19:58:02 -04:00
Alyssa Rosenzweig
da3e3fc7a3
OpcodeDispatcher: Optimize some selects
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-09 15:21:10 -04:00
Alyssa Rosenzweig
1ce3c16b30
OpcodeDispatcher: Make "not" not garbage
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-09 12:02:20 -04:00
Alyssa Rosenzweig
bdaa70405f
OpcodeDispatcher: simplify flag control ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig
87cac09477
OpcodeDispatcher: optimize cmc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-09 09:40:51 -04:00
Alyssa Rosenzweig
73958b9163
OpcodeDispatcher: Use DeriveOp
...
Replace every instance of the Op overwrite pattern, and ban that anti-pattern
from the codebase in the future. This will prevent piles of NZCV related
regressions.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-07 12:05:00 -04:00
Ryan Houdek
5103f2d92b
Merge pull request #3247 from alyssarosenzweig/refactor/nzcv-prereq
...
Preparatory patches for nzcv
2023-11-01 14:07:14 -07:00
Alyssa Rosenzweig
5522c6db9c
OpcodeDispatcher: Use jump wrappers
...
Mostly automated replacement + renaming for build fixing.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-01 15:44:36 -04:00
Ryan Houdek
bab96b9441
OpcodeDispatcher: Allow garbage in upper bits for more ALU ops
...
Secondary ALU operations were missed and when the operation is 4-bytes
in size we can also allow garbage upper bits since the JIT will emit a
32-bit operation for this instruction which is safe.
Optimizes some bad codegen around 32-bit ALU operations.
2023-10-25 20:44:36 -07:00
Ryan Houdek
95c756b466
OpcodeDispatcher: Optimize DF pointer offset calculation
...
Previously this moved two constant, did a compare and a csel. Four
instructions in total. It also corrupts NZCV which we want to use for
other things.
This new codegen emits one constant and one subtract instruction, two
instructions total and doesn't touch NZCV.
More optimal!
2023-10-23 09:27:41 -07:00
Alyssa Rosenzweig
e455996dbd
OpcodeDispatcher: Remove silly shift branching
...
The flag generation code does this internally and more efficiently.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-10-23 10:16:38 -04:00
Lioncache
4b356a7c2c
OpcodeDispatcher: Have MOVNTSD go down the non-temporal path
...
For some reason this was using the regular unaligned path.
2023-10-18 14:59:02 +02:00
Lioncache
2b67f87054
OpcodeDispatcher: Handle SSE vector moves into themselves a little better
...
Obviously, it's silly to do this, but we should still be generating
optimal code for this case (which is none at all).
2023-10-18 14:58:57 +02:00
Lioncache
47a0f14537
OpcodeDispatcher: Remove unnecessary 128-bit truncating moves from StoreResult
...
Removes the truncating move that we perform inside the StoreResult
function and instead delegates the responsibility to the instruction
implementations themselves.
This removes a lot of redundant moves that occur on 128-bit variants
of AVX instructions.
Also fixes a weird case where we were handling 128-bit SVE
in VBroadcastFromMem when we already have AdvSIMD instructions
that will perfom the zero-extension behavior for us.
2023-10-17 11:07:04 +02:00
Lioncache
2304cfc530
OpcodeDispatcher: Remove prefixing from MemoryAccessType enum
...
Since this is an enum class, we don't need to add a prefix.
2023-10-16 03:10:33 +02:00
Lioncache
1a39de4509
OpcodeDispatcher: Put extra LoadSource options in a struct
...
Allows for easier expansion without needing to expand the function definitons.
Also makes a few usages significantly less verbose and makes specifying
options a little more declarative, rather than having to memorize what
each argument is specifying.
2023-10-15 21:18:00 +02:00