Alyssa Rosenzweig
d1e43d94e9
OpcodeDispatcher: use axflag for fcmp faster
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
1b490e0e53
CodeEmitter: add ax/xaflag
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
094146d630
IR: add axflag
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
23c2a53683
OpcodeDispatcher: move fcmp flag fixup to dispatcher
...
simpler *and* much faster
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
82b7689ca4
OpcodeDispatcher: remove fcmp deferral
...
no longer load bearing, delete the abstraction.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
149f3e6f6d
OpcodeDispatcher: rm flagsOp unused since select rework
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Alyssa Rosenzweig
2dcae23776
Arm64Emitter: Dedicate registers for PF/AF
...
Many flag-generating instructions like cmp need to save calculations for
deferred PF and AF flag calculation. Currently, they require a store per flag,
which is prohibitively expensive for hot instructions like cmp. By instead
pinning PF/AF temporary results to registers (x26/x27 by convention here), we
eliminate many stores altogether and turn the rest into zero-cycle moves (on
64-bit at least, this isn't optimal for 32-bit emulation due to CTX->GetGPRSize
shenanigans, need to check if this requirement can be lifted..).
To implement, we model as SRA and then the existing SRA code is able to generate
good code with little manual tuning. (Future work will get us to excellent code
with more tuning ;) ).
The tradeoff is reducing the working dynamic GPR set by 2 registers, which might
increase spilling in some cases. I think it's worth it in practice, though.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-17 17:37:24 -04:00
Ryan Houdek
c1d5fae018
Merge pull request #3273 from alyssarosenzweig/opt/shifts
...
Optimize shifts/rotates
2023-11-14 14:13:56 -08:00
Alyssa Rosenzweig
56841f0e50
OpcodeDispatcher: avoid moves with 64bit imul
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-14 08:40:41 -04:00
Alyssa Rosenzweig
723146050b
OpcodeDispatcher: allow garbage with multiplies
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-14 08:40:41 -04:00
Alyssa Rosenzweig
85b1aa4c2d
OpcodeDispatcher: optimize mul flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-14 08:40:41 -04:00
Ryan Houdek
c69082b1a4
OpcodeDispatcher: Optimize three sha instructions
...
- sha1nexte
- Takes advantage of sha1h if supported
- Does the operation in a vector otherwise
- sha1msg2
- Instead of dumping everything to GPRs, we can do this with vectors
- Mostly matches ARM's sha1su1 instruction, but it is /just/
different enough to be annoying.
- sha256msg1
- Directly matches sha256u0
- Leaves the previous implementation alone
2023-11-13 18:38:02 -08:00
Ryan Houdek
74b2548982
IR: Implement support for VUSHRAI IR op
...
This matches Arm64 usra semantics. This instruction is useful for
implementing vector element rotate.
2023-11-13 18:22:45 -08:00
Ryan Houdek
f31656ec65
IR: Support sha1h and sha256u0
...
These match our needs so wire them up
2023-11-13 18:22:45 -08:00
Ryan Houdek
e91420c405
HostFeatures: fixup SHA checks
...
- Simulator doesn't support SHA
- Use DisableCrypto option to disable sha as well
- Only enable SHA if ARM cpu supports both SHA1 and SHA2
2023-11-13 18:22:45 -08:00
Ryan Houdek
25df59a65d
ArmEmitter: Fixes sha256u1 emitter
...
Noticed this was actually emitting sha256h2
2023-11-13 18:22:45 -08:00
Alyssa Rosenzweig
83fdd5720f
IR: Add ccmn
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 22:05:02 -04:00
Alyssa Rosenzweig
89b00c89aa
OpcodeDispatcher: optimize rcl 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
d38917b5f0
OpcodeDispatcher: rm pointless constant
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
910e0242c1
OpcodeDispatcher: optimize rcr 1-bit
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
651b7bb75d
OpcodeDispatcher: optimize RCL
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
c9f13ae1dd
OpcodeDispatcher: optimize RCR the usual ways
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
862e575100
OpcodeDispatcher: avoid some ubfx for rcr with flagm
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
4669c4541c
OpcodeDispatcher: don't zero for flagm ror
...
missed earlier in the PR, would be annoying to rebase in.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
769a8c41c4
OpcodeDispatcher: Branch over shift=0 flags
...
Not supposed to touch flags at all, so don't! instead of making a terrible mess
of csels. a lot less instructions, and probably faster because the branch should
be predicted correctly in practice in hot loops.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
bec9dba2b1
OpcodeDispatcher: use shifted xor + rmif for rotates
...
eliminates lots of Bfe on flagm.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
205ba2ea13
IR: add shifted xor
...
for rotates.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
2073f6d287
OpcodeDispatcher: don't zero nzcv for flagm shifts
...
Faster for flagm. would be slower for !flagm because bfi slowness...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
0f25a960ee
OpcodeDispatcher: remove bfe for small shl imm
...
We allow the garbage in flags calculation, it's ignored.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
282ed3e309
OpcodeDispatcher: optimize bsf/bsr
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
cd031a7d38
OpcodeDispatcher: avoid some ubfx for flagm
...
Do the masking as part of the rmif, for free.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
e1885ed0bd
OpcodeDispatcher: Use 64-bit ubfx
...
for larger shifts.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-13 21:21:01 -04:00
Alyssa Rosenzweig
238e52f74a
OpcodeDispatcher: Don't mask 32-bit bzhi either
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:36:46 -04:00
Alyssa Rosenzweig
9398b931fb
Arm64Emitter: Handle 32-bit negatives
...
Noticed in the area.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:36:42 -04:00
Alyssa Rosenzweig
b2a9785959
OpcodeDispatcher: optimize bzhi
...
Trickery to save an instruction :')
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
224a1f19a3
OpcodeDispatcher: improve bzhi flag gen
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
e25849b2cb
OpcodeDispatcher: fix BZHI flag calculation
...
needs SF.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
57978accc1
OpcodeDispatcher: Add heuristic to prefer rmif for InsertNZCV
...
Decent instcountci win.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
ff37177f4d
IR: Remove AND(N)Z conditions
...
Unused since TestNZ+NZCVSelect accomplishes the same and good riddance. Might
bring them back later for tbz/tbnz, but certainly not in this Selectful form.
(I added them when I thought we were going to RA the flags. With the more
effective static approach, we don't need this for that.)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
472d143021
OpcodeDispatcher: Use TestNZ directly for SelectBit
...
Lets us drop the bitwise select variants, this was the last use.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
157f95b08f
IR: Extend TestNZ to two sources
...
Unlock the power of AND.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
6bf7ab0778
IR: Remove complex branches
...
Performance footgun and now unused. Don't bring it back.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
0de958be2a
IR: Remove Abs
...
Now unused. If we bring it back, it should be brought back as CSSC only. On
non-CSSC platforms, an explicit cmp + predicated neg can be better.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
ef544fecf2
OpcodeDispatcher: Use predicated neg for x87 fild
...
Saves an instruction on non-CSSC platforms by deleting a redundant cmp.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
5eea68d6c6
IR: Support predicated Neg
...
i.e. cneg. will be used for x87 hell op
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
5471367db1
OpcodeDispatcher: Optimize branches
...
For native cases. Big perf win.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
b8265b1067
OpcodeDispatcher: Rework SelectCC for branches
...
Separate out the NZCV bits from the more complex stuff so we can specially
optimize the branches.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
099c683a5a
IR: Add FromNZCV mode to CondJump
...
In this mode, rather than the branch comparing its arguments and then jumping
based on the result, the branch simply jumps by the native comparison based on
the NZCV value. This allows us to map x86 branches to arm64 branches 1:1.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-12 17:32:24 -04:00
Alyssa Rosenzweig
ec14a65e23
OpcodeDispatcher: optimize LOOP invert
...
need to reorder since the select clobbers nzcv. before/after diff on the loopne
unit test:
> 4308: [INFO] cset w20, ne
40c41
< 4308: [INFO] mrs x20, nzcv
---
> 4308: [INFO] mrs x21, nzcv
42,49c43,48
< 4308: [INFO] cset x21, ne
< 4308: [INFO] ubfx w22, w20, #30 , #1
< 4308: [INFO] eor x22, x22, #0x1
< 4308: [INFO] and x21, x21, x22
< 4308: [INFO] msr nzcv, x20
< 4308: [INFO] cbnz x21, #+0x8 (addr 0xfffed66f8094)
< 4308: [INFO] b #+0x1c (addr 0xfffed66f80ac)
< 4308: [INFO] ldr x0, pc+8 (addr 0xfffed66f809c)
---
> 4308: [INFO] cset x22, ne
> 4308: [INFO] and x20, x22, x20
> 4308: [INFO] msr nzcv, x21
> 4308: [INFO] cbnz x20, #+0x8 (addr 0xfffec94e8090)
> 4308: [INFO] b #+0x1c (addr 0xfffec94e80a8)
> 4308: [INFO] ldr x0, pc+8 (addr 0xfffec94e8098)
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 21:03:58 -04:00
Alyssa Rosenzweig
0d70c6a0d0
OpcodeDispatcher: Use cset for getting nzcv flags
...
Usually better in practice... some rotates are slightly regressed by this but
they were already terrible.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-11-10 21:03:58 -04:00