Commit Graph
265 Commits
Author SHA1 Message Date
Ryan Houdek ca1c33047c Arm64: Only allocate vixl::Decoder if enabled
This class is very expensive to initialize so if you happen to have the
disassembler configuration enabled you were eating a very bad
initialization cost for no reason.

Only initialize the data member if any disassembler runtime option is
enabled, this completely removes the overhead.
2023-09-03 12:53:18 -07:00
Ryan Houdek b18592f153 Merge pull request #3050 from Sonicadvance1/fix_flag_reconstruction
OpcodeDispatcher: Fixes NZCV and PF flag compacting
2023-09-03 10:15:44 -07:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Ryan Houdek 44a14e7fd0 OpcodeDispatcher: Cleans up RFLAGS size handling
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.

Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
2023-09-03 01:41:04 -07:00
Mai 2d22176699 Merge pull request #3052 from Sonicadvance1/remove_todo_shlimm
OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
2023-09-03 04:38:52 -04:00
Ryan Houdek c425db7284 OpcodeDispatcher/Flags: Update SHLimm to use Opsize upfront
This one was easy, barely anything changes behaviour, as seen by
InstCountCI changes.
2023-09-03 00:27:28 -07:00
Ryan Houdek d0595b5f13 OpcodeDispatcher/Flags: Update ShiftLeft to use Opsize upfront 2023-09-03 00:16:26 -07:00
Ryan Houdek cb5d665046 OpcodeDispatcher: Fixes NZCV and PF flag compacting
Currently in main today, FEX fails to compact OF/CF/ZF/SF and PF.

This is due to recent optimizations with flag calculations on each of
these. Now that we have a centralized location where we compact and set
our internal representation of flags we can do this in one location.
2023-09-02 23:10:57 -07:00
Billy Laws 13b8f95f85 X87: Switch all stack pointer accesses to 32-bit OpSize 2023-09-02 09:17:33 -07:00
Billy Laws cb49373f47 FEXCore: Rework X87 tag word handling
The FXSAVE and FSAVE tag words are written out in different formats,
with FXSAVE using an abridged version that lacks the zero/special/valid
distinction. Switch to using this abridged version internally for
simplicity, and to allow the calculation of zero/special/valid
distinction to be deferred until an fxsave instruction (in the future,
currently the distinction is ignored and only valid/empty states are
possible).
2023-09-02 09:17:33 -07:00
Ryan Houdek 435f03c703 Context: Adds helper to reconstruct and consume packed EFLAGS
Currently FEX's internal EFLAGS representation is a perfect 1:1 mapping
between bit offset and byte offset. This is going to change with #3038.
There should be no reason that the frontend needs to understand how to
reconstruct the compacted flags from the internal representation.

Adds context helpers and moves all the logic to FEXCore. The locations
that previously needed to handle this have been converted over to use
this.
2023-09-02 07:05:54 -07:00
Alyssa Rosenzweig 659568ec1b ConstProp: Propagate 0 to first argument of SubNZCV
This allows inlining a constant into the comparison for Neg.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:06:01 -04:00
Alyssa Rosenzweig dcb09b085b OpcodeDispatcher: Use AddNZCV/SubNZCV
32-bit or 64-bit addition without carry-in. This matches the baseline hardware
semantic. Generalizing to support other cases can come later, this should be a
win already.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:41 -04:00
Alyssa Rosenzweig 45d99a2cce OpcodeDispatcher: Fix source sizes for Sub flags
For correct carry/overflow behaviour, we need to use a compare of the right size. The existing logic
to look at the source sizes doesn't work for this, since a 32-bit NEG instruction will compare a
32-bit source with a 64-bit _Constant(0) .. which needs a 32-bit compare but the existing logic
would use a 64-bit compare. This is not yet a bug fix, since the overflow code is currently in
software for 32-bit negates so it's irrelevant. But it should prevent regressions from using native
compares later in this series. Presumably this was intended all along but left as-is to avoid
disturbing instcountci once noticed. Time to disturb CI!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 09:05:39 -04:00
Alyssa Rosenzweig f688364bf2 IR: Add AddNZCV/SubNZCV op
AddNZCV is a new op to return the NZCV for an addition directly, which lets us
skip software flag calculation in some cases. In the future it would be nice to
fuse this into the Add itself as a second destination to avoid repeating the
addition, but that's a very involved change and right now I'm building FEX on an
old Chromebook because my M1 kernel is FUBAR.

Similarly, SubNZCV returns flags for Sub. This has the extra twist of needing to
invert the carry bit due to the inverted definition between arm64 and x86_64.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-31 07:52:12 -04:00
Mai 7c81a0d4fe Merge pull request #3042 from Sonicadvance1/optimize_call
OpcodeDispatcher: Optimize calls with push
2023-08-31 02:23:06 -04:00
Ryan Houdek 61df7a576a OpcodeDispatcher: Be super defensive when starting a new block
Ensure all cached data is correct.
2023-08-30 20:43:11 -07:00
Ryan Houdek ffaa908475 OpcodeDispatcher: Use the named zero register for each usage
This will let us reuse it in some cases. In some of these
implementations there is a bad code smell around using the zero register
but that isn't going to get solved in this commit.
2023-08-30 20:43:11 -07:00
Ryan Houdek a2307f28d6 Arm64: Optimize AES operations by caching a zero register
A bunch of the AES operations take a zero register upfront and we
currently materialize it for each instruction.
Considering that most AES operations are used back to back, we can
eliminate these materializations by caching it between instructions.

Additionally removes a move in the optimal case when destination matches
the state register, which is exactly what the SSE operation ends up
doing.

AESKeyGenAssist has an edge case that if the destination RA overlaps the
zero register then we still need to eat a move, hopefully doesn't happen
too frequently in practice. This is also the lesser used instruction so
it isn't a big deal. RA constraints could solve that still.
2023-08-30 19:00:43 -07:00
Ryan Houdek 1446d4fe12 IR: Adds support for named vector zero
This is useful for caching a zero register vector which we use in
various locations. This will be abused soon.
2023-08-30 18:59:38 -07:00
Ryan Houdek 24215f7ad0 OpcodeDispatcher: Optimize BLENDV when xmm0 is one of the sources
This instruction has xmm0 be one of the implicit sources. We were
loading xmm0 twice. #2700 would also fix this but that breaks other
things for some reason.
2023-08-30 17:09:34 -07:00
Ryan Houdek b20c518bf0 OpcodeDispatcher: Optimize calls with push
InstCountCI doesn't cover branch instructions so needs manual
inspection.
2023-08-30 16:25:01 -07:00
Ryan Houdek 81a32c3998 FEXCore: Allows disabling telemetry at runtime
This is useful for InstCountCI so you can disable the telemetry
gathering even if enabled so it doesn't affect the CI system.
2023-08-30 12:59:41 -07:00
Mai 02b891c0fe Merge pull request #3040 from Sonicadvance1/optimize_aeskeygen
Arm64: Optimize AESKeyGenAssist
2023-08-30 15:58:29 -04:00
Ryan Houdek d8f131fa3d Arm64: Optimize AESKeyGenAssist
We can load the swizzle table from our constant pool now. This removes
the only usage of VTMP3 from our Arm64 JIT.

I would say the this is now optimal for the version without RCON set.
With RCON we could technically make some of the move of the constant
more optimal.
2023-08-30 12:15:09 -07:00
Ryan Houdek 9bfd4b650f OpcodeDispatcher: Removes erroneous debug log 2023-08-30 11:12:12 -07:00
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00
Ryan Houdek 8bc84c202c IR: Removes implicit sized xor 2023-08-28 19:51:14 -07:00
Ryan Houdek e9a3848602 Merge pull request #3027 from Sonicadvance1/remove_implicit_andn
IR: Removes implicit sized andn
2023-08-28 19:39:49 -07:00
Ryan Houdek 1699ec9a76 IR: Removes implicit sized andn 2023-08-28 19:16:16 -07:00
Ryan Houdek db6c8852fc IR: Removes implicit sized or 2023-08-28 19:06:05 -07:00
Ryan Houdek 65dc6f3e90 IR: Removes implicit sized lshr 2023-08-28 18:16:56 -07:00
Ryan Houdek 60c4438780 IR: Removes implicit sized lshl 2023-08-28 17:50:41 -07:00
Ryan Houdek 898ce1ce8f IR: Removes implicit sized UMulH 2023-08-28 17:20:55 -07:00
Ryan Houdek aa8dfd6af1 IR: Removes implicit sized UMul 2023-08-28 17:20:55 -07:00
Ryan Houdek 6a6d808b0d IR: Removes implicit sized MulH 2023-08-28 17:20:55 -07:00
Ryan Houdek fac5b2ac72 IR: Removes implicit sized Mul 2023-08-28 17:20:55 -07:00
Ryan Houdek b9e4a1423f IR: Removes sext IR helper
You hold no power here IR operation.
2023-08-28 17:03:38 -07:00
Ryan Houdek c0bb6a053f IR: Removes implicit sized {Create,Extract}ElementPair 2023-08-28 16:50:00 -07:00
Ryan Houdek bc1e89d91d Merge pull request #3020 from Sonicadvance1/remove_implicit_ops_pt_atomic
Remove implicit sized IR ops part atomic
2023-08-28 07:27:47 -07:00
Ryan Houdek c1b4c11e54 Merge pull request #3018 from Sonicadvance1/opcodedispatcher_sizeopt
OpcodeDispatcher: Optimize Get{Src,Dst}Size
2023-08-28 07:13:41 -07:00
Ryan Houdek a36427d01e IR: Removes implicit sized CAS 2023-08-28 07:12:31 -07:00
Ryan Houdek 594baff705 IR: Removes implicit sized CASPair 2023-08-28 07:12:31 -07:00
Ryan Houdek 6dcfd6eb73 IR: Removes non-opsize AtomicXor 2023-08-28 07:12:31 -07:00
Ryan Houdek 4f8a63459c IR: Removes non-opsize AtomicSwap 2023-08-28 07:12:31 -07:00
Ryan Houdek ab230bf527 IR: Removes non-opsize AtomicFetchAdd 2023-08-28 07:12:31 -07:00