Commit Graph
115 Commits
Author SHA1 Message Date
Ryan Houdek 4cff3e5f1f FEXCore/IR: Changes over to automated IR dispatch generation
Suggested by Alyssa. Adding an IR operation can be a little tedious
since you need to add the definition to JIT.cpp for the dispatch switch,
JITClass.h for the function declared, and then actually defining the
implementation in the correct file.

Instead support the common case where an IR operation just gets
dispatched through to the regular handler. This lets the developer just
put the function definition in to the json and the relevent cpp file and
it just gets picked up.

Some minor things:
- Needs to support dynamic dispatch for {Load,Store}Register and
  {Load,Store}Mem
   - This is just a bool in the json
- It needs to not output JIT dispatch for some IR operations
   - SSE4.2 string instructions and x87 operations
   - These go down the "Unhandled" path
- Needs to support a Dispatcher function override
   - This is just for handling NoOp IR operations that get used for
     other reasons.
- Finally removes VSMul and VUMul, consolidating to VMul
   - Unlike V{U,S}Mull, signed or unsigned doesn't change behaviour here
- Fixed a couple random handler names not matching the IR operation
  name.
2023-10-07 15:01:47 -07:00
Ryan Houdek 6b4ff4ae81 Merge pull request #3163 from alyssarosenzweig/opt/ascii-flags
Optimize ASCII flags
2023-09-27 10:42:47 -07:00
Alyssa Rosenzweig 3efac9646c OpcodeDispatcher: Optimize ASCII flags
Make the zeroing of undefined NZCV more obvious. Mitigates regressions from
future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:31:31 -04:00
Alyssa Rosenzweig 3bb64c64e3 OpcodeDispatcher: Don't mask for TEST
Like AND.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:02 -04:00
Alyssa Rosenzweig a4de164944 OpcodeDispatcher: Use lshr for ah/bh with AllowUpperGarbage
If we ever get around to fusing ops with shifts in the ConstProp optimizer (may
or may not be worthwhile), this will delete an instruction from things like "or
al, bh".

Even though lsr is the same speed as bfe on Firestorm, I feel if you ask for
garbage you should get garbage C:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:28:01 -04:00
Alyssa Rosenzweig 45a645fbbc OpcodeDispatcher: Don't mask logic op inputs
Pointless, upper bits ignored anyway. Deletes piles of uxt and even some 32-bit
instruction moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:12:22 -04:00
Alyssa Rosenzweig 92211bf8c6 OpcodeDispatcher: Add AllowUpperGarbage option
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:08:20 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Ryan Houdek ff24f64b2a PassManager: Optimize out CPUID and XGetBV calls
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.

This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.

Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.

Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
2023-09-24 17:25:38 -07:00
Alyssa Rosenzweig 699aa85c4b OpcodeDispatcher: Opt PF selection
Fold the and in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Alyssa Rosenzweig 2d65a3677b OpcodeDispatcher: Optimize NZCV selects
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Billy Laws d641d3f61e OpcodeDispatcher: Avoid redundantly passing args to WIN32 ABI syscalls 2023-09-22 10:12:39 -07:00
Ryan Houdek 1a4d1d820b OpcodeDispatcher: Optimize lock btr
This is an atomicFetchCLR, removes two mvn instructions that are back to
back negating the source.

We didn't have this instruction combination in InstCountCI so will be a
bit hard to see.
2023-09-21 14:54:51 -07:00
Ryan Houdek 8e9e87f631 OpcodeDispatcher: Removes non-explicit SelectCC function
Renames the explicit sized one to `SelectCC`
Cleans up a bit of duplicated code.
2023-09-21 01:56:38 -07:00
Ryan Houdek 67680d71a4 Merge pull request #3125 from Sonicadvance1/spdx_fexcore
FEXCore: Adds SPDX identifier
2023-09-19 17:42:07 -07:00
Ryan Houdek e4613477b1 FEXCore/Interface/Core: Adds SPDX identifier 2023-09-19 17:33:15 -07:00
Alyssa Rosenzweig 25943d1d17 OpcodeDispatcher: Sigh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek ad8b0c673f Merge pull request #3109 from lioncash/shlx
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
2023-09-15 18:49:36 -07:00
Lioncache e9be291cec OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
We can remove some unnecessary moves for the 32-bit cases and
collapse the operations down to a single instruction.
2023-09-15 21:05:50 -04:00
Lioncache d4f87c7db1 OpcodeDispatcher: Improve output of MULX
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
2023-09-15 20:52:02 -04:00
Lioncache 4a37ea4819 OpcodeDispatcher: Handle RORX corner cases better
There are a few cases where we were emitting code when we
didn't really need to, or could emit less.
2023-09-15 17:36:36 -04:00
Ryan Houdek d5b58eebaf OpcodeDispatcher: Optimize cmov
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason

A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
2023-09-15 10:09:37 -07:00
Ryan Houdek 6dbbd9ecfc OpcodeDispatcher: Duplicate SelectCC but with Explicit result size
This is a temporary measure as we are moving Select operations over to
explicit sizes. Once we remove all uses of SelectCC then it will get
removed.
2023-09-15 10:09:37 -07:00
Mai 96bbd01ad6 Merge pull request #3096 from Sonicadvance1/optimal_crc
OpcodeDispatcher: Optimize CRC32
2023-09-15 05:19:05 -04:00
Ryan Houdek a3115d4699 OpcodeDispatcher: Optimize CRC32
The only version of this instruction that was generating optimal code
was the one with 64-bit destination and source.

Optimizes the rest of the operating sizes so that they are all optimal
at one instruction translations
2023-09-14 16:21:42 -07:00
Ryan Houdek 80cda1bb18 OpcodeDispatcher: Optimize 16-bit MOVBE
16-bit MOVBE is a bit of a special case where it loads 16-bits in to the
bottom of the GPR without clearing the upper bits of the register.
Which means 32-bits or 64-bits depending on operating mode.

Arm64 doesn't support a 16-bit bswap so it needs to operate at 32-bits
instead. We then can insert the resulting bits of the 32-bit rev with a
bfxil in to the lower bits of the resulting destination register.

This allows 16-bit movbe to be optimal now.
2023-09-14 16:05:16 -07:00
Ryan Houdek c362d3a9d8 OpcodeDispatcher: Optimize 32-bit bswap
Removes a redundant move, making it optimal now.
2023-09-12 16:19:10 -07:00
Mai 48521a4416 Merge pull request #3075 from Sonicadvance1/optimize_bt_ops
OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
2023-09-11 16:05:33 -04:00
Mai 950007c815 Merge pull request #3071 from Sonicadvance1/update_rcl_opsize
OpcodeDispatcher: Update 32/64-bit RCL for operating size
2023-09-11 15:34:06 -04:00
Ryan Houdek 3d1bbe505d OpcodeDispatcher: Minor optimization to BT/BTC/BTR/BTS
These instructions set all the flags to undefined and moves the
resulting bit in to CF. No need to calculate the deferred flags when
we are about to write over them.
2023-09-10 20:16:29 -07:00
Ryan Houdek 6c62691af0 OpcodeDispatcher: Update 32/64-bit RCL for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:40:12 -07:00
Ryan Houdek 47f50a7008 OpcodeDispatcher: Update 32/64-bit RCR for operating size
Removes todo from explicit size PR. Saves one instruction.
2023-09-09 09:33:27 -07:00
Alyssa Rosenzweig e6db2d0b96 IR: Remove phi nodes
It turns out that pure SSA isn't a great choice for the sort of emulation we do.
On one hand, it discards information from the guest binary's register allocation
that would let us skip stuff. On the other hand, it doesn't have nearly as many
benefits in this setting as in a traditional compiler... We really *don't* want
to do global RA or really any global optimization. We assume the guest optimizer
did its job for x86, we just need to clean up the mess left from going x86 ->
arm. So we just need enough SSA to peephole optimize.

My concrete IR proposals are that:

  * SSA values must be killed in the same block that they are defined.
  * Explicit LoadGPR/StoreGPR instructions can be used for global persistence.
  * LoadGPR/StoreGPR are eliminated in favour of SSA within a block.

This has a lot of nice properties for our setting:

  * Except for some internal REP instruction emulation (etc), we already have
    registers for everything that escapes block boundaries, so this form is very
    easy to go into -- straightforward local value numbering, not a full into
    SSA pass.

  * Spilling is entirely local (if it happens at all), since everything is in
    registers at block boundaries. This is excellent, because Belady's algorithm
    lets us spill nearly optimally in linear-time for individual blocks. (And
    the global version of Belady's algorithm is massively more complicated...)
    A nice fit for a JIT.

    Relatedly, it turns out allowing spilling is probably a decent decision,
    since the same spiller code can be used to rematerialize constants in a
    straightforward way. This is an issue with the current RA.

  * Register assignment is entirely local. For the same reason, we can assign
    registers "optimally" in linear time & memory (e.g. with linear scan). And
    the impl is massively simpler than a full blown SSA-based tree scan RA. For
    example, we don't have to worry about parallel copies or coalescing phis or
    anything. Massively nicer algorithm to deal with.

  * SSA value names can be block local which makes the validation implicit :~)

It also has remarkably few drawbacks, because we didn't want to do CFG global
optimization anyway given our time budget and the diminishng returns. The few
global optimizations we might want (flag escape analysis?) don't necessarily
benefit from pure SSA anyway.

Anyway, we explicitly don't want phi nodes in any of this. They're currently
unused. Let's just remove them so nobody gets the bright idea of changing that.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 16:35:12 -04:00
Alyssa Rosenzweig 79a20b899b Remove ABINoPF option
Now that PF calculation is deferred, the cost of calculating PF correctly should
be tolerable. Remove the speed hack to skip PF. It's fundamentally broken, and
there are enough broken things in FEX as it is that we don't need to maintain
this one ;-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:56:43 -04:00
Alyssa Rosenzweig 02c864d837 OpcodeDispatcher: Defer second XOR for AF
AF is calculated as:

  ((Src1 ^ Src2) ^ Res)[4]

Due to the extract, this is equivalent to

  ((Src1 ^ Src2) ^ (Res ^ 1))[4]

We already store (Res ^ 1) as the PF byte. So, it suffices to store

  AF Byte = Src1 ^ Src2

and then we can recover the flag value

  AF = (AF Byte ^ PF Byte)[4]

This saves an instruction from the AF calculation. It does couple PF/AF writes.
In practice, most instructions fall into one of these categories:

  * Both PF and AF written together, the coupling is correct.
  * PF written but AF invalidated, irrelevant.
  * Both invalidated, irrelevant.

None of these require special handling. Where we do need special handling is
when we want to write them separately, in which case we can fix-up the value of
AF as appropriate.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 14:21:18 -04:00
Alyssa Rosenzweig b6462ee854 OpcodeDispatcher: Optimize CMC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 13:55:05 -04:00
Alyssa Rosenzweig 5edd9cb35b OpcodeDispatcher: Use SetAF
For constants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:34:12 -04:00
Alyssa Rosenzweig f46ba52e0e OpcodeDispatcher: Use LoadAF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-05 11:33:36 -04:00
Ryan Houdek 5cc6eff62c OpcodeDispatcher: Remove final assumptions about small IR operating sizes
With Or, Orlshl, Bfe, and Bfi there were some assumptions made that i8
and i16 operations made sense. Which required us to disable the IR
validation for these operations when it was just added.

This removes the final assumptions about these IR operations supporting
these small operating sizes allowing us to enable the IR validation.

Also a very minor optimization by moving a couple extracts from source
before trying to BFI from it, making RA more optimal.
2023-09-03 02:25:34 -07:00
Ryan Houdek 44a14e7fd0 OpcodeDispatcher: Cleans up RFLAGS size handling
When moving everything away from implicit size handling, I kept this the
same codegen even though it was uglier.

Now that implicit stuff is mostly done, switch this over to 32-bit
operations. The behaviour of these changes is no functional change, just
cleans up the operations.
2023-09-03 01:41:04 -07:00
Mai 7c81a0d4fe Merge pull request #3042 from Sonicadvance1/optimize_call
OpcodeDispatcher: Optimize calls with push
2023-08-31 02:23:06 -04:00
Ryan Houdek 61df7a576a OpcodeDispatcher: Be super defensive when starting a new block
Ensure all cached data is correct.
2023-08-30 20:43:11 -07:00
Ryan Houdek b20c518bf0 OpcodeDispatcher: Optimize calls with push
InstCountCI doesn't cover branch instructions so needs manual
inspection.
2023-08-30 16:25:01 -07:00
Ryan Houdek 81a32c3998 FEXCore: Allows disabling telemetry at runtime
This is useful for InstCountCI so you can disable the telemetry
gathering even if enabled so it doesn't affect the CI system.
2023-08-30 12:59:41 -07:00
Ryan Houdek f741ebf970 IR: Removes implicit sized add
Saw a few locations in here that we operate things at 64-bit
unconditionally around pointer calculation. Will be coming back for
those when running in 32-bit mode.

This is the last of the implicit sized ALU operations! After this I'll
be going through the IR more individually to try and remove any
stragglers.
Then should be able to start cleaning up and actually optimizing GPR
operations.
2023-08-29 22:26:51 -07:00
Ryan Houdek e8b767b553 IR: Removes implicit sized bfe
This one is a bit of a mess, looking forward to coming back and cleaning
this up.
2023-08-29 19:43:39 -07:00
Ryan Houdek 9e70aa4192 IR: Removes implicit sized and 2023-08-28 22:43:21 -07:00
Ryan Houdek b5dc6a69c7 IR: Removes implicit sized sub 2023-08-28 22:05:02 -07:00
Ryan Houdek a276b37252 IR: Removes bfi from variable size
This one was already explicit sized. Just convert it over to OpSize.
2023-08-28 21:31:37 -07:00