Commit Graph
467 Commits
Author SHA1 Message Date
Ryan Houdek cd83d3eb24 InstCountCI: Support multiple instructions in the tests
There are some cases where we want to test multiple instructions where
we can do optimizations that would overwise be hard to see.

eg:
```asm
; Can be optimized to a single stp
push eax
push ebx

; Can remove half of the copy since we know the direction
cld
rep movsb

; Can remove a redundant insert
addss xmm0, xmm1
addss xmm0, xmm2
```

This lets us have arbitrary sized code in instruction count CI, with the
original json key becoming only a label if the instruction array is
provided.

There are still some major limitations to this, instructions that
generate side-effects might have "garbage" after the end of the block
that isn't correctly accounted for. So care must be taken.

Example in the json
```json
"push ax, bx": {
  "ExpectedInstructionCount": 4,
  "Optimal": "No",
  "Comment": "0x50",
  "x86Insts": [
    "push ax",
    "push bx"
  ],
  "ExpectedArm64ASM": [
    "uxth w20, w4",
    "strh w20, [x8, #-2]!",
    "uxth w20, w7",
    "strh w20, [x8, #-2]!"
  ]
}
```
2023-10-09 21:49:53 -07:00
Ryan Houdek 462fff2c67 Merge pull request #3189 from Sonicadvance1/remove_warnings_15
FEXCore: Removes a warning about assume discarding side-effects
2023-10-09 17:26:14 -07:00
Ryan Houdek a1a479e69f FEXCore: Removes a warning about assume discarding side-effects 2023-10-09 16:04:57 -07:00
Ryan Houdek 6403290019 FEXCore: Renames raw FLAGS location names to signify they can't be used directly
Six of the EFLAGS can't be used directly in a bitmask because they are
either contained in a different flags location or has multiple bits
stored in it.

SF, ZF, CF, OF are stored in ARM's NZCV format in offset 24.
PF calculation is deferred but stored in the regular offset.
AF is also deferred in relation to the PF but stored in the regular
offset.

These /need/ to be reconstructed using the `ReconstructCompactedEFLAGS`
function when wanting to read the EFLAGS.

When setting these flags they /need/ to be set using
`SetFlagsFromCompactedEFLAGS`.

If either of these functions are not used when managing EFLAGs then the
internal representation will get mangled and the state will be
corrupted.

Having a little `_RAW` on these to signify that these aren't just
regular single bit representations like the other flags in EFLAGS should
make us puzzle about this issue before writing more broken code that
tries accessing it directly.
2023-10-08 11:51:11 -07:00
Ryan Houdek 22590dde77 FEXCore: Implements support for RPRES
This allows us to use reciprocal instructions which matches precision of
what x86 expects rather than converting everything to float divides.

Currently no hardware supports this, and even the upcoming X4/A720/A520
won't support it, but it was trivial to implement so wire it up.
2023-10-07 23:13:47 -07:00
Ryan Houdek 6543a80ff9 Merge pull request #3185 from Sonicadvance1/ir_dispatcher_emit
FEXCore/IR: Changes over to automated IR dispatch generation
2023-10-07 21:21:44 -07:00
Ryan Houdek 4cff3e5f1f FEXCore/IR: Changes over to automated IR dispatch generation
Suggested by Alyssa. Adding an IR operation can be a little tedious
since you need to add the definition to JIT.cpp for the dispatch switch,
JITClass.h for the function declared, and then actually defining the
implementation in the correct file.

Instead support the common case where an IR operation just gets
dispatched through to the regular handler. This lets the developer just
put the function definition in to the json and the relevent cpp file and
it just gets picked up.

Some minor things:
- Needs to support dynamic dispatch for {Load,Store}Register and
  {Load,Store}Mem
   - This is just a bool in the json
- It needs to not output JIT dispatch for some IR operations
   - SSE4.2 string instructions and x87 operations
   - These go down the "Unhandled" path
- Needs to support a Dispatcher function override
   - This is just for handling NoOp IR operations that get used for
     other reasons.
- Finally removes VSMul and VUMul, consolidating to VMul
   - Unlike V{U,S}Mull, signed or unsigned doesn't change behaviour here
- Fixed a couple random handler names not matching the IR operation
  name.
2023-10-07 15:01:47 -07:00
Ryan Houdek 559cf6491a InstCountCI: Support overriding AFP features
Also disable AFP under the vixl simulator by default since it doesn't support it.
2023-10-07 11:48:42 -07:00
Ryan Houdek 439a3b9c3a HostFeatures: Use a define 2023-10-06 09:33:41 -07:00
Ryan Houdek 5b7ba06d5c FEXCore: Support crypto extensions in HostFeatures override
Enables in InstCountCI so Pi users can run InstCountCI can run the tests
without breaking on crypto operations.

When crypto is enabled or disabled just wholesale change AES, CRC32, and
PMULL 128-bit in one step. We don't really care about partial support
here.
2023-10-05 17:41:08 -07:00
Ryan Houdek 8a51bb7a61 FEXCore: Support CpuState relative vector named constants
The motivation towards just having a pointer array in CpuState was that
initialization was fairly cheap and that we have limited space inside
the encoding depending on what we want to do.

Initialization cost is still a concern but doing a memcpy of 128-bytes
isn't that big of a deal.

Limited space in CpuState, while a concern isn't a significant one.
   - Needs to currently be less than 1 page in size
   - Needs to be under the architectural offset limitations of loadstore
     scaled offsets. Which is 65KB for 128-bit vectors

Still keeps the pointer array around for cases when we would need
synthesize an address offset and it's just easier to load the
process-wide table.

The performance improvement here is removing the dependency in the
ldr+ldr chain. In microbenchmarks this has shown to have an improvement
of ~4% by removing this dependency chain on Cortex-X1C.
2023-10-04 20:56:29 -07:00
Ryan Houdek fba7c4bedc IR/RA: Fixes register aliasing and pre-colouring for AVX
This is the cause of a bunch of redundant moves that shows up in
InstCountCI. Fixing this aliasing and pre-colouring issue causes a ton
of 256-bit operations to become optimal.
2023-10-04 10:04:06 -07:00
Ryan Houdek c52753e9c8 OpcodeDispatcher: Minor optimization in vzeroall
Using the cached zero value is less efficient than loading it in to the
register for all these cases.

Lets us use rename hardware more efficiently and removes a dependency
chain on a single register.

Original:
```
movi v2.2d, #0x0
mov z16.d, p7/m, z2.d
<... 16 more times>
mov z31.d, p7/m, z2.d
```

Result:
```
movi v16.2d, #0x0
<... 16 more times>
movi v31.2d, #0x0
```
2023-10-04 10:01:13 -07:00
Ryan Houdek e39634d314 Arm64: Fixes assert in VSQSHL/VSQSHR with SVE
When Dst != Vector then we need to pass Dst in to both Zd and Zdn.
Would have worked fine in a release build but assert build managed to
capture it.
2023-10-04 09:59:59 -07:00
Ryan Houdek 6964e65660 HostFeatures: Hardcode icache and dcache line size on x86
64-byte is effectively part of x86's ABI anyway. No need to query it for
our uses.
2023-10-02 16:26:14 -07:00
Ryan Houdek 11db8e7506 FEXCore: Wire up the new option to disable vixl indirect runtimes
Also so it compiles without the vixl simulator enabled.
2023-10-02 16:26:12 -07:00
Ryan Houdek b6b5e93dbb Config: Adds an option to disable vixl sim indirect runtime calls 2023-10-02 16:23:11 -07:00
Ryan Houdek 935b3a313a Merge pull request #3171 from Sonicadvance1/merge_dispatcher
FEXCore: Merge Arm64Dispatcher in to Dispatcher
2023-10-02 16:22:36 -07:00
CallumDev 9c25db83d9 JIT: VectorOps remove extraneous element size logs 2023-10-01 15:03:21 +10:30
CallumDev c42b581378 X87F64: Implement FABS with vector instruction 2023-10-01 14:39:55 +10:30
CallumDev d4a623a3fb InstCountCI Update 2023-10-01 11:22:18 +10:30
CallumDev c09c25005e X87F64: Use Bfe for rounding mode, FCHS use float instruction 2023-10-01 11:11:33 +10:30
Ryan Houdek 90570fd5f4 FEXCore: Merge Arm64Dispatcher in to Dispatcher
With the removal of the x86 JIT, there is no need to have these be
independent classes.

Merges the Arm64Dispatcher in to the base Dispatcher class.
No functional change, just moving code.
2023-09-30 09:31:55 -07:00
Ryan Houdek 98789a8039 FEXCore: Implement support for AVX2 feature detection 2023-09-28 19:57:08 -07:00
Ryan Houdek 6b4ff4ae81 Merge pull request #3163 from alyssarosenzweig/opt/ascii-flags
Optimize ASCII flags
2023-09-27 10:42:47 -07:00
Alyssa Rosenzweig 711583aa76 OpcodeDispatcher: Optimize PTEST flags
Zero NZCV first to avoid RMW.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:55:57 -04:00
Alyssa Rosenzweig 3efac9646c OpcodeDispatcher: Optimize ASCII flags
Make the zeroing of undefined NZCV more obvious. Mitigates regressions from
future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-27 10:31:31 -04:00
Alyssa Rosenzweig 3bb64c64e3 OpcodeDispatcher: Don't mask for TEST
Like AND.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:30:02 -04:00
Alyssa Rosenzweig a4de164944 OpcodeDispatcher: Use lshr for ah/bh with AllowUpperGarbage
If we ever get around to fusing ops with shifts in the ConstProp optimizer (may
or may not be worthwhile), this will delete an instruction from things like "or
al, bh".

Even though lsr is the same speed as bfe on Firestorm, I feel if you ask for
garbage you should get garbage C:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 20:28:01 -04:00
Alyssa Rosenzweig 45a645fbbc OpcodeDispatcher: Don't mask logic op inputs
Pointless, upper bits ignored anyway. Deletes piles of uxt and even some 32-bit
instruction moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:12:22 -04:00
Alyssa Rosenzweig 92211bf8c6 OpcodeDispatcher: Add AllowUpperGarbage option
To load 8-bit sources without bfe'ing for al/bl/cl if the caller knows it
doesn't need masking behaviour, but without lying about the size so the extract
for ah/bh/ch will still work properly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 19:08:20 -04:00
Alyssa Rosenzweig 7a06cc9727 IR: Use adcs/sbcs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-26 09:06:46 -04:00
Alyssa Rosenzweig 5facb21d30 OpcodeDispatcher: Don't mask small add/sub carries
For the GPR result, the masking already happens as part of the bfi. So the only
point of masking is for the flag calculation. But actually, every flag except
carry will ignore the upper bits anyway. And the carry calculation actually
WANTS the upper bit as a faster impl.

Deletes a pile of code both in FEX and the output :-)

ADC/SBC could probably get similar treatment later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-25 18:25:30 -04:00
Ryan Houdek 234e029391 Merge pull request #3145 from Sonicadvance1/optimize_inline_calls
PassManager: Optimize out CPUID and XGetBV calls
2023-09-24 18:09:18 -07:00
Alyssa Rosenzweig c8519b0b87 OpcodeDispatcher: Remove LoadPF
Now unused, its former users all prefer LoadPFRaw since they can fold in some of
this math into the use.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 68d32ad70d OpcodeDispatcher: Optimize PF in lahf
Use the raw popcount rather than the final PF and use some sneaky bit math to
come out 1 instruction ahead.

Closes #3117

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:59:28 -04:00
Alyssa Rosenzweig 1f02a6da34 IR: Add Ornror op
Mostly copypaste of Orlshl... we really should deduplicate this mess somehow.
Maybe a shift enum on the core Or op?

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 20:47:50 -04:00
Alyssa Rosenzweig 86063411dc Revert "OpcodeDispatcher: Use plain Lshl for flags"
This logic is unused since 8adfaa9aa ("OpcodeDispatcher: Use SelectCC for x87"),
which addressed the underlying issue.

This reverts commit df3833edbe.
2023-09-24 20:47:50 -04:00
Ryan Houdek 9968e6431f Passes: Rename SyscallOptimization
This is now inlining multiple external calls out of the JIT. Rename it
to InlineCallOptimization.
2023-09-24 17:25:38 -07:00
Ryan Houdek ff24f64b2a PassManager: Optimize out CPUID and XGetBV calls
If we const-prop the required functions and leafs then we can directly
encode the CPUID information rather than jumping out of the JIT.
In testing almost all CPUID executions const-prop which function is
getting called. Worst case that I found was only 85% const-prop rate.

This isn't quite 100% optimal since we need to call the RCLSE and
Constprop passes after we optimize these, which would remove some
redundant moves.

Sadly there seems to be a bug in the constprop pass that starts crashing
applications if that is done.
Easily enough tested by running Half-Life 2 and it immediately hitting
SIGILL.

Even without this optimization, this is stil a significant savings since
we aren't jumping out of the JIT anymore for these optimized CPUIDs.
2023-09-24 17:25:38 -07:00
Ryan Houdek e9a7ef2534 CPUID: Describe CPUID functions if they return constant state or not
Most CPUID routines return constant data, there are four that don't.
Some CPUID functions also need the leaf descriptor, so we need to
describe that as well.

Functions that don't return constant data:
- function 1Ah - Returns different data depending on current CPU core
- function 8000_000{2,3,4} - Different data based on CPU core

Functions that need leaf constprop:
- 4h, 7h, Dh, 4000_0001h, 8000_001Dh
2023-09-24 17:25:38 -07:00
Ryan Houdek 842c57e221 CPUID: Constify some functions
These don't modify CPUIDEmu state.
2023-09-24 17:25:38 -07:00
Alyssa Rosenzweig 8798e0cba0 Arm64: Rewrite Set/GetRoundingMode
I went auditing for places to use cset and what I found was hot garbage.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig c5fc03dac4 OpcodeDispatcher: Use cset for blsr/etc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig e63871ed2e OpcodeDispatcher: Handle sub in CalculateOF
Gets us the constant source optimization without more code duplication. And
honestly I prefer the combined presentation.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 19:52:35 -04:00
Alyssa Rosenzweig ea8b7633eb OpcodeDispatcher: Optimize OF calc of immediates
If we know the sign of one of the sources, we can do better when calculating OF.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-24 18:16:09 -04:00
Ryan Houdek 9ab2967d71 Arm64: Fixes wide shifts
movprfx is invalid to use when the source register matches the movprfx
destination.

This was getting picked up on by `TwoByte/0F_D1.asm` now that RCLSE is
working better now.
2023-09-23 06:06:18 -07:00
Ryan Houdek d01b457727 RCLSE: Optimize redundant store->load operations
The bug that was causing crashes with this was due to inline syscalls.
Now that this is fixed we can re-enable store->load operations.

This allows constant propagation to work significantly better, which
means inline syscalls start working again. This can significantly
improve syscall performance in some cases.

This is most likely to improve performance in dxsetup and vc_redist but
hard to get a real profile.

Additionally this will let us inline cpuid results in the future which
is pretty nice.
2023-09-23 06:06:18 -07:00
Mai 4e9a114858 Merge pull request #3142 from Sonicadvance1/inline_syscall_fix
Arm64: Fixes inline syscalls
2023-09-23 09:03:49 -04:00
Mai 72d092e951 Merge pull request #3141 from Sonicadvance1/fix_simm9_range
ConstProp: Fixes unscaled signed 9-bit range
2023-09-23 09:03:01 -04:00