Commit Graph
7759 Commits
Author SHA1 Message Date
Ryan Houdek 0d9dce987d Merge pull request #3126 from neobrain/feature_better_wayland_thunks64
Thunks/wayland: Add support for APIs required by zink and Super Meat Boy
2023-09-19 10:44:53 -07:00
Ryan Houdek 65d558b2c4 Merge pull request #3119 from alyssarosenzweig/opt/x87-sel
Make x87 FCMOV slightly less terrible
2023-09-19 10:34:29 -07:00
Tony Wasserka b00d413961 Thunks/wayland: Add more message signatures required by Super Meat Boy with zink 2023-09-19 17:33:24 +02:00
Tony Wasserka 6b54540756 Thunks/wayland: Add support for message signatures with nullable arguments 2023-09-19 17:33:24 +02:00
Tony Wasserka 356a42d330 Thunks/wayland: Reorder listener signatures alphabetically 2023-09-19 17:33:24 +02:00
Tony Wasserka 8fcf419183 Thunks/wayland: Add more functions required by Super Meat Boy via libdecor and SDL 2023-09-19 17:33:23 +02:00
Alyssa Rosenzweig 83c8b64c50 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 25943d1d17 OpcodeDispatcher: Sigh.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig bf03dab295 Arm64: Use csetm
Saves some moves.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig 8adfaa9aa6 OpcodeDispatcher: Use SelectCC for x87
Better code gen and will benefit from future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-19 08:37:54 -04:00
Alyssa Rosenzweig 3b188b7f49 Merge pull request #3118 from Sonicadvance1/spdx_fhu
FHU: Prepend SPDX identifier
2023-09-18 18:48:30 -04:00
Ryan Houdek 94bbd415a2 FHU: Prepend SPDX identifier
Added with `sed -i '1 i\\/\/ SPDX-License-Identifier: MIT' *.h`
2023-09-18 11:45:18 -07:00
Ryan Houdek 2ea2300408 Merge pull request #3110 from Sonicadvance1/buffered_jit_symbols
FEXCore/JitSymbols: Buffer writes to reduce overhead
2023-09-18 11:38:06 -07:00
Ryan Houdek 000fb2efae Merge pull request #3068 from neobrain/feature_thunk_testlib
unittests: Add test thunk library
2023-09-18 10:31:42 -07:00
Ryan Houdek 8b523082af Merge pull request #3116 from alyssarosenzweig/minor/flag-opts
Minor/flag opts
2023-09-18 10:28:51 -07:00
Alyssa Rosenzweig 5d2a3cd322 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Alyssa Rosenzweig df3833edbe OpcodeDispatcher: Use plain Lshl for flags
If we have PF but no CF this simplifies the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 11:01:46 -04:00
Tony Wasserka 527b65648f unittests: Enable logging to stderr when invoking FEXLoader 2023-09-18 16:53:35 +02:00
Tony Wasserka bef64c53f8 unittests: Add test thunk library 2023-09-18 16:53:35 +02:00
Alyssa Rosenzweig 8edcd31404 OpcodeDispatcher: Avoid inverting PF
..if we can fold the invert into the reader.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-18 10:35:39 -04:00
Ryan Houdek fd1b639ad9 Merge pull request #3115 from lioncash/sqxtun
Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
2023-09-17 14:51:27 -07:00
Ryan Houdek 950a8dbfe7 Merge pull request #3114 from lioncash/ins
Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
2023-09-17 14:40:38 -07:00
Lioncache 26e4d8ad59 Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
If the destination and lower data alias, we can
avoid needing to move into a temporary.
2023-09-17 17:37:36 -04:00
Lioncache d54f590b14 Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
If none of the vectors alias the destination, then we can eliminate
an extra move and usage of a temporary.
2023-09-17 16:56:23 -04:00
Ryan Houdek b3269f20ef Merge pull request #3113 from lioncash/shrn
Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
2023-09-17 13:03:42 -07:00
Lioncache 047646be6d Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
In the event the destination and lower source are the same, then
we don't need to perform any moves.
2023-09-17 15:49:54 -04:00
Ryan Houdek 8168a49d10 Merge pull request #3112 from lioncash/assert
Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
2023-09-17 12:40:47 -07:00
Lioncache 7f2fd4e9a0 Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
Ensures that if our temp vectors change in the future that this is
caught at compile-time rather than runtime.
2023-09-17 15:22:39 -04:00
Lioncache 4ea9f08425 ARMEmitter: Mark index and conversion ops as constexpr
Will be used for assertions. Also makes registers more flexible
for compile-time stuff in general.
2023-09-17 15:19:44 -04:00
Ryan Houdek ffb58761c1 Merge pull request #3111 from lioncash/shift
Arm64/VectorOps: Fix SVE aliasing-path  move in VSShr
2023-09-17 12:15:49 -07:00
Lioncache 8ecdb341e2 Arm64/VectorOps: Fix SVE aliasing-path move in VSShr
Seems like this was a typo from 8d11073, since we'd be moving
into a temporary and then never use it.
2023-09-17 14:43:05 -04:00
Ryan Houdek 0c5c146fcf FEXCore/JitSymbols: Buffer writes to reduce overhead
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.

The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.

To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
  clang benchmark and seeing how frequently it was still writing.
   - 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
  symbols.
   - In most cases 100ms is enough that you won't notice the blip in
     perf.

One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Ryan Houdek ad8b0c673f Merge pull request #3109 from lioncash/shlx
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
2023-09-15 18:49:36 -07:00
Ryan Houdek e574cfe681 Merge pull request #3108 from lioncash/mulx
OpcodeDispatcher: Improve output of MULX
2023-09-15 18:09:02 -07:00
Lioncache e9be291cec OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
We can remove some unnecessary moves for the 32-bit cases and
collapse the operations down to a single instruction.
2023-09-15 21:05:50 -04:00
Lioncache d4f87c7db1 OpcodeDispatcher: Improve output of MULX
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
2023-09-15 20:52:02 -04:00
Ryan Houdek 4604c01986 Merge pull request #3107 from lioncash/pext
Arm64/ALUOps: Remove spills in PEXT
2023-09-15 17:40:44 -07:00
Lioncache b0c8ff0ea6 Arm64/ALUOps: Remove spills in PEXT
Reduces the number of emitted instructions for a
corresponding PEXT instruction.

We no longer spill for this IR op.
2023-09-15 19:39:51 -04:00
Ryan Houdek 647629ac23 Merge pull request #3105 from lioncash/rorx
OpcodeDispatcher: Handle RORX corner cases better
2023-09-15 14:55:17 -07:00
Lioncache be90e76422 Arm64/ALUOps: mov in the case of full 32-bit/64-bit BFE
Allows register-renaming mechanisms to be invoked more frequently
2023-09-15 17:38:01 -04:00
Lioncache 4a37ea4819 OpcodeDispatcher: Handle RORX corner cases better
There are a few cases where we were emitting code when we
didn't really need to, or could emit less.
2023-09-15 17:36:36 -04:00
Ryan Houdek 6e08ac65b9 Merge pull request #3106 from lioncash/clwb
HostFeatures: Fix x86 CLWB support check
2023-09-15 14:06:54 -07:00
Lioncache 8705de1893 HostFeatures: Fix x86 CLWB support check
This was clobbering the BMI2 boolean unintentionally.
2023-09-15 16:38:45 -04:00
Alyssa Rosenzweig c8e7c347c3 Merge pull request #3100 from Sonicadvance1/optimize_cmov
OpcodeDispatcher: Optimize cmov
2023-09-15 15:17:51 -04:00
Ryan Houdek 3d0b66407e Merge pull request #3104 from lioncash/vperm2
InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
2023-09-15 11:27:11 -07:00
Lioncache c86b6dc690 InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
Allows viewing the codegen for cases where conditional zeroing is performed.

Also fixes up the vperm2f variants shorthanding one of the registers
to make everything a little more explicit.
2023-09-15 14:10:11 -04:00
Ryan Houdek 773e9465bc Merge pull request #3103 from lioncash/warn
DeadContextStoreElimination: Silence unused function warning
2023-09-15 10:57:40 -07:00
Lioncache e1ed7f43fd DeadContextStoreElimination: Turn LastAccessType into an enum class
Makes the type stricter in terms of implicit conversions.
2023-09-15 13:23:00 -04:00
Ryan Houdek 3eb501aa27 InstCountCI: Update for optimized NZCV and cmov 2023-09-15 10:11:50 -07:00
Ryan Houdek d5b58eebaf OpcodeDispatcher: Optimize cmov
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason

A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
2023-09-15 10:09:37 -07:00