Ryan Houdek
0d9dce987d
Merge pull request #3126 from neobrain/feature_better_wayland_thunks64
...
Thunks/wayland: Add support for APIs required by zink and Super Meat Boy
2023-09-19 10:44:53 -07:00
Ryan Houdek
65d558b2c4
Merge pull request #3119 from alyssarosenzweig/opt/x87-sel
...
Make x87 FCMOV slightly less terrible
2023-09-19 10:34:29 -07:00
Tony Wasserka
b00d413961
Thunks/wayland: Add more message signatures required by Super Meat Boy with zink
2023-09-19 17:33:24 +02:00
Tony Wasserka
6b54540756
Thunks/wayland: Add support for message signatures with nullable arguments
2023-09-19 17:33:24 +02:00
Tony Wasserka
356a42d330
Thunks/wayland: Reorder listener signatures alphabetically
2023-09-19 17:33:24 +02:00
Tony Wasserka
8fcf419183
Thunks/wayland: Add more functions required by Super Meat Boy via libdecor and SDL
2023-09-19 17:33:23 +02:00
Alyssa Rosenzweig
83c8b64c50
InstCountCI: Update
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig
25943d1d17
OpcodeDispatcher: Sigh.
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig
bf03dab295
Arm64: Use csetm
...
Saves some moves.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-19 08:50:40 -04:00
Alyssa Rosenzweig
8adfaa9aa6
OpcodeDispatcher: Use SelectCC for x87
...
Better code gen and will benefit from future work.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-19 08:37:54 -04:00
Alyssa Rosenzweig
3b188b7f49
Merge pull request #3118 from Sonicadvance1/spdx_fhu
...
FHU: Prepend SPDX identifier
2023-09-18 18:48:30 -04:00
Ryan Houdek
94bbd415a2
FHU: Prepend SPDX identifier
...
Added with `sed -i '1 i\\/\/ SPDX-License-Identifier: MIT' *.h`
2023-09-18 11:45:18 -07:00
Ryan Houdek
2ea2300408
Merge pull request #3110 from Sonicadvance1/buffered_jit_symbols
...
FEXCore/JitSymbols: Buffer writes to reduce overhead
2023-09-18 11:38:06 -07:00
Ryan Houdek
000fb2efae
Merge pull request #3068 from neobrain/feature_thunk_testlib
...
unittests: Add test thunk library
2023-09-18 10:31:42 -07:00
Ryan Houdek
8b523082af
Merge pull request #3116 from alyssarosenzweig/minor/flag-opts
...
Minor/flag opts
2023-09-18 10:28:51 -07:00
Alyssa Rosenzweig
5d2a3cd322
InstCountCI: Update
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-18 11:01:46 -04:00
Alyssa Rosenzweig
df3833edbe
OpcodeDispatcher: Use plain Lshl for flags
...
If we have PF but no CF this simplifies the IR.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-18 11:01:46 -04:00
Tony Wasserka
527b65648f
unittests: Enable logging to stderr when invoking FEXLoader
2023-09-18 16:53:35 +02:00
Tony Wasserka
bef64c53f8
unittests: Add test thunk library
2023-09-18 16:53:35 +02:00
Alyssa Rosenzweig
8edcd31404
OpcodeDispatcher: Avoid inverting PF
...
..if we can fold the invert into the reader.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2023-09-18 10:35:39 -04:00
Ryan Houdek
fd1b639ad9
Merge pull request #3115 from lioncash/sqxtun
...
Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
2023-09-17 14:51:27 -07:00
Ryan Houdek
950a8dbfe7
Merge pull request #3114 from lioncash/ins
...
Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
2023-09-17 14:40:38 -07:00
Lioncache
26e4d8ad59
Arm64/VectorOps: Elide moves where applicable in 128-bit VSQXTUN2
...
If the destination and lower data alias, we can
avoid needing to move into a temporary.
2023-09-17 17:37:36 -04:00
Lioncache
d54f590b14
Arm64/VectorOps: Improve handling of 128-bit vector VInsElement
...
If none of the vectors alias the destination, then we can eliminate
an extra move and usage of a temporary.
2023-09-17 16:56:23 -04:00
Ryan Houdek
b3269f20ef
Merge pull request #3113 from lioncash/shrn
...
Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
2023-09-17 13:03:42 -07:00
Lioncache
047646be6d
Arm64/VectorOps: Elide moves in ASIMD VUShrNI2 if possible
...
In the event the destination and lower source are the same, then
we don't need to perform any moves.
2023-09-17 15:49:54 -04:00
Ryan Houdek
8168a49d10
Merge pull request #3112 from lioncash/assert
...
Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
2023-09-17 12:40:47 -07:00
Lioncache
7f2fd4e9a0
Arm64/VectorOps: Assert VTMP1 and VTMP2 are sequential in VTBL2
...
Ensures that if our temp vectors change in the future that this is
caught at compile-time rather than runtime.
2023-09-17 15:22:39 -04:00
Lioncache
4ea9f08425
ARMEmitter: Mark index and conversion ops as constexpr
...
Will be used for assertions. Also makes registers more flexible
for compile-time stuff in general.
2023-09-17 15:19:44 -04:00
Ryan Houdek
ffb58761c1
Merge pull request #3111 from lioncash/shift
...
Arm64/VectorOps: Fix SVE aliasing-path move in VSShr
2023-09-17 12:15:49 -07:00
Lioncache
8ecdb341e2
Arm64/VectorOps: Fix SVE aliasing-path move in VSShr
...
Seems like this was a typo from 8d11073 , since we'd be moving
into a temporary and then never use it.
2023-09-17 14:43:05 -04:00
Ryan Houdek
0c5c146fcf
FEXCore/JitSymbols: Buffer writes to reduce overhead
...
While this interface is usually pretty fast because it is a write and
forget operation, this has issues when there are multiple threads
hitting the perf map file at the same time. In particular this interface
becomes a bottleneck due to a locking mutex on writes in the kernel.
The situations when this bottleneck occurs is when a bunch of threads
get spawned and they are all jitting code as quickly as possible. In
particular Geekbench's clang benchmark hits this hard where each CPU
thread spends ~40% CPU time on all eight CPU threads because they are
stalled waiting for this mutex to unlock.
To work around this issue, buffer the writes a small amount. Either up
to a page-ish of data or 100ms of time. This completely eliminates
threads waiting on the kernel mutex.
- Around a page of buffer space was chosen by profiling Geekbench's
clang benchmark and seeing how frequently it was still writing.
- 1024 bytes was still fairly aggressive, 4096 seemed fine.
- 100ms was chosen to ensure we don't wait /too/ long to write JIT
symbols.
- In most cases 100ms is enough that you won't notice the blip in
perf.
One thing of note is that with profiling enabled and checking the time
on every JIT block still ends up with 2-3% CPUtime in vdso
clock_gettime. We can improve this by using the cyclecounter directly
since that is still guaranteed to be monotonic. Maybe we'll come back to
that if it is actually an issue here.
2023-09-16 17:52:46 -07:00
Ryan Houdek
ad8b0c673f
Merge pull request #3109 from lioncash/shlx
...
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
2023-09-15 18:49:36 -07:00
Ryan Houdek
e574cfe681
Merge pull request #3108 from lioncash/mulx
...
OpcodeDispatcher: Improve output of MULX
2023-09-15 18:09:02 -07:00
Lioncache
e9be291cec
OpcodeDispatcher: Improve output of SHLX/SHRX/SARX
...
We can remove some unnecessary moves for the 32-bit cases and
collapse the operations down to a single instruction.
2023-09-15 21:05:50 -04:00
Lioncache
d4f87c7db1
OpcodeDispatcher: Improve output of MULX
...
We can cut down on a few of the generated moves. For
the case where both destinations alias one another,
we can just calculate the high part instead of both of them.
2023-09-15 20:52:02 -04:00
Ryan Houdek
4604c01986
Merge pull request #3107 from lioncash/pext
...
Arm64/ALUOps: Remove spills in PEXT
2023-09-15 17:40:44 -07:00
Lioncache
b0c8ff0ea6
Arm64/ALUOps: Remove spills in PEXT
...
Reduces the number of emitted instructions for a
corresponding PEXT instruction.
We no longer spill for this IR op.
2023-09-15 19:39:51 -04:00
Ryan Houdek
647629ac23
Merge pull request #3105 from lioncash/rorx
...
OpcodeDispatcher: Handle RORX corner cases better
2023-09-15 14:55:17 -07:00
Lioncache
be90e76422
Arm64/ALUOps: mov in the case of full 32-bit/64-bit BFE
...
Allows register-renaming mechanisms to be invoked more frequently
2023-09-15 17:38:01 -04:00
Lioncache
4a37ea4819
OpcodeDispatcher: Handle RORX corner cases better
...
There are a few cases where we were emitting code when we
didn't really need to, or could emit less.
2023-09-15 17:36:36 -04:00
Ryan Houdek
6e08ac65b9
Merge pull request #3106 from lioncash/clwb
...
HostFeatures: Fix x86 CLWB support check
2023-09-15 14:06:54 -07:00
Lioncache
8705de1893
HostFeatures: Fix x86 CLWB support check
...
This was clobbering the BMI2 boolean unintentionally.
2023-09-15 16:38:45 -04:00
Alyssa Rosenzweig
c8e7c347c3
Merge pull request #3100 from Sonicadvance1/optimize_cmov
...
OpcodeDispatcher: Optimize cmov
2023-09-15 15:17:51 -04:00
Ryan Houdek
3d0b66407e
Merge pull request #3104 from lioncash/vperm2
...
InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
2023-09-15 11:27:11 -07:00
Lioncache
c86b6dc690
InstCountCI/VEX_map3: Add missing zeroing vperm2f128/vperm2i128 test cases
...
Allows viewing the codegen for cases where conditional zeroing is performed.
Also fixes up the vperm2f variants shorthanding one of the registers
to make everything a little more explicit.
2023-09-15 14:10:11 -04:00
Ryan Houdek
773e9465bc
Merge pull request #3103 from lioncash/warn
...
DeadContextStoreElimination: Silence unused function warning
2023-09-15 10:57:40 -07:00
Lioncache
e1ed7f43fd
DeadContextStoreElimination: Turn LastAccessType into an enum class
...
Makes the type stricter in terms of implicit conversions.
2023-09-15 13:23:00 -04:00
Ryan Houdek
3eb501aa27
InstCountCI: Update for optimized NZCV and cmov
2023-09-15 10:11:50 -07:00
Ryan Houdek
d5b58eebaf
OpcodeDispatcher: Optimize cmov
...
cmov was quite terrible in its implementation. Some things of note:
- NZCV cache would cause store for no reason
- {16,32}-bit would zero extend sources for no reason
- 16-bit would zero extend result for no reason
A bunch of flag testing is still doing a ubfx plus compare against zero
when it could end up being a tst instead, but this is a step in the
right direction and switches over to explicit sized selects.
2023-09-15 10:09:37 -07:00