Commit Graph
1576 Commits
Author SHA1 Message Date
Ryan Houdek f51832ca6b FEXCore: Fixes a bug with VPSRLDQ/VPSLLDQ with >= 16-byte shifts
When the shift amount is >= 16-bytes then we need to zero the register.
We had a bug where we were assigning `Result.High` to itself, which
effectively made the top 128-bits of the ymm register not modify itself.

Adds a unit test to ensure that doesn't happen again.
2024-09-08 15:39:22 -07:00
Alyssa Rosenzweig 67c751d3e5 JIT: always use 64-bit moves for SRA
I see no reason not to. Simpler code and it's slightly faster on Firestorm.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-08 13:02:57 -04:00
Ryan Houdek 1c59bfeb9f Merge pull request #4027 from alyssarosenzweig/opt/global-flag
Global flag optimizations
2024-09-08 09:42:18 -07:00
Ryan Houdek 59643db331 CPUBackend: Remove unused functions
Some of these were bad ideas, some ideas were something we effectively
grew out of.
2024-09-07 08:07:00 -07:00
Alyssa Rosenzweig d17f33a922 OpcodeDispatcher: don't emit fake 0 for condjump
not needed and getting in the way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig 50a3ca0d6d IR: introduce dedicated PF/AF instructions
this makes reasoning about them a little easier, e.g. for flags.  about 1% win
in nodejs.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 10:59:05 -04:00
Alyssa Rosenzweig 8745455a5b IR: track whether parity is read
so we can gate optimizations efficiently

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig e9ab514962 IR: push parity evaluation down
so we can optimize it globally

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig 4eb0948451 IR: push down AXFLAG lowering
so we can get the new axflag optimizations on billy's x13s.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Alyssa Rosenzweig 8d9f19bd73 IR: add TestZ op
more optimized than TestNZ if we don't care about the sign bit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-07 08:26:04 -04:00
Billy Laws f7b911ca43 OpcodeDispatcher: Do not forbid INT 2E syscalls on 64-bit Windows
This works fine on real Windows and is relied on by wine as SystemCall
is set to 1 in KUSER_SHARED_DATA, which causes the ntdll thunks to use
it over `syscall`
2024-09-06 15:58:17 +00:00
Ryan Houdek a66fac614b Merge pull request #4034 from alyssarosenzweig/fix-tied-fma
IR: fix scalar FMA tied sources
2024-09-04 09:14:24 -07:00
Alyssa Rosenzweig 6d4693cbc1 IR: fix scalar FMA tied sources
needs to be modelled explicitly or else we lose information when translating

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-04 07:42:44 -04:00
Ryan Houdek fd4f6b8020 FEXCore: Dynamically scale TSC
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.

Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.

Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.

Fixes #4026
2024-09-03 13:41:31 -07:00
Alyssa Rosenzweig 74e95df661 Merge pull request #3974 from Sonicadvance1/strict_inprocess_splitlocks
Arm64: Implement support for strict in-process split-locks
2024-09-02 09:26:24 -04:00
Alyssa Rosenzweig 4baeffe84f Merge pull request #4022 from Sonicadvance1/move_sigreturn_to_frontend
FEX: Moves sigreturn symbols to frontend
2024-09-02 09:20:24 -04:00
James Calligeros f588304b12 CPUID: add missing Apple core part numbers
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.

Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
2024-09-01 12:14:39 +10:00
Ryan Houdek c748dbf0e3 FEX: Moves sigreturn symbols to frontend
These are a Linux construct and should live here. Removes a weird
passthrough API from FEXCore and keeps it in the frontend instead.
This isn't even typically allocated in a real setup, as it's only a
fallback for if VDSO isn't loaded.

The CallbackReturn function stays in FEXCore because it would have
caused an API in the other direction instead.
2024-08-31 07:43:04 -07:00
Ryan Houdek 92ddc0041b Merge pull request #4003 from Sonicadvance1/shortcircuit_invalid_inst
Frontend: short-circuit code generation on invalid instructions with multiblock
2024-08-29 04:52:14 -07:00
Ryan Houdek 90f7cc925d Merge pull request #4007 from Sonicadvance1/ensure_no_256bit_operations
Arm64: Ensure 256-bit operations always assert without 256-bit SVE
2024-08-27 20:14:06 -07:00
Alyssa Rosenzweig 335cd9180e OpcodeDispatcher: optimize mul rax
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-27 12:32:43 -04:00
Alyssa Rosenzweig 812224a0ef Merge pull request #4011 from bylaws/arm64-mb
Fix multiblock on ARM64EC
2024-08-27 08:01:26 -04:00
Alyssa Rosenzweig 42f2851575 Merge pull request #4009 from alyssarosenzweig/opt/axflag
Optimize AXFLAG-less systems
2024-08-27 08:01:05 -04:00
Billy Laws fe6dbbfa63 Frontend: Don't explore native ARM64EC jump targets with multiblock 2024-08-26 13:01:06 +00:00
Ryan Houdek b19440c78f Arm64: Ensure 256-bit operations always assert without 256-bit SVE
Our JIT will happily consume incorrectly formed 256-bit vector operations in a lot of cases when the host CPU doesn't support 256-bit SVE.
This is what caused the bug in #4006. For every vector operation that
can consume a 256-bit size, add an assert that always checks if 256-bit
SVE is supported in those cases.

This will ensure that #4006 doesn't happen again.
2024-08-25 18:16:31 -07:00
Alyssa Rosenzweig 8d6b454455 OpcodeDispatcher: optimize AXFLAG emulation
this should help on x13s which has flagm but not flagm2.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:38:03 -04:00
Alyssa Rosenzweig 205ec3e14d OpcodeDispatcher: refactor axflag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:35:27 -04:00
Mai 2478abba29 Merge pull request #4008 from Sonicadvance1/fix_avx128_vfcmp
AVX128: Fixes 256-bit float compares
2024-08-25 12:37:04 -04:00
Ryan Houdek 8bf4a124c8 AVX128: Fixes 256-bit float compares
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.

PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek 7e6ba184f9 AVX128: Fixes incorrect size usage in AVX128_Vector_CVT_Int_To_Float
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".

The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.

Fixes Aperture Desk Job in FEX.
2024-08-25 05:18:53 -07:00
Ryan Houdek c82a683987 Frontend: short-circuit code generation on invalid instructions with multiblock
A source of overhead with multiblock is hitting instructions through a
conditional branch that can never be executed. Usually AVX512
instructions in glibc. This causes us to emit partial blocks for a ton
of targets that will never get executed.

Instead, when we have multiblock enabled, if a block hits an instruction
encoding we don't support, then remove all the decoded instructions from
the block and early terminate it if it isn't the entry block. This
resolves the issue of emitting a bunch of IR and code for blocks never
executed.

If the block of code has an invalid instruction in the entry block for
decoding then it'll still emit code up to the invalid instruction and
raise a SIGILL. This has the potential for generating some additional
blocks of code if a game is abusing SIGILL, but since that's unlikely
it's a good trade-off.

Also removes a few log instructions that don't really provide anything
anymore and just show up as confusing messages when multiblock is
enabled.
2024-08-23 17:22:33 -07:00
Ryan Houdek 54d332935e Arm64: Implement support for strict in-process split-locks
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.

I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.

One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.

This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
2024-08-23 16:18:19 -07:00
Mai 2829ad56a1 Merge pull request #3996 from Sonicadvance1/more_bind
OpcodeDispatcher: Convert more template handlers to Bind handlers
2024-08-23 18:27:59 -04:00
Ryan Houdek fbf62f1296 Merge pull request #3998 from Sonicadvance1/move_midr_fetch
HostFeatures: Moves MIDR querying to the frontend
2024-08-23 15:19:18 -07:00
Billy Laws ef823ce82b OpcodeDispatcher: Allow x86 code to read CNTVCT on ARM64EC
Required by newer insider preview versions, I noticed many crashes with
this exception number and QueryPeformanceCounter in the backtrace,
testing with XTA found it to not be passed through and instead write
the host CNTVCT (unscaled) into RAX. No other registers seem to be
affected.
2024-08-23 21:46:11 +00:00
Ryan Houdek 0ec724cf1f HostFeatures: Moves MIDR querying to the frontend
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.

This lets us keep the Linux specific code in the frontend.
2024-08-23 00:53:53 -07:00
Ryan Houdek 1d00ad6030 OpcodeDispatcher: Convert VectorVariableBlend to Bind handler 2024-08-22 15:09:44 -07:00
Ryan Houdek 23a076c313 OpcodeDispatcher: Convert packed vector shifts to Bind handler 2024-08-22 15:07:09 -07:00
Ryan Houdek 57eacab654 OpcodeDispatcher: Convert VBROADCASTOp to Bind handler 2024-08-22 14:59:25 -07:00
Ryan Houdek b4093a8888 OpcodeDispatcher: Convert packed HSub to Bind handler 2024-08-22 14:57:40 -07:00
Ryan Houdek 7c5a9b5d6a OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler 2024-08-22 14:55:34 -07:00
Ryan Houdek 12b3c82d83 OpcodeDispatcher: Convert PExtr to Bind handler 2024-08-22 14:54:26 -07:00
Ryan Houdek 66520bce0a OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler 2024-08-22 14:52:41 -07:00
Ryan Houdek 25cc2bdcb8 OpcodeDispatcher: Convert VPERMILImm to Bind handler 2024-08-22 14:51:02 -07:00
Ryan Houdek f98b18800c OpcodeDispatcher: Convert MOVMSK to Bind handler 2024-08-22 14:49:57 -07:00
Ryan Houdek 0dd687a7a1 OpcodeDispatcher: Convert SHUFOp to Bind handler 2024-08-22 14:47:30 -07:00
Ryan Houdek 1aff3acbb9 OpcodeDispatcher: Convert PSHUFW to Bind handler 2024-08-22 14:45:04 -07:00
Ryan Houdek a6ab2ca30d OpcodeDispatcher: Convert PUNPCKH to Bind handler 2024-08-22 14:42:24 -07:00
Ryan Houdek ca43e2a61c OpcodeDispatcher: Convert PUNPCKL to Bind handler 2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig ce8e6e5c0c OpcodeDispatcher: optimize adc 0
clang generates this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-22 10:51:06 -04:00