When the shift amount is >= 16-bytes then we need to zero the register.
We had a bug where we were assigning `Result.High` to itself, which
effectively made the top 128-bits of the ymm register not modify itself.
Adds a unit test to ensure that doesn't happen again.
This works fine on real Windows and is relied on by wine as SystemCall
is set to 1 in KUSER_SHARED_DATA, which causes the ntdll thunks to use
it over `syscall`
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.
Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.
Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.
Fixes#4026
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.
Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
These are a Linux construct and should live here. Removes a weird
passthrough API from FEXCore and keeps it in the frontend instead.
This isn't even typically allocated in a real setup, as it's only a
fallback for if VDSO isn't loaded.
The CallbackReturn function stays in FEXCore because it would have
caused an API in the other direction instead.
Our JIT will happily consume incorrectly formed 256-bit vector operations in a lot of cases when the host CPU doesn't support 256-bit SVE.
This is what caused the bug in #4006. For every vector operation that
can consume a 256-bit size, add an assert that always checks if 256-bit
SVE is supported in those cases.
This will ensure that #4006 doesn't happen again.
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.
PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".
The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.
Fixes Aperture Desk Job in FEX.
A source of overhead with multiblock is hitting instructions through a
conditional branch that can never be executed. Usually AVX512
instructions in glibc. This causes us to emit partial blocks for a ton
of targets that will never get executed.
Instead, when we have multiblock enabled, if a block hits an instruction
encoding we don't support, then remove all the decoded instructions from
the block and early terminate it if it isn't the entry block. This
resolves the issue of emitting a bunch of IR and code for blocks never
executed.
If the block of code has an invalid instruction in the entry block for
decoding then it'll still emit code up to the invalid instruction and
raise a SIGILL. This has the potential for generating some additional
blocks of code if a game is abusing SIGILL, but since that's unlikely
it's a good trade-off.
Also removes a few log instructions that don't really provide anything
anymore and just show up as confusing messages when multiblock is
enabled.
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.
I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.
One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.
This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
Required by newer insider preview versions, I noticed many crashes with
this exception number and QueryPeformanceCounter in the backtrace,
testing with XTA found it to not be passed through and instead write
the host CNTVCT (unscaled) into RAX. No other registers seem to be
affected.
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.
This lets us keep the Linux specific code in the frontend.