Commit Graph
1764 Commits
Author SHA1 Message Date
Ryan Houdek a66fac614b Merge pull request #4034 from alyssarosenzweig/fix-tied-fma
IR: fix scalar FMA tied sources
2024-09-04 09:14:24 -07:00
Alyssa Rosenzweig 6d4693cbc1 IR: fix scalar FMA tied sources
needs to be modelled explicitly or else we lose information when translating

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-09-04 07:42:44 -04:00
Ryan Houdek fd4f6b8020 FEXCore: Dynamically scale TSC
When I implemented TSC scaling originally, I chose a scale factor of 128
because it basically covered the range of devices we cared about without
going too high. I also only tested devices that had a TSC scale factor
from 19.2Mhz to 34Mhz. Turns out there is hardware that also has a 48Mhz
cycle counter, which cause them to effectively have a 6.1Ghz cycle
counter, which is kind of absurd.

Instead of a fixed scale, just calculate the amount of scaling we need
to get >= the minimum threshold of 1Ghz. This will change the shift from
7 to 5 or 6 for the faster cycle counter devices.

Of course if someone wants to know the scale factor they can still use
cpuid function 15h to know it.

Fixes #4026
2024-09-03 13:41:31 -07:00
Ryan Houdek ac32876e4e LinuxEmulation: Implement support for seccomp
Seccomp is a relatively complex feature that was added to Linux back in
2005, and was further extended in 2013 to support BPF based protections.
Once seccomp is enabled, you can no longer disable seccomp but
additional protections can be placed on top of existing seccomp filters.
Additionally seccomp filters are inherited in child processes, which
ensures the process tree can't escape from the secure computing
environment through child processes.

The basis of this feature is a shim that lives between userspace and the
kernel at the syscall entrypoint.
In "strict" mode, seccomp only allows read, write, exit, exit_group, and {rt_,}sigreturn to function.
When in "filter" mode, a BPF filter is run on syscall entrypoint and
returns state about if the syscall should be allowed or not. Multiple
filters can be installed in this mode, all of which get executed. The
result that is the most restricted is the action that occurs at the end.

There are some significant limitations in filter mode that must be
adhered to which makes executing this code inside of kernel space a
non-issue and effectively limits how much cpu time is spent in the filters.
Although these filters are free to do basically anything with the
provided data, just can't do any loops.

FEX needs to implement seccomp because there are multiple applications
using the feature, the primary one being Chromium which some games embed
without disabling the sandbox. WINE also uses seccomp for capturing
games that do raw Windows system calls. Apparently Red Dead Redemption
is one of the games that requires this.

While FEX implements seccomp, it is not yet all encompassing, which is
one of the reasons why it isn't enabled by default and requires a config
option.

**seccomp_unotify is not implemented**
This is a relatively new feature for seccomp which lets the seccomp
filter signal an FD for multiple things. Luckily Chromium and WINE don't
use this. This will be tricky to implement under FEX since it
requires ioctl trapping and some other behaviour

**ptrace isn't supported**
One feature of seccomp is that it can raise ptrace events. Since FEX
doesn't support ptrace at all, this isn't handled. Again Chromium and
WINE don't use this.

**kill-thread not quite correct**
This isn't directly related to seccomp but more about how we do thread
shutdown in FEX. This will require some more changes around thread state
tracking before fully supporting this. Chromium and WINE don't use this.
kill-process also falls under this

Features that are supported:
- Strict mode and seccomp-bpf mode supported
- All BFP instructions that seccomp-bpf understands
- Inheriting seccomp through execve
   - This means we serialize and deserialize the calling thread's
     seccomp filters
   - An execve that escapes FEX will also escape seccomp. Not much we
     can do about it
- TSync - Allowing post-mortem seccomp insertion which allows threads to
  synchronize seccomp filters after the fact

Features that are not supported:
- Different arch qualifiers depending on syscall entrypoint
  - Just like our syscall handler, we are hardcoded to the arch that the
    application starts with
- user_notif
- ptrace
- Runtime code cache invalidation when seccomp is installed
  - Currently we must ensure all syscalls go through the frontend
    syscall handler
  - Runtime invalidation of code cache with inline syscalls will get
    fixed in the future.

This currently isn't enabled by default because of the minor feature
problems that haven't been resolved. Currently the Linux Kernel's test
application works for the features that FEX supports, and WINE's usage
can be handled by FEX. Chromium's sandbox doesn't yet work with this PR,
but it only fails due to features unrelated to seccomp.

Having this open for merging now so we can work to resolve the remaining
issues without this bitrotting.
2024-09-02 14:07:53 -07:00
Alyssa Rosenzweig 74e95df661 Merge pull request #3974 from Sonicadvance1/strict_inprocess_splitlocks
Arm64: Implement support for strict in-process split-locks
2024-09-02 09:26:24 -04:00
Alyssa Rosenzweig 4baeffe84f Merge pull request #4022 from Sonicadvance1/move_sigreturn_to_frontend
FEX: Moves sigreturn symbols to frontend
2024-09-02 09:20:24 -04:00
Ryan Houdek b4a67a6178 Merge pull request #4021 from Sonicadvance1/remove_global_initializer_thunks
Thunks: Removes global static initializer
2024-08-31 22:54:22 -07:00
James Calligeros f588304b12 CPUID: add missing Apple core part numbers
The Ultra-class SoCs are two Max-class SoCs connected via
Apple's fabric, and thus use the same core revisions as the
Max-class SoCs for both big and LITTLE cores.

Signed-off-by: James Calligeros <jcalligeros99@gmail.com>
2024-09-01 12:14:39 +10:00
Ryan Houdek c748dbf0e3 FEX: Moves sigreturn symbols to frontend
These are a Linux construct and should live here. Removes a weird
passthrough API from FEXCore and keeps it in the frontend instead.
This isn't even typically allocated in a real setup, as it's only a
fallback for if VDSO isn't loaded.

The CallbackReturn function stays in FEXCore because it would have
caused an API in the other direction instead.
2024-08-31 07:43:04 -07:00
Ryan Houdek e2a7fef742 Thunks: Removes global static initializer
This variable can't be constexpr initialized since it requires linker
fix-ups, which changes it in to a global static initializer instead.

While this benign as it doesn't allocate any memory, just move it next
to its single use. Removes the static initializer that was there for no
reason.
2024-08-31 01:58:35 -07:00
Ryan Houdek 92ddc0041b Merge pull request #4003 from Sonicadvance1/shortcircuit_invalid_inst
Frontend: short-circuit code generation on invalid instructions with multiblock
2024-08-29 04:52:14 -07:00
Ryan Houdek 90f7cc925d Merge pull request #4007 from Sonicadvance1/ensure_no_256bit_operations
Arm64: Ensure 256-bit operations always assert without 256-bit SVE
2024-08-27 20:14:06 -07:00
Alyssa Rosenzweig 335cd9180e OpcodeDispatcher: optimize mul rax
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-27 12:32:43 -04:00
Alyssa Rosenzweig 03832b2523 Merge pull request #4012 from alyssarosenzweig/opt/pf-af-kill
DeadStoreElimination: handle PF/AF invalidate
2024-08-27 08:01:54 -04:00
Alyssa Rosenzweig 812224a0ef Merge pull request #4011 from bylaws/arm64-mb
Fix multiblock on ARM64EC
2024-08-27 08:01:26 -04:00
Alyssa Rosenzweig 42f2851575 Merge pull request #4009 from alyssarosenzweig/opt/axflag
Optimize AXFLAG-less systems
2024-08-27 08:01:05 -04:00
Alyssa Rosenzweig faf1b85904 DeadStoreElimination: handle PF/AF invalidate
lets us kill some AF calculations with multiblock.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-26 11:28:16 -04:00
Billy Laws bec5f4fe2a DeadStoreElimination: Fix uint64 typing for register bitmask
unsigned longs are 32 bits wide on Windows.
2024-08-26 13:12:58 +00:00
Billy Laws fe6dbbfa63 Frontend: Don't explore native ARM64EC jump targets with multiblock 2024-08-26 13:01:06 +00:00
Ryan Houdek b19440c78f Arm64: Ensure 256-bit operations always assert without 256-bit SVE
Our JIT will happily consume incorrectly formed 256-bit vector operations in a lot of cases when the host CPU doesn't support 256-bit SVE.
This is what caused the bug in #4006. For every vector operation that
can consume a 256-bit size, add an assert that always checks if 256-bit
SVE is supported in those cases.

This will ensure that #4006 doesn't happen again.
2024-08-25 18:16:31 -07:00
Alyssa Rosenzweig 8d6b454455 OpcodeDispatcher: optimize AXFLAG emulation
this should help on x13s which has flagm but not flagm2.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:38:03 -04:00
Alyssa Rosenzweig 205ec3e14d OpcodeDispatcher: refactor axflag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-25 18:35:27 -04:00
Mai 2478abba29 Merge pull request #4008 from Sonicadvance1/fix_avx128_vfcmp
AVX128: Fixes 256-bit float compares
2024-08-25 12:37:04 -04:00
Ryan Houdek 8bf4a124c8 AVX128: Fixes 256-bit float compares
Just like the bug in #4006, we were incorrectly passing a 256-bit
comparison in to the float compare IR operations. Once again it would
"safely" decompose in to 128-bit operations without issue.

PR #4007 added asserts to ensure that 256-bit operations aren't emitted
when the host CPU doesn't support 256-bit SVE and detected this.
2024-08-25 06:33:31 -07:00
Ryan Houdek 7e6ba184f9 AVX128: Fixes incorrect size usage in AVX128_Vector_CVT_Int_To_Float
This handler was incorrectly using 256-bit IR operation sizes. Due to a
quirk with our IR handling, this would "safely" fall back to a 128-bit
operation and work "correctly".

The problem encountered is that since the IR operation is claiming to be
256-bit, when the value got spilled due to register pressure then a
true 256-bit store and load operation would be generated. This would
then emit an SVE load and store, with the expectation of 256-bit SVE
loadstores. This caused a SIGILL on Oryon since it doesn't support SVE,
but even would generate an invalid predicated loadstore on SVE 128-bit
hardware.

Fixes Aperture Desk Job in FEX.
2024-08-25 05:18:53 -07:00
Ryan Houdek c82a683987 Frontend: short-circuit code generation on invalid instructions with multiblock
A source of overhead with multiblock is hitting instructions through a
conditional branch that can never be executed. Usually AVX512
instructions in glibc. This causes us to emit partial blocks for a ton
of targets that will never get executed.

Instead, when we have multiblock enabled, if a block hits an instruction
encoding we don't support, then remove all the decoded instructions from
the block and early terminate it if it isn't the entry block. This
resolves the issue of emitting a bunch of IR and code for blocks never
executed.

If the block of code has an invalid instruction in the entry block for
decoding then it'll still emit code up to the invalid instruction and
raise a SIGILL. This has the potential for generating some additional
blocks of code if a game is abusing SIGILL, but since that's unlikely
it's a good trade-off.

Also removes a few log instructions that don't really provide anything
anymore and just show up as confusing messages when multiblock is
enabled.
2024-08-23 17:22:33 -07:00
Ryan Houdek 54d332935e Arm64: Implement support for strict in-process split-locks
For atomics that cross the 16-byte or 64-byte granularity, we need to
lock a mutex to ensure strict emulation of split-locks.

I took another look at these when I found out that Zen3 actually
implements split-locks. Not sure which architecture actually added
support for from them, but I wanted to ensure we have the ability to
handle this.

One thing that we can't handle in user-space is cross-process
split-locks through shared memory. This requires a kernel SIGBUS handler
to ensure a crashing/SIGKILL'd process doesn't lock all FEX processes in
the system.

This fixes a little split-lock abusing test that I have locally. It's a
bit flakey so it isn't viable to run in CI. Considering it is explicitly
testing a race problem.
2024-08-23 16:18:19 -07:00
Mai 2829ad56a1 Merge pull request #3996 from Sonicadvance1/more_bind
OpcodeDispatcher: Convert more template handlers to Bind handlers
2024-08-23 18:27:59 -04:00
Ryan Houdek fbf62f1296 Merge pull request #3998 from Sonicadvance1/move_midr_fetch
HostFeatures: Moves MIDR querying to the frontend
2024-08-23 15:19:18 -07:00
Billy Laws ef823ce82b OpcodeDispatcher: Allow x86 code to read CNTVCT on ARM64EC
Required by newer insider preview versions, I noticed many crashes with
this exception number and QueryPeformanceCounter in the backtrace,
testing with XTA found it to not be passed through and instead write
the host CNTVCT (unscaled) into RAX. No other registers seem to be
affected.
2024-08-23 21:46:11 +00:00
Ryan Houdek 0ec724cf1f HostFeatures: Moves MIDR querying to the frontend
The MIDR querying is inherently OS specific and needs a bit of special
casing. Instead let the frontend inform FEXCore how many CPU cores there
are and their MIDRs instead.

This lets us keep the Linux specific code in the frontend.
2024-08-23 00:53:53 -07:00
Ryan Houdek 1d00ad6030 OpcodeDispatcher: Convert VectorVariableBlend to Bind handler 2024-08-22 15:09:44 -07:00
Ryan Houdek 23a076c313 OpcodeDispatcher: Convert packed vector shifts to Bind handler 2024-08-22 15:07:09 -07:00
Ryan Houdek 57eacab654 OpcodeDispatcher: Convert VBROADCASTOp to Bind handler 2024-08-22 14:59:25 -07:00
Ryan Houdek b4093a8888 OpcodeDispatcher: Convert packed HSub to Bind handler 2024-08-22 14:57:40 -07:00
Ryan Houdek 7c5a9b5d6a OpcodeDispatcher: Convert AVXVectorVariableBlend to Bind handler 2024-08-22 14:55:34 -07:00
Ryan Houdek 12b3c82d83 OpcodeDispatcher: Convert PExtr to Bind handler 2024-08-22 14:54:26 -07:00
Ryan Houdek 66520bce0a OpcodeDispatcher: Convert VPACK{U,S}S to Bind handler 2024-08-22 14:52:41 -07:00
Ryan Houdek 25cc2bdcb8 OpcodeDispatcher: Convert VPERMILImm to Bind handler 2024-08-22 14:51:02 -07:00
Ryan Houdek f98b18800c OpcodeDispatcher: Convert MOVMSK to Bind handler 2024-08-22 14:49:57 -07:00
Ryan Houdek 0dd687a7a1 OpcodeDispatcher: Convert SHUFOp to Bind handler 2024-08-22 14:47:30 -07:00
Ryan Houdek 1aff3acbb9 OpcodeDispatcher: Convert PSHUFW to Bind handler 2024-08-22 14:45:04 -07:00
Ryan Houdek a6ab2ca30d OpcodeDispatcher: Convert PUNPCKH to Bind handler 2024-08-22 14:42:24 -07:00
Ryan Houdek ca43e2a61c OpcodeDispatcher: Convert PUNPCKL to Bind handler 2024-08-22 14:39:45 -07:00
Alyssa Rosenzweig ce8e6e5c0c OpcodeDispatcher: optimize adc 0
clang generates this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-22 10:51:06 -04:00
Alyssa Rosenzweig 6951284924 OpcodeDispatcher: optimize ADCX/ADOX
rewrite to use native adcs with flag fixups.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig b03c613fbf OpcodeDispatcher: optimize JP/JNP
fuse and+cbnz into tbz/tbnz.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig b893bdb8df OpcodeDispatcher: optimize fcmovu
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig db45f6eec8 OpcodeDispatcher: optimize test with small immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig d53e689e22 IR: model tbz/tbnz
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00