Commit Graph
1283 Commits
Author SHA1 Message Date
Alyssa Rosenzweig e4ee35a925 JIT: factor out sub reg size conversion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 15:56:05 -04:00
Ryan Houdek 926eefc86c Merge pull request #3635 from alyssarosenzweig/opt/flag-store
OpcodeDispatcher: reorder some moves
2024-05-16 10:58:40 -07:00
Ryan Houdek 3eb7a5b998 Merge pull request #3632 from pmatos/RemoveBlocks
Use erase-remove idiom to remove element
2024-05-16 10:50:50 -07:00
Ryan Houdek 58614ff131 Merge pull request #3634 from alyssarosenzweig/constprop/leftover
ConstProp: remove x86 jit leftover
2024-05-16 10:48:31 -07:00
Alyssa Rosenzweig bf3a09e5e3 ConstProp: remove x86 jit leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 09:21:41 -04:00
Alyssa Rosenzweig 7b39e57e72 OpcodeDispatcher: defer overwritten store
this can save moves, as it's a bit easier to reason about the live ranges.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 08:31:58 -04:00
Paulo Matos 1eb7be9870 Check BitOffset instead of zeroext 2024-05-16 10:50:40 +02:00
Paulo Matos 93e4288c57 Fix left shift undefined behaviour
Here BitOffset can have values higher than 32.
2024-05-16 10:22:03 +02:00
Paulo Matos 9ca4868833 Use erase-remove idiom to remove element
Fixes #3631
2024-05-16 10:16:01 +02:00
Ryan Houdek efe7c54374 Merge pull request #3625 from Sonicadvance1/restricted_inst
FEXCore: Fixes the difference between CPL-0 and undefined instructions
2024-05-14 07:11:19 -07:00
Ryan Houdek 9e1840e974 FEXCore: Moves CodeEmitter to FHU
Now that the vixl dependency is gone, this gets moved to FHU since the
frontend is going to need it for a microjit.
2024-05-13 12:48:10 -07:00
Ryan Houdek 1f40590f9a Emitter: Inline IsImmLogical from vixl
The only core vixl usage we use in the emitter. Is a complete pain to
reimplement so keep it around.
2024-05-13 12:48:10 -07:00
Ryan Houdek 7d9af246ea CodeEmitter: Removes vixl Float16 usage
Creating local Float16 helper which handles our needs
2024-05-13 12:48:09 -07:00
Ryan Houdek 010028e381 FEXCore: Fixes the difference between CPL-0 and undefined instructions
undefined instructions are expected to return SIGILL, while implemented
instructions that aren't available in CPL-3 are expected to SIGSEGV.

Noticed this while testing out CPU-Z, it installs a kernel module and
does a bunch of `RDMSR` and `OUTS` instructions. Decided to walk through
the rest of the instructions in the `System Instruction Reference`
section.

Turns out there's a bunch of oddities in there that we don't support.
First step is to go through all the explicitl SIGILL and SIGSEGV and
implement a test for them.

Next step will be implementing the remaining operations that are
considered "System" operations but are still available in CPL-3.
This list includes:
- lar
- lgdt
- lsl
- sidt
- sldt
- stac
- clac
- verr
- verw
2024-05-13 11:12:26 -07:00
Ryan Houdek d19b57a52e FEXCore: Get rid of DeferredSignalFaultAddress and use the InterruptFaultPage
Arm64ec introduced the InterruptFaultPage which is lower overhead since
instead of ldr+str it just turns in to a single str. We were already
allocating the space, FEXCore and the frontend signal delegator just
needed to be updated to understand the new location.

We can additionally use this in the future if we want to make deferred
async signals INSIDE the JIT only cost a single str as well.
2024-05-10 15:31:28 -07:00
Ryan Houdek 2cae2f2462 Merge pull request #3617 from bylaws/arm64ec-dispatcher
FEXCore: ARM64EC x64 entry/exit support
2024-05-08 12:25:26 -07:00
Ryan Houdek 1fde5d7fca Merge pull request #3621 from alyssarosenzweig/ra/drop-avx
RegisterAllocationPass: drop AVX flag
2024-05-08 11:42:53 -07:00
Ryan Houdek 10de2f83ac Merge pull request #3620 from alyssarosenzweig/ir/burn-parser
IR: drop IRParser
2024-05-08 11:29:37 -07:00
Alyssa Rosenzweig 9d86e11a47 RegisterAllocationPass: drop AVX flag
RA should not depend on whether we support AVX, that's a huge layering
violation! and fortunately, it does not.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:26:31 -04:00
Alyssa Rosenzweig 4d503d3155 RegisterAllocationPass: drop prewritable check
always true.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:24:41 -04:00
Alyssa Rosenzweig 7e663b91df IR: drop IRParser
Aside from its own self-test, the parser is unused and should remain that way,
since it's a maintenance burden with no real benefit. Burn it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:16:54 -04:00
Alyssa Rosenzweig 3c3ba62c10 MemoryOps: optimize 32-bit SRA case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 34fe56dfb2 DeadStoreElimination: CSE block info
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig ecf8cde5e0 DeadStoreElimination: group common logic
slightly less obnoxious copypaste.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 55284aad7e DeadStoreElimination: don't handle partial stores
SRA replaces the whole contents of the destination.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3afc35f7b4 DeadStoreElimination: simplify
use registers internally, not synthesized offsets

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1058428a51 IR: document invariant on SRA
This lets us simplify a lot!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig a2fc51fc7b IR: specify registers, not offsets for SRA
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1848629ba5 RegisterAllocationPass: drop aliasable check
always true with the new ir invariants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3399577330 JIT: clean up fpr sra
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 18bfc8afd0 JIT: clean up gpr sra handling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 76b023ed3e JIT: drop unaligned and partial SRA handling
This is all dead, assert as much so it stays that way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig b91b0e9d65 IR: infer SRA static class
no need to stick it in the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 74489a4177 IR: remove dead SRA flags
I don't know what these were meant for, and I don't care (-:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Billy Laws bd24364c1b FEXCore: Switch stacks before exiting the JIT on ARM64EC
This removes the need for the frontend to have any knowledge of FEX's
SRA layout.
2024-05-06 15:41:34 +00:00
Billy Laws ab516d7b79 Dispatcher: Implement ARM64EC SRA setup entrypoints
While the ARM64EC ABI mostly matches FEX's SRA, the stack still needs to
be switched to the emulator stack and target RIP stored into the FEX
context before jumping to the dispatcher loop.
2024-05-06 15:41:34 +00:00
Billy Laws d25ed4b0bf Dispatcher: Block system call callbacks when compiling code
These callbacks are used for code invalidation and setting the right
emulated CPU features, neither of which are necessary for syscalls made
from within FEX. Avoid calling them to prevent deadlocks caused by
nested locks during compilation.
2024-05-06 15:41:28 +00:00
Ryan Houdek 5f0427c253 CPUID: Adds Qualcomm Oryon product name
From https://github.com/llvm/llvm-project/pull/91022

Easy enough
2024-05-03 20:16:46 -07:00
Ryan Houdek faa494c288 Merge pull request #3605 from Sonicadvance1/move_fex_versionstring_cpuid
CPUID: Removes FEX version string from CPU model name
2024-05-02 11:20:49 -07:00
Ryan Houdek 6228226c08 CPUID: Fix inverted RDTSCP check
This was inverted and always enabling the RDTSCP cpuid bit for wine.
Thus always disabling it elsewhere.
2024-05-01 18:31:41 -07:00
Ryan Houdek 31341bb7c2 CPUID: Removes FEX version string from CPU model name
Moves it to the hypervisor leafs.

Before:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-A78C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
model name      : FEX-2404-101-gf9effcb           Cortex-X1C
```

After:
```bash
$ FEXBash 'cat /proc/cpuinfo | grep "model name"'
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-A78C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
model name      : Cortex-X1C
```

Now the FEX string is in the hypervisor functions as a leaf, so if some
utility wants the FEX version they can query that directly

Ex:
```bash
$ ./Bin/FEXInterpreter get_cpuid_fex
Maximum 4000_0001h sub-leaf: 2
We are running under FEX on host: 2
FEX version string is: 'FEX-2404-113-g820494d'
```
2024-05-01 16:27:13 -07:00
Alyssa Rosenzweig 76b5ca4bcc OpcodeDispatcher: optimize 8/16-bit adc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 20:00:34 -04:00
Alyssa Rosenzweig 28fa88ff39 OpcodeDispatcher: fix 8/16-bit adc/sbc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-29 20:00:34 -04:00
Ryan Houdek 1069cabad0 Merge pull request #3598 from Sonicadvance1/half_barrier_delete_hack
Arm64: Adds another TSO hack to disable half-barrier TSO
2024-04-26 18:24:49 -07:00
Ryan Houdek fe70ec7277 Merge pull request #3599 from alyssarosenzweig/jit/fix-faddv
JIT: fix neon vec4 faddv
2024-04-24 18:20:11 -07:00
Alyssa Rosenzweig 4a4fa64254 JIT: fix neon vec4 faddv
We were previously genrating nonsense code if the destination != source:

         faddp v2.4s, v4.4s, v4.4s
         faddp s2, v4.2s

The result of the first faddp is ignored, so the second merely calculates the
sum of the first 2 sources (not all 4 as needed).

The correct fix is to feed the first add into the second, regardless of the
final destination:

         faddp v2.4s, v4.4s, v4.4s
         faddp s2, v2.2s

Hit in an ASM test with new RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-24 21:13:02 -04:00
Ryan Houdek 6463054fa3 Arm64: Adds another TSO hack to disable half-barrier TSO
A feature of FEX's JIT is that when an unaligned atomic load/store
operation occurs, the instructions will be backpatched in to a barrier
plus a non-atomic memory instruction. This is the half-barrier technique
that still ensures correct visibility of loadstores in an unaligned
context.

The problem with this approach is that the dmb instructions are HEAVY,
because they effectively stop the world until all memory operations in
flight are visible. But it is a necessary evil since unaligned atomics
aren't a thing on ARM processors. FEAT_LSE only gives you unaligned
atomics inside of a 16-byte granularity, which doesn't match x86
behaviour of cacheline size (effectively always 64B).

This adds a new TSO option to disable the half-barrier on unaligned
atomic and instead only convert it to a regular loadstore instruction,
ommiting the half-barrier. This gives more insight in to how well a
CPU's LRCPC implementation is by not stalling on DMB instructions when
possible.

Originally implemented as a test to see if this makes Sonic Adventure 2
run full speed with TSO enabled (but all available TSO options disabled)
on NVIDIA Orin. Unfortunately this basically makes the code no longer
stall on dmb instructions and instead just showing how bad the LRCPC
implementation is, since the stalls show up on `ldapur` instructions
instead.

Tested Sonic Adventure 2 on X13s and it ran at 60FPS there without the
hack anyway.
2024-04-24 13:09:00 -07:00
Ryan Houdek a0bf6a4255 Merge pull request #3595 from alyssarosenzweig/ir/before
Factor out SetWriteCursorBefore
2024-04-23 13:34:05 -07:00
Alyssa Rosenzweig 2372c9458b ConstProp: use SetWriteCursorBefore
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-23 13:09:45 -04:00
Alyssa Rosenzweig 1a11343f34 IREmitter: add SetWriteCursorBefore helper
This is subtle, add an ergonomic helper for it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-04-23 13:09:45 -04:00