Commit Graph
56 Commits
Author SHA1 Message Date
Justin Becker 638329278c Inlined PCMPXSTRX
Improves performance for SSE4.2 String operations by ~10x by emitting inline ASM instead of jumping to a C++ helper.

Performance numbers from a simple microbenchmark vs the existing C++ implementation:
|               | Equal Any | Ranges | Equal Each | Equal Ordered |
|---------------|-----------|--------|------------|---------------|
| **pcmpestri** | 13.40×    | 9.71×  | 19.38×     | 14.07×        |
| **pcmpestrm** | 13.12×    | 9.54×  | 18.73×     | 13.80×        |
| **pcmpistri** | 11.75×    | 9.19×  | 17.63×     | 12.80×        |
| **pcmpistrm** | 11.39×    | 8.83×  | 16.31×     | 12.26×        |
2026-09-22 12:50:02 -07:00
Ryan Houdek 021c4fa4bf FEXCore/Allocator: Support naming VMA regions from rpmalloc
This allows WTF to catch the allocations just like on Linux. Punch our
unixlib path all the way through to rpmalloc so it gets named and
tracked properly.
2026-09-07 15:13:33 -07:00
LC 04d06d386f IRDumper: stringstream -> ostringstream
These are purely output operations, so we don't need to use the more
heavyweight class.
2026-07-21 04:29:03 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Ryan Houdek cb432548bf IR: Allow passing padding and MaxBytes through Constant IR 2025-12-29 11:04:34 -08:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Lioncache 3841cc4aa5 IRDumper: Add remaining missing enum values
Now that the compiler can warn against these, we can fill the
remaining list in.
2025-10-02 12:18:36 -04:00
Lioncache 468f2f2d2d IRDumper: Convert remaining printers to lambda style invocations
This allows easily moving the default case out of the switch, making it
easier for compilers to warn about missing values in switches if any
enum members are added in the future but aren't added to the
formatters.
2025-10-02 12:18:28 -04:00
Lioncache be5db91c13 IRDumper: Add missing CheckTF BranchHint 2025-10-02 03:01:11 -04:00
Lioncache 6e7d0a520a IRDumper: Fix error message for FloatCompareOp
If ever printed this would give a misleading error that it was an
unrecognized OpSize type.
2025-10-02 02:57:29 -04:00
Lioncache e37996873a IRDumper: Add missing FenceType entry 2025-10-02 02:55:34 -04:00
Lioncache 25e851b570 IRDumper: Add missing SyscallFlags entry 2025-10-02 02:55:31 -04:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Lioncache 863d1e0007 IR: Convert FenceType to an enum class
Same thing minus an extra struct lingering around.
2025-10-01 23:13:08 -04:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Ryan Houdek 6ae94581bb Merge pull request #4821 from bylaws/monof
Frontend: Fix tailcall handling when mono hacks are enabled
2025-09-06 16:51:08 -07:00
Lioncache dfc42fa7c5 IRDumper: Add missing PrintArg for IndexNamedVectorConstant
Results in better named output for these.
2025-09-04 12:04:48 -04:00
Billy Laws 20fcbb9c62 FEXCore: Avoid potential OOB reads, and flag clobbers in ValidateCode
An instruction could be on the edge of a page and less than 16 bytes
long.
2025-09-02 22:41:59 +01:00
Ryan Houdek 67e9b40bab Merge pull request #4803 from neobrain/refactor_irdumper_cleanup
IRDumper: Clean up formatting using fmt
2025-08-25 11:43:36 -07:00
Tony Wasserka 99cfe05ee5 IRDumper: Clean up formatting using fmt 2025-08-25 17:05:10 +02:00
Tony Wasserka 887f586874 IRDumper: Remove unneeded maybe_unused attribute 2025-08-25 10:36:01 +02:00
Billy Laws 261b7f1110 IR: Support call/ret hints in ExitFunction 2025-07-24 14:53:09 +01:00
Tony Wasserka 9b503f3702 IRDumper: Remove unnecessary use of maybe_unused 2025-07-15 17:10:37 +02:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 49f8332c5b JIT: use registers directly from the IR
This is the flag day change from the series, using all the new shiny
infrastructre we added.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Ryan Houdek 0b519b29d9 IR/IRDumper: static analysis warnings 2025-03-27 04:26:40 -07:00
Tony Wasserka ebb7137839 IRDumper: Allow const RA data 2025-01-31 15:20:23 +01:00
Paulo Matos 44c65c35c8 Revert "Enable RA of SVE Predicate Registers"
This reverts commit fcbf0de05a.

The initial user of this code has been re-implemented in  b148cc6c.
This is not needed any longer so we're removing it.
2025-01-29 11:56:19 +01:00
Paulo Matos 159ed07e68 Print arg type f80Bit 2025-01-10 09:14:49 +01:00
Billy Laws efd6e95059 OpcodeDispatcher: Match x86 overflow behaviour for F2I conversions
ARM behaviour here is to saturate on overflow or NaN inputs, whereas
X86 returns a sentinel value of 2^(bitsize-1), explicitly emulate this.
2024-12-30 00:42:55 +00:00
Paulo Matos fcbf0de05a Enable RA of SVE Predicate Registers 2024-12-02 18:35:31 +01:00
Ryan Houdek 9b6cc8f7e0 IR: Convert OpSize over to enum class
NFC

Do the final mopping up to convert the OpSize enum to an enum class!
2024-10-29 16:52:16 -07:00
Ryan Houdek 82f936cb6d IR: Converts base IR operations to store OpSize sizes
NFC

Finally converts the IR operations themselves to store the OpSize for
the IR operation size and element sizes.

This also finally, FINALLY, converts that remaining `_Constant` helper
to stop using a size field that is specified in bits rather than bytes
like all the other IR op handlers. That thing was so confusing and now
it's gone.
2024-10-28 21:26:59 -07:00
Alyssa Rosenzweig d53e689e22 IR: model tbz/tbnz
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-21 17:52:22 -04:00
Alyssa Rosenzweig 64a45c0d29 IR: remove pairs
They're now unused. And won't be missed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-08-14 09:37:06 -04:00
Ryan Houdek efa05ba19d IR: Adds support for new SUBADD FMA constants
ADDSUB didn't cover this new variant.
2024-06-25 11:22:22 -07:00
Paulo Matos 2b4ec88dae Whole-tree reformat
This follows discussions from #3413.
Followup commits add clang-format file, script and blame ignore lists.
2024-04-12 16:26:02 +02:00
Lioncache 27ba66a181 IRDumper: Extend printer for NamedVectorConstant
Makes it aware of the x87 constants.
2024-04-09 10:13:35 -04:00
Ryan Houdek ed3af580c5 FEXCore: Move nearly all IR definitions to internal
It has been a long time coming that FEX no longer needed to leak IR
implementation details to the frontend, this was legacy due to IR CI and
various other problems.

Now that the last bits of IR leaking has been removed, move everything
that we can internally to the implementation.
We still have a couple of minor details in the exposed IR.h to the
frontend, but these are limited to a few enums and some thunking struct
information rather than all the implementation details.

No functional change with this, just moving headers around.
2024-03-29 17:20:18 -07:00
Ryan Houdek 1b41304fc1 IRDumper: Fixes missing conditional name
When COND_AL was added it wasn't added to this helper. Since there is a
gap between the last condition, just early check the value.

Fixes reading beyond the end of the array
2023-11-29 08:39:35 -08:00
Ryan Houdek 99465faf63 IR: Implements support for subtract with shifted register
Will be used soon.
2023-10-23 09:27:41 -07:00
Ryan Houdek 2671246fef IR: Adds scalar vector insert operations
These IR operations are required to support AFP's NEP mode which does
vector insert in to the destination register. Additionally it gives us
tracking information to allow optimizing out redundant inserts on
devices that don't support AFP natively.

In order to match x86 semantics we need to support binary and unary
scalar operations that do a final insert in to a vector. With optional
zeroing of the top 128-bits for AVX variants.

A tricky thing is that in binary operations this means that the
destination and first source have an intrinsically linked property
depending on if it is SSE or AVX.

SSE example:
- addss xmm0, xmm1
   - xmm0 is both the destination and the first source.
   - This means xmm0[31:0] = xmm0[31:0] + xmm1[31:0]
   - Bits [127:32] are UNMODIFIED.

FEX's JIT jumps through some hoops so that if the destination register
equals the first source register, then it hits the optimal path the
AFP.NEP will insert in to the result. AVX throws a small wrench in to
this due to changed behaviour

AVX example:
- vaddss xmm0, xmm1, xmm2
  - xmm0 is ONLY the destination, xmm1 and xmm2 are the sources
  - This operation copies the bits above the scalar result from the
    first source (xmm1).
  - Additionally this will zero bits above the original 128-bit xmm
    register.
  - xmm0[31:0] = xmm1[31:0] + xmm2[31:0]
  - xmm0[127:32] = xmm1[127:32]
  - ymm0[255:127] = 0

This causes these instructions to support a fairly large table depending
on if the instruction is an SSE or AVX instruction, plus if the host CPU
supports AFP or not.

So while fairly complex, it's handling all the edge cases and gives us
optimization opportunities as we move forward. Currently on non-AFP
supporting devices this has a minor benefit that these IR operations
remove one temporary register, lowering the Register Allocation
overhead.

In the coming weeks I am likely to introduce an optimization pass that
removes redundant inserts because FEX currently does /really/ badly with
scalar code loops.

Needs #3184 merged first.
2023-10-10 03:17:19 -07:00
Alyssa Rosenzweig 2a2619c0f5 IR: Add bit masking selects
Add new synthetic condition codes that do an AND as their relational operator,
testing the result. This is 1 IR op for things like

  (A & B) == 0 ? C : D

This can translate to

  tst A, B
  csel A, B, eq

In the future, if A is the NZCV register and B is a supported immediate, eg

  (NZCV & 0x80000000) == 0 ? C : D

this will be able to translate to a single instruction with the appropriate
condition

  csel A, B, pl

but that needs RA support.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-09-22 19:07:42 -04:00
Ryan Houdek afda4d6b7a FEXCore/Interface/IR: Adds SPDX identifier 2023-09-19 17:33:14 -07:00