Commit Graph
857 Commits
Author SHA1 Message Date
LC 6bc67609a3 [SVE256] Handle 256-bit blend operations much more efficiently
We can massage a given selector into a valid predicate register bitmask
and then simply perform a merging move, which eliminates most busywork
around optimizing 256-bit blends.

In the future, once we drop SVE2.1 support in, we can use PMOV to
eliminate the load from memory and related constant management.
2026-07-08 17:35:12 -04:00
LC 201216ba54 HostFeatures: Put SVE support querying into single function
Lets us avoid open-coding long checks for the existence of either
SVE-128 or SVE-256.
2026-06-30 03:34:17 -04:00
LC ef19242be3 IR: Add constant for swapping midsections of 256-bit vectors around
This'll let us eliminate some pretty gnarly inserts while
ZIP1Q/UZP1Q/etc are still unavailable.
2026-06-29 12:50:53 -04:00
Paulo Matos 37b010795e Re-optimize FYL2X for reduced precision x87 path 2026-06-30 15:00:16 -07:00
Egor Lazarchuk ff5dfff5bb IR: fix typo in CacheLineClean description 2026-06-24 00:38:59 +01:00
Ryan Houdek 280568df2f Merge pull request #5581 from lioncash/fwd
Passes: Trim unnecessary forward declarations
2026-06-21 20:05:23 -07:00
LC a1f90dd8d3 Passes: Trim unnecessary forward declarations
Less visual noise and lingering types left in the header.
2026-06-20 13:59:18 -04:00
LC 2a67261eac IntrusiveIRList: Amend signature for PostRA()
PostRA is a bool, not an unsigned value. We can also adjust SpillSlots()
to use uint32_t like its returned data member.
2026-06-20 13:39:35 -04:00
Paulo Matos 84d968c7d2 JIT-inline FPREM/FPREM1 for reduced precision x87 path 2026-05-11 10:27:08 +02:00
Paulo Matos db50b04663 JIT-inline F64ATAN and F64FYL2X for reduced precision x87 path 2026-04-21 12:09:55 +02:00
Paulo Matos 6bf19ffe37 JIT-inline F64SCALE and F64F2XM1 for reduced precision x87 path 2026-04-16 11:24:19 +02:00
Ryan Houdek 4428aea5ee IR: Adds support for printing strings
I utilize this functionality quite heavily when debugging and I need
bread crumbs spread around. Instead of reimplementing it a dozen times,
just have it upstreamed.
2026-03-28 17:53:33 -07:00
Paulo Matos 30e853305d JIT-inline F64SIN, F64COS, and F64TAN for reduced precision x87 path 2026-03-19 18:58:46 +01:00
Tony Wasserka 91c55facb2 IRDumper: Avoid passing packed member data as references
Fixes "Cannot bind packed field to reference type" build errors on GCC.
2026-03-16 19:15:05 +01:00
Ryan Houdek 63d52c2a1a OpcodeDispatcher: Fixes memory leak in IREmitter pool allocator
While not a leak in the traditional sense, we were causing pool
allocations to never become free until the thread was closed.

This meant in the case of a game running with >200 threads or so, these
would add up very quickly. So some minor reworking so the IREmitter
doesn't allocate a buffer until first JIT, and making sure to actually
disown the buffer on dispatch error resolved the problems.

Fixes an edge case where Ender Lilies was consuming 409MB with THP
enabled on my desktop, and now it is something like 6MB once idling for
a bit to have the pool allocations do its magic.
2026-03-10 20:55:09 -07:00
Ryan Houdek 9d86c3270c IR: Implement new VOrn operation 2026-02-23 14:16:45 -08:00
Tony Wasserka 35bba14b2c IR: Remove malformatted comments from json
Presumably this was done as a hack to highlight the line in red when rendering
documentation to markdown. Since it's only used for two instructions, drop this
use to ease generation of C++ docstrings.
2026-02-23 17:45:05 +01:00
Ryan Houdek 12fcf96e93 IR: Adds a ContextClear operation
This will be useful for zeroing parts of the context at CLZero alignment
and sizes.
2026-02-21 15:24:13 -08:00
Ryan Houdek 217bbf423b FEXCore: Switch constant emission to default to NoPad
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.

The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
2025-12-29 11:45:51 -08:00
Ryan Houdek 5bcdb3d478 IREmitter: Remove Pad default argument
Default argument is no longer used
2025-12-29 11:29:02 -08:00
Ryan Houdek 8269d04b57 IR/IREmitter: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek a1a30cd9a6 IR: Remove default argument for padding in Constant op
All use cases now pass a pad type in to this.
2025-12-29 11:29:01 -08:00
Ryan Houdek a480793708 Passes/RegisterAllocationPass: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 8588c22170 IREmitter: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek c2d5ee43c6 Passes/x87StackOptimizationPass: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek c7c6855740 IREmitter: Allow the constant pool to understand padtype and bytes
Fixes an issue that a potential same constant could be padded for one
use and not padded for another.
2025-12-29 11:04:34 -08:00
Ryan Houdek cb432548bf IR: Allow passing padding and MaxBytes through Constant IR 2025-12-29 11:04:34 -08:00
Ryan Houdek 90c8fcf393 Merge pull request #5094 from pmatos/fix/issue5084
Set current code block in x87 pass
2025-12-02 14:43:54 -08:00
Paulo Matos f91ac09f87 Set current code block in x87 pass
This resets the constant pool in IREmit used by SelectAddressMode().

Fixes #5084.
2025-12-02 15:33:08 +01:00
Tony Wasserka 096c408ef6 IR: Replace hand-written operators with three-way comparison 2025-12-01 17:10:06 +01:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Anton Kesy 013ac1e627 Align code with clang-format
Automatically done by running:
`find . \( -path './External' -prune \) -o \
  \( -iname '*.cc' -o -iname '*.cpp' -o -iname '*.hpp' -o \
     -iname '*.h' -o -iname '*.c' \) -print | \
  xargs clang-format --style=file -i`
2025-11-13 13:29:49 +01:00
Ryan Houdek ff25e9a92e FEXCore: Remove usage of "remote atomic" xor
This is the only usage of LSE atomics that isn't the fetch variety.
[This article](https://www.phoronix.com/news/Linux-6.18-ARM64-Atomics-Issue)
reminded me that this was a thing and that I should double check the IR.
This was the only IR operation remaining that still didn't use the fetch
variety. Convert it over to the fetch to avoid the expectation that it
can be a "remote atomic". Change is going to fall in to noise, but might
as well as be consistent.
2025-11-11 17:16:03 -08:00
Paulo Matos 6ee9984280 f80 stack xchg optimization for fast path 2025-11-05 11:20:02 +01:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
Ryan Houdek 81fc502c6c FEXCore: Remove Paranoid TSO mode.
This mode has been broken for a long time because it's mostly untested.
Barriers, and backpatching while slow have proven that they work.
Maintain the one TSO path, at least until all ARM hardware gains support for
x86-TSO memory model mode.
2025-10-21 10:53:34 -07:00
Paulo Matos eb1689e79f instcountci: Revert Fix quiet and signalling nan propagation 2025-10-16 09:11:01 +02:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Lioncache b98d5f30c0 RegisterAllocationPass: Ensure relevant members are initialized
Ensures that they have deterministic values on construction
2025-10-09 04:17:21 -04:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Ryan Houdek 06ff0a45a7 Merge pull request #4935 from lioncash/move
IR: Remove Swap1/Swap2 ops
2025-10-05 11:57:35 -07:00
Ryan Houdek c948d532a7 Merge pull request #4934 from lioncash/select
OpcodeDispatcher: Move off implicit _Select
2025-10-05 11:54:59 -07:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Lioncache c84801adb3 RedundantFlagCalc: Avoid vector copy in OptimizeParity()
Previously this was making a copy of the vector, when we only
need to read from it.
2025-10-05 12:17:45 -04:00
Lioncache 69f990c1f8 RedundantFlagCalc: Organize headers
Also remove incorrect comment that this pass isn't used.
2025-10-05 12:17:45 -04:00
Lioncache 8a83564678 RedundantFlagCalc: Remove unimplemented prototype
Just tidies the interface a little.
2025-10-05 12:17:45 -04:00