Commit Graph
2827 Commits
Author SHA1 Message Date
Ryan Houdek 854a741ea4 OpcodeDispatcher: Fix #4982
Forgot to move the OpcodeDispatcher
2025-10-17 14:01:47 -07:00
Ryan Houdek ed1952a79a Merge pull request #4983 from Sonicadvance1/detect_partial_decode
Frontend: Detect partial decoded instructions
2025-10-17 13:19:27 -07:00
Ryan Houdek 39e8f5122f Merge pull request #4982 from Sonicadvance1/fix_fex_conflict
OpcodeDispatcher: Move FEX reserved instruction
2025-10-17 11:23:09 -07:00
Ryan Houdek eb41cb2261 Frontend: Detect partial decoded instructions
Currently FEX doesn't properly support partial decoded instructions,
which behave slightly differently than full noexec or invalid
instruction decodings. Before this commit we didn't even have a way to
detect the difference.

Primary difference is that the faulting RIP is the beginning of
instruction decode, while the fault address is the first byte that
couldn't be fetched due to memory permissions. This shows up as a
difference between the RIP in mcontext and si_addr in siginfo in the
Linux signal handler.

Right now just change the log so we can determine if we need to support
this edge case.
2025-10-16 13:14:26 -07:00
Ryan Houdek d654f55c3c OpcodeDispatcher: Move FEX reserved instruction
This now conflicts with an SMX instruction, so move it over to another
bytecode that is unlikely to be used.
2025-10-16 11:20:58 -07:00
Paulo Matos eb1689e79f instcountci: Revert Fix quiet and signalling nan propagation 2025-10-16 09:11:01 +02:00
Paulo Matos bc6295a78d Revert "Fix quiet and signalling nan propagation"
This reverts commit e7a47a647c.
2025-10-16 08:53:44 +02:00
Ryan Houdek 40a29ca9a7 FEXCore: Adds option to disable L2 cache lookups
This saves a whole bunch of memory. Cutting `Just Cause 2`'s title
screen from 1132MB anonymous FEX memory down to 438MB. 629MB in L2
alone.

L2 is primarily a means to reduce overhead in map queries, so it's all
about performance. But because it consumes a lot of people it's kind of
hard.

One idea is that the L2 lookups can be moved to shared data structures,
since we already pull the shared lock when doing an L2 lookup this is
already halfway there.

Side note, we're using unique locks even with read-only code paths
which we can't use the shared lock because this terrible recursive
mutex!

Instead of outright changing L2 behaviour and potentially wrecking
havoc, add a config option for now so testing can happen over time.

before:
```
Total FEX Anon memory resident: 1132 mB
    JIT resident:             60 mB
    OpDispatcher resident:    97 mB
    Frontend resident:        37 mB
    CPUBackend resident:      500 kB
    Lookup cache resident:    629 mB
    Lookup L1 cache resident: 108 mB
    ThreadStates resident:    436 kB
```

after:
```
Total FEX Anon memory resident: 438 mB
    JIT resident:             62 mB
    OpDispatcher resident:    56 mB
    Frontend resident:        22 mB
    CPUBackend resident:      496 kB
    Lookup cache resident:    0 (null)
    Lookup L1 cache resident: 109 mB
    ThreadStates resident:    436 kB
```
2025-10-15 13:22:19 -07:00
Ryan Houdek f2841ccb5e FEXCore: Remove the last recursive_mutex
Every time I see this recursive mutex I glare at it. Remove the last one
so that we no longer need to deal with it.

The only reason why this recursive mutex still existed today was because
it is fairly intertwined with the ContextImpl and tracing it all was a
pain.

Peel back the layers and follow the idiom to have ContextImpl pull the
write mutex when requiredand pass it through by reference to ensure it stays alive.
This allows us to entirely give rid of the recursive nature of the
mutex, which means that `FindBlock` can eventually be switched over to a
read-lock to improve multiple threads reading the caches at the same
time.

I didn't do that exercise since that can be followed up in a subsequent
PR.
2025-10-15 08:42:36 -07:00
Ryan Houdek 474f2dc267 FEX: Name remaining allocations as "Misc"
This captures the remaining FEX allocations that /aren't/ coming from
JEMalloc, allowing us to separate our mapped regions versus just
jemalloc allocations.

With some additional naming in jemalloc (which I'm not adding here) this
gets us interesting results:
```
        Misc resident:        54 MiB
    JEMalloc resident:        208 MiB
```

So 208MB of active jemalloc allocations in this particular case. These will be able to be tracked in heaptrack-like applications if careful.
This should let us target down whatever live allocations we're keeping
large amounts of data around if possible.
2025-10-10 17:34:03 -07:00
Ryan Houdek 5f390c16be OpcodeDispatcher: Fixes Scalar FMA size calculation
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.

Correct the implementation to use the correct register size for FMA.
2025-10-09 15:18:26 -07:00
Lioncache c6e60ff3f5 Core: Add missing std::move in AddForceTSOInformation()
All callsites move the instructions into the function, but we weren't
further passing the rvalue-reference to merge().
2025-10-09 12:36:44 -04:00
Lioncache f5d450b95c OpcodeDispatcher: Amend wonky bitwise AND usage in LoadMemPairAutoTSO/_StoreMemPairAutoTSO 2025-10-09 11:39:17 -04:00
Lioncache b98d5f30c0 RegisterAllocationPass: Ensure relevant members are initialized
Ensures that they have deterministic values on construction
2025-10-09 04:17:21 -04:00
LC f8ff46f3e3 Merge pull request #4952 from Sonicadvance1/naming_block_links
FEXCore/fexl: Support a named monotonic_buffer_resource
2025-10-08 23:04:41 -04:00
Ryan Houdek 673e826e46 FEX: Remove InlineSyscall and related flags
FEXCore no longer optimizes syscalls to be inline.

NFC, just avoids passing around a bunch of data structures for no
reason.
2025-10-08 19:33:58 -07:00
Ryan Houdek f1f81f9de2 FEXCore: Remove InlineSyscall
Due to IR changes we can no longer do this, its use was fairly limited
anyway.
2025-10-08 18:47:18 -07:00
Ryan Houdek 282f091e85 FEXCore/fexl: Support a named monotonic_buffer_resource
Lets us track our memory usage of our block links.
2025-10-08 16:43:28 -07:00
Ryan Houdek 8647033029 FEXCore: Support naming a bunch of VMA regions
Useful for memory usage tracking.
2025-10-08 16:43:28 -07:00
Lioncache 7aa5bc0503 EnumUtils: Further simplify enum passthrough formatting
Turns out a simpler way was added to the docs at some point and I never
noticed.

Before:
   text     data      bss      dec      hex  filename
4159895  1471360  4336824  9968079   9819cf  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4157159  1471360  4336824  9965343   980f1f  Bin/FEX
2025-10-07 01:52:05 -04:00
Ryan Houdek 9626a64340 Merge pull request #4944 from lioncash/type
Addressing: Shave 8 bytes off AddressMode
2025-10-06 14:01:09 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Lioncache b7117b86ea X86Tables: Remove unused LateInitCopyTable() 2025-10-06 15:04:36 -04:00
Lioncache 72cb29d36b OpcodeDispatcher: Remove unimplemented function prototypes
Just cleans out the interface a little.
2025-10-06 13:08:07 -04:00
Lioncache a16d4ff1f3 OpcodeDispatcher: Deduplicate in LEAOp/SMSWOp
We can shorten a few lines here by just storing the op addr value
to a local variable.
2025-10-06 11:13:00 -04:00
Lioncache 305b1ecf2d OpcodeDispatcher: Default alignment parameters for store helpers
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.

Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
2025-10-06 11:04:12 -04:00
Lioncache 137aa59254 PrctlUtils: Move to include folder
We can group more prctl value handling in here.
2025-10-06 02:12:29 -04:00
Ryan Houdek 45978474f3 Allocator: Name FEX's VMA regions for allocation
Will allow external tools to track how much memory FEX allocates.
Necessary since we can't use traditional memory usage tools to track FEX
memory allocation independently of guest allocations. Plus most tools
like heaptrack hook allocation symbols, which break under jemalloc.

Using this information I can see with Steam loaded with my library that
FEX consumes ~825MB. Total process resident memory is 1204M, accounting
for around 379MB being used by steam itself. This is /relatively/ close
to my desktop running steam at around 261MB. There's a bit of variance
due to what Steam chooses to do at startup.

This tracking will be the first step towards seeing where our memory
usage is going.
2025-10-05 19:30:59 -07:00
Ryan Houdek 06ff0a45a7 Merge pull request #4935 from lioncash/move
IR: Remove Swap1/Swap2 ops
2025-10-05 11:57:35 -07:00
Ryan Houdek c948d532a7 Merge pull request #4934 from lioncash/select
OpcodeDispatcher: Move off implicit _Select
2025-10-05 11:54:59 -07:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Lioncache 4f6800b768 IR: Remove Swap1/Swap2 ops
These are no longer used.
2025-10-05 13:26:51 -04:00
Lioncache c84801adb3 RedundantFlagCalc: Avoid vector copy in OptimizeParity()
Previously this was making a copy of the vector, when we only
need to read from it.
2025-10-05 12:17:45 -04:00
Lioncache 69f990c1f8 RedundantFlagCalc: Organize headers
Also remove incorrect comment that this pass isn't used.
2025-10-05 12:17:45 -04:00
Lioncache 8a83564678 RedundantFlagCalc: Remove unimplemented prototype
Just tidies the interface a little.
2025-10-05 12:17:45 -04:00
Lioncache 43d93f836f RedundantFlagCalc: Mark members as const where applicable
These don't affect member state.
2025-10-05 03:14:21 -04:00
Ryan Houdek 3baa598b9e HostFeatures: Adds flag for SSE4a 2025-10-04 02:51:28 -07:00
Ryan Houdek 9cef8ff7ce OpcodeDispatcher: Implement support for SSE4a variable extrq/insertq 2025-10-04 02:44:56 -07:00
Ryan Houdek a499ad6404 IR: Implement two new IR ops
rbit is useful in generic algorithms.
`MaskGenerateFromBitWidth` is only really useful for SSE4a, but
implemented in the OpcodeDispatcher is rough, so add an operation.
2025-10-04 02:43:39 -07:00
Ryan Houdek 7c8767ab32 OpcodeDispatcher: Minor improvement to inserting constant to vector 2025-10-04 02:43:21 -07:00
Lioncache 48c6acbc87 IR: Move RegisterAllocationPass forward decl to JITClass
This isn't actually used anywhere in the IR header, so we can
move it to where it's actually used.

Now the IR interface header doesn't have anything related to the
independent passes in it.
2025-10-04 00:08:58 -04:00
Ryan Houdek b5d93bc4a0 Merge pull request #4928 from lioncash/passthru
EnumUtils: Add define for default passthrough formatting
2025-10-03 19:45:41 -07:00
Ryan Houdek f45be1f59e OpcodeDispatcher: Implement support for imm extrq/insertq
Fairly straightforward to implement, but not exciting in their
performance.
2025-10-03 13:05:43 -07:00
Ryan Houdek 012c2ba851 IR: Fixes vector 64-bit binops
We were only supporting operating size of 256-bit and 128-bit. These
also support 64-bit which wasn't wired up.

SSE4a will want to use 64-bit.
2025-10-03 13:04:09 -07:00
Lioncache a0c2ce0068 EnumUtils: Add define for default passthrough formatting
Handles a normal case where printing an enum type as an integral value
is still desirable.

Mainly just a way to reduce boilerplate.
2025-10-03 14:21:19 -04:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Ryan Houdek c35b81d2c5 Merge pull request #4925 from lioncash/regclass
IREmitter: Add register class helpers
2025-10-02 22:14:41 -07:00
Lioncache fc63dcedb6 IR: Migrate to new helpers
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
2025-10-03 00:53:02 -04:00
Lioncache c879650c4b IR: Add GPR/FPR helpers
These will be utilized in follow up PRs
2025-10-02 23:53:50 -04:00
Lioncache 3841cc4aa5 IRDumper: Add remaining missing enum values
Now that the compiler can warn against these, we can fill the
remaining list in.
2025-10-02 12:18:36 -04:00