Commit Graph
815 Commits
Author SHA1 Message Date
Ryan Houdek 217bbf423b FEXCore: Switch constant emission to default to NoPad
Most constants don't need to be padded for relocations. So now that
these have all been audited, switch to defaulting to NoPad to reduce
verbosity.

The number of constant that need to be explicitly padded are now marked
and with all the prior changes, this allows bisecting if something has
gone wrong.
2025-12-29 11:45:51 -08:00
Ryan Houdek b794b9ed2c Core/Vector: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 3fd86a953b Core/AVX_128: Constant audit 2025-12-29 11:29:02 -08:00
Ryan Houdek 5e782cc1c2 Core/Flags: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek 0ff3fb7f47 Core/X87: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek aa631c5585 Core/X87F64: Constant audit 2025-12-29 11:29:01 -08:00
Ryan Houdek f63ba7e3be OpcodeDispatcher/Vector: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 54dca47e09 OpcodeDispatcher/X87: _Constant audit 2025-12-29 11:04:35 -08:00
Ryan Houdek 3025a10808 FEXCore: Revert literal optimization from #4884
This no longer does anything due to #5123
2025-12-22 10:00:55 -08:00
Billy Laws 2878583627 OpcodeDispatcher: Support relocated operand type variants
In order to support code caching of 32-bit libraries, any library-base
relative relocations on the guest must be transformed into FEX
relocations so e.g. absolute jumps or loads refer to the correct
location when the library is loaded at a different base address.
2025-12-22 16:19:02 +00:00
Ryan Houdek 9f584c8014 SVE256: Fixes AVX scalar round with insert
We were using the incorrect source registers on SVE256 implementation of
these instructions.

Fixes #5100
2025-12-04 01:43:07 -08:00
Paulo Matos 39dbf46422 Refactoring of storing code in x87 opt. stack pass
Enables memcpy optimization of 80bit floats on reduced precision.

Also uncovered a bug where if we had done 80bit memcpy
optimization, we wouldn't have properly stored the 80bits.
This was caught by the existing tests when we enabled the optimization.
2025-11-17 10:14:29 +01:00
Paulo Matos 70b6bc2bae Remove InterpretAsFloat from x87StackOptimizationPass
The InterpretAsFloat was never properly made use of. There's a couple of issues
that are fixed more easily with this gone, so lets remove it.

If there's a specific optimization that requires this, we can bring it back
at a later time. This should not have any effect on the current code generation.
2025-10-28 11:26:28 +01:00
Ryan Houdek 854a741ea4 OpcodeDispatcher: Fix #4982
Forgot to move the OpcodeDispatcher
2025-10-17 14:01:47 -07:00
Ryan Houdek 5f390c16be OpcodeDispatcher: Fixes Scalar FMA size calculation
The frontend did a quirky widening check which was accidentally working
in this case, but it is supposed to be for the couple of GPR handling
AVX instructions.

Correct the implementation to use the correct register size for FMA.
2025-10-09 15:18:26 -07:00
Lioncache c1cfd4db83 Addressing: Shave 8 bytes off AddressMode
Just a trivial rearrangement, Drops the struct down to 40 bytes.
We can also make some functions take it by const reference so we
aren't churning some copies.

Before:
   text     data      bss      dec      hex  filename
4160559  1471360  4336824  9968743   981c67  Bin/FEX

After:
   text     data      bss      dec      hex  filename
4159927  1471360  4336824  9968111   9819ef  Bin/FEX
2025-10-06 15:52:15 -04:00
Lioncache 305b1ecf2d OpcodeDispatcher: Default alignment parameters for store helpers
Avoids actively doing this wonky thing where we're passing
iInvalid all over the place to mean variable alignment depending
on store element size or GPR size.

Makes using the API a little more visibly straightforward and makes
cases where alignment matters more explicit.
2025-10-06 11:04:12 -04:00
Lioncache 16b09a9dd3 OpcodeDispatcher: Move off implicit _Select
Resolves a lingering TODO.
2025-10-05 13:57:36 -04:00
Ryan Houdek 9cef8ff7ce OpcodeDispatcher: Implement support for SSE4a variable extrq/insertq 2025-10-04 02:44:56 -07:00
Ryan Houdek 7c8767ab32 OpcodeDispatcher: Minor improvement to inserting constant to vector 2025-10-04 02:43:21 -07:00
Ryan Houdek f45be1f59e OpcodeDispatcher: Implement support for imm extrq/insertq
Fairly straightforward to implement, but not exciting in their
performance.
2025-10-03 13:05:43 -07:00
Lioncache 0258e914e0 IR: Convert RegisterClassType to an enum class
Removes the last of the wrapper value types in favor of enum classes.
2025-10-03 03:54:56 -04:00
Lioncache fc63dcedb6 IR: Migrate to new helpers
Reduces a bunch of noise related to the register classes and hoists them
out so that converting the classes over to enums should be fairly
straightforward.
2025-10-03 00:53:02 -04:00
Lioncache f76e8c7185 IR: Convert MemOffsetType to enum class
Gets rid of another wrapper struct.
2025-10-02 00:24:42 -04:00
Lioncache 530821fc4a IR: Convert rounding modes to enum class
Same behavior, but more compact
2025-10-01 23:37:57 -04:00
Lioncache a798880ac8 IR: Convert CondClassType over to enum class
Instead of having this sort of odd indirection through a struct type,
we can add support for defining custom enums in the IR description.

This lets us both get strong typing (and allow for weak typing, should
any enum in the future need it), without needing a struct for a basic
value type.

Even then, if we do need a struct for anything in the future, then
we still allow strong typing for values themselves while allowing
them to be used in various ways.
2025-10-01 10:45:00 -04:00
Ryan Houdek 7f8fffbb24 FEXCore: Implement support for AVX gathers with overflow
Falls down the emulated path so we can zero extend the address
calculation. No way for SVE gathers to use 32-bit addressing.
2025-09-29 16:06:02 -07:00
Ryan Houdek efe401c7a4 OpcodeDispatcher: Fixes fxsave x87 register storing
These get stored based on the current rotation of TOP. So we need to be
a bit careful with how we do this storing. A smidge of overhead, but
nothing unexpected.
2025-09-19 14:04:32 -07:00
Ryan Houdek 3fa400bc55 Frontend: Improve DecodeInst size from 128 bytes to 80
We were paying a large cost per Literal type that we can special case
for the two class of instructions that use a 64-bit literal.

If we packed this would get to a further 62 bytes but probably not worth
it.
2025-09-12 16:07:01 -07:00
Ryan Houdek 0eedd55dfd Merge pull request #4866 from neobrain/refactor_reformat
Update code formatting
2025-09-11 10:13:42 -07:00
Ryan Houdek 4a9170e981 Merge pull request #4837 from bylaws/x87opt
X87: Better cache intermediate results in the X87 slowpath
2025-09-11 10:12:36 -07:00
Tony Wasserka 9fdd96af61 Update code formatting 2025-09-11 10:40:29 +02:00
Ryan Houdek 58b5620f97 FEXCore: Fixes OOB access on x87 reduced precision loads
x87 80-bit loads, both BCD and regular tword were loading 128-bits of
data when reduced precision was enabled. This was an oversight from the
previous fix a while ago.

Adds a specific reduced precision test for this, and updates the current
test to ensure stores are still tested as well.
2025-09-10 19:19:55 -07:00
Billy Laws a2286cb00a OpDispatcher: Only use the MM register cache when in MMX state
Since the X87 pass now performs its own MM register caching, these must
be mutually exclusive.
2025-09-09 21:30:33 +01:00
Billy Laws 481787d45c OpDispatcher: Avoid using the register cache for AbridgedFTW
This is now cached internally in the X87 pass, and all operations
outside of that that use it are rare so can afford loading/storing
directly from context after flushing x87 regs.
2025-09-09 21:30:30 +01:00
Paulo Matos 3d92c1fb32 Fix IEEE 754 unordered comparison detection in x87 floating-point operations
Fix SoftFloat IsNan - custom detection matches IEEE754 semantics.
Sets Invalid Operation flags properly for NaN comparisons.

Fixes: GCC-C-execute-ieee-fp-cmp-8l test
__builtin_isunordered() now returns correct values for both NaN and normal operands
2025-09-01 18:26:48 +02:00
Ryan Houdek add54b8089 X86Tables: Convert AVX tables to constexpr
One set of tables for 128-bit and one set of tables for 256-bit.
This one took a bit longer since I needed to convert a few handlers over
to `Bind`. With this all of our x86 tables are costexpr so they end up
in RO mapped memory which is great.
2025-08-27 14:55:10 -07:00
Ryan Houdek 5ed5566106 X86Tables/AVX128: Stop installing state functions a second time 2025-08-27 13:09:53 -07:00
Ryan Houdek 2d9a4412a4 X86Tables: Have AVX128 table always have PCLMUL 2025-08-27 13:08:27 -07:00
Ryan Houdek 982abd5719 X86Tables: Have AVX256 table always have PCLMUL 2025-08-27 13:06:43 -07:00
Ryan Houdek 1a2b2f8870 X86Tables: Move Second ModRM table to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek e8c576047f X86Tables: Moves Base ops to be constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 995d3152eb X86Tables: Convert Secondary tables to constexpr 2025-08-27 12:24:25 -07:00
Ryan Houdek 5ee311b831 X86Tables: Converts H0F3A table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek ea90296e22 X86Tables: Converts H0F38 table to constexpr 2025-08-27 12:24:24 -07:00
Ryan Houdek 87a19c7938 FEXCore: Move SecondaryGroupTables Arch specific ops to a table
No runtime-installation necessary.
2025-08-27 12:24:24 -07:00
Tony Wasserka 51e64c69f3 LogManager: Unconditionally evaluate assertion conditions
A prevalent pattern in the FEX codebase is to compute some data and store it
in a maybe_unused variable that's only ever passed to LOGMAN_THROW_A_FMT.
Besides few exceptions, we never compute expensive data in the macro
arguments themselves, so we can remove a lot of code noise by unconditionally
evaluating the condition even in assertion-disabled builds.
2025-08-25 10:36:01 +02:00
Paulo Matos 3bf323aa2b Clear IE flag on fninit 2025-08-18 15:40:46 +02:00
Ryan Houdek 13da6102b3 OpcodeDispatcher: Implement support for CALLF/RETF
Similar to the previous far jmp, if the CS changes operating mode then
things will still explode with other FEX asserts. But this gets another
change out of my stashes.

These instructions go hand-in-hand obviously so they get implemented as
a pair.
2025-08-13 16:18:12 -07:00
Ryan Houdek 2b7d03d4f9 OpcodeDispatcher: Implement support for far jump
If anything actually attempts to use this to change CS then it'll very
quickly hit some other asserts, but it gets one more change off my list.
2025-08-13 12:19:16 -07:00