Commit Graph
43 Commits
Author SHA1 Message Date
Ryan Houdek 25ce57ef92 IR: Removes non-opsize AtomicFetchOr 2023-08-28 07:12:31 -07:00
Ryan Houdek 436e0f6f86 IR: Removes non-opsize AtomicFetchXor 2023-08-28 07:12:31 -07:00
Ryan Houdek f44c6f394f IR: Removes non-opsize AtomicFetchNeg 2023-08-28 07:12:31 -07:00
Ryan Houdek 51baf594c9 OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek 16481c0e55 IR: Removes implicit sized rev 2023-08-28 05:04:32 -07:00
Ryan Houdek b00310674a IR: Removes implicit sized not 2023-08-28 05:04:32 -07:00
Ryan Houdek 1f1473eb74 IR: Removes implicit sized neg 2023-08-28 05:04:32 -07:00
Ryan Houdek cccd7001cb IR: Removes implicit sized ashr 2023-08-28 05:04:32 -07:00
Ryan Houdek ce8392d5ae IR: Removes implicit sized ror 2023-08-28 05:02:01 -07:00
Ryan Houdek 386cf36cfd IR: Removes implicit sized sbfe
This one is a bit weird since currently it /always/ assumes a 64-bit
operating size.

We'll likely need to revisit this.
2023-08-28 05:02:01 -07:00
Ryan Houdek 0ddc23a5c9 IR: Removes implicit sized Popcount 2023-08-28 05:02:01 -07:00
Ryan Houdek b95648a4ab IR: Removes implicit sized FindMSB 2023-08-28 05:02:01 -07:00
Ryan Houdek bf18672999 IR: Removes implicit sized FindLSB 2023-08-28 05:02:01 -07:00
Ryan Houdek f405c1be69 IR: Removes inverted EntrypointOffset/InlineEntrypointOffset 2023-08-28 05:02:01 -07:00
Ryan Houdek f3cd115fa7 IR: Removes implicit sized CountLeadingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek a14720e130 IR: Removes implicit sized FindTrailingZeroes 2023-08-28 05:02:00 -07:00
Ryan Houdek 86ed909de8 IR: Removes implicit sized DIV/REM 2023-08-28 05:02:00 -07:00
Ryan Houdek ac7e75c06b IR: Removes implicit sized UDIV/UREM 2023-08-28 05:02:00 -07:00
Ryan Houdek ec3e7ceeb5 IR: Removes implicit sized EXTR 2023-08-28 05:02:00 -07:00
Ryan Houdek c8c8ddbd4f IR: Removes implicit sized PDEP/PEXT 2023-08-28 05:02:00 -07:00
Ryan Houdek ea6d068cc5 IR: Removes implicit sized LDIV/LREM 2023-08-28 05:02:00 -07:00
Ryan Houdek 5013473ec0 IR: Removes implicit sized LUDIV/LUREM 2023-08-28 05:02:00 -07:00
Ryan Houdek 6f2b3e76ac Merge pull request #3013 from Sonicadvance1/32bit_sra
IR/Passes/RA: Enable SRA for 32-bit GPRs
2023-08-27 21:30:39 -07:00
Ryan Houdek e4bb0df486 IR: Convert all Move+Atomic+ALU ops from implicit to explicit size
The number of times the implicit size calculation in GPR operations has
bit us is immeasurable and was a mistake from the start of the project.
The vector based operations never had this problem since they were
explicitly sized for a long time now.

This converts the base IR operations to be explicitly sized, but adds
implicit sized helpers for the moment while we work on removing implicit
usage from the OpcodeDispatcher.

Should be NFC at this moment but it is a big enough change that I want
it in before the "real" work starts.
2023-08-27 01:35:08 -07:00
Ryan Houdek 572d6cd3e6 OpcodeDispatcher: Fixes ADC and SBB 2023-08-26 18:22:50 -07:00
Ryan Houdek fcc37bf6a8 OpcodeDispatcher: Fixes RCR and ADOX 32-bit
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek eace648fa9 OpcodeDispatcher: Fixes bug in UMUL
This was trying to operating on a 32-bit value but BFE the upper
32-bits.

Actually fixes this so it is operating on the 64-bit multiply result.
2023-08-26 18:22:50 -07:00
Ryan Houdek a76c2c57b0 OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek f3679a99ec OpcodeDispatcher: Optimize nontemporal moves
These are now optimal.
2023-08-25 03:13:10 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Lioncache bcba3700c8 OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
Zero-extension will automatically occur if necessary upon storing.

We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache c409ea78bc OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Lioncache d99bcbf01b OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
Zero-extension will already occur if necessary upon storing.

Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Lioncache e5f5629ffc OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Lioncache fa17d9fae9 OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
We already do a zero-extend if necessary in StoreResult.

This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Ryan Houdek a523858f66 Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek 8b051b5e63 OpcodeDispatcher: Implement support for push IR operation
This paves the way to optimizing pushes in to both push operations and
push pair operations to more optimally match Arm64 push support.

While this does the first step for supporting the base push, we'll leave
optimizing push pairs to future work.
2023-08-18 14:19:16 -07:00
Ryan Houdek d19e2507e5 FEXCore: Adds telemetry around legacy segment register setting
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.

Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.

It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.

Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.

A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.

[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
   - 3.6 - Restricted Subset of Segmentation
      - `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
        registers; the base for CS, DS, ES, and SS is ignored for 32-bit
        mode, same as 64-bit mode (treated as zero).`
   - 4.2.17 - MOV to Segment Register
      - Will fault if SS is written (Breaking anything that writes to
        SS).
      - Will not fault if CS, DS, ES are written (Thus it sets the
        segment but gets ignored due to 3.6).
2023-08-17 17:00:41 -07:00
Lioncache 764c844225 OpcodeDispatcher: Improve output of {V}MOVSHDUP 2023-08-17 17:01:50 -04:00
Lioncache 31719aac6a OpcodeDispatcher: Improve output of {V}MOVSLDUP 2023-08-17 17:01:46 -04:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00