Commit Graph
28 Commits
Author SHA1 Message Date
Ryan Houdek 51baf594c9 OpcodeDispatchers: Adds OpSizeFromSrc/Dst helpers
To reduce how cluttered `IR::SizeToOpSize(GetSrcSize(Op))` is.
2023-08-28 05:11:32 -07:00
Ryan Houdek 48669b7006 IR: Removes implicit sized orlshl/orlshr 2023-08-28 05:04:32 -07:00
Ryan Houdek cccd7001cb IR: Removes implicit sized ashr 2023-08-28 05:04:32 -07:00
Ryan Houdek f405c1be69 IR: Removes inverted EntrypointOffset/InlineEntrypointOffset 2023-08-28 05:02:01 -07:00
Ryan Houdek fcc37bf6a8 OpcodeDispatcher: Fixes RCR and ADOX 32-bit
Automatic size inheritance was breaking these operations.
2023-08-26 18:22:50 -07:00
Ryan Houdek a4ac21a4e4 OpcodeDispatcher: Fixes bug in GetRFLAG with CachedNZCV
This was only operating at byte size but it was attempting to get bit
offsets at greater than operating size.
Change operating size over to 64-bit.
2023-08-26 18:22:50 -07:00
Ryan Houdek a76c2c57b0 OpcodeDispatcher: Optimize PSHUF{LW, HW, D}!
This is way more optimal!
2023-08-25 12:59:40 -07:00
Ryan Houdek c441b238c7 OpcodeDispatcher: Optimize MMX conversion operation
These instructions are now optimal
2023-08-24 15:46:19 -07:00
Ryan Houdek 565b30e15e OpcodeDispatcher: Cache named vector constants in the block
If the named constant of that size gets used multiple times then just
use the previous value if it was in scope.

Makes addsubp{s,d} and phminposuw more optimal for each that are in a
block.

Needs #2993 merged first.
2023-08-24 14:46:37 -07:00
Lioncache 26c81224ac OpcodeDispatcher: Remove redundant moves from AESIMC
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 21:34:37 -04:00
Lioncache 52ab3f6a1e OpcodeDispatcher: Remove redundant moves from VAESKeyGenAssist
Zero-extension will occur upon storing if necessary.

We can also join the AVX implementation with the SSE one.
2023-08-23 21:22:00 -04:00
Mai 4b06069c0d Merge pull request #2972 from Sonicadvance1/optimize_scalar_mov
OpcodeDispatcher: Optimizes scalar movd/movq
2023-08-23 16:30:45 -04:00
Ryan Houdek 8836ab8988 OpcodeDispatcher: Optimizes scalar movd/movq
MMX and SSE versions are now optimal.
2023-08-23 12:56:11 -07:00
Ryan Houdek 86ef6fe48d Merge pull request #2976 from lioncash/mov
OpcodeDispatcher: Remove unnecessary moves from AVX move ops where applicable
2023-08-23 12:02:05 -07:00
Lioncache bcba3700c8 OpcodeDispatcher: Remove redundant moves from VMOVVectorNTOp
Zero-extension will automatically occur if necessary upon storing.

We can also join the SSE and AVX implementations.
2023-08-23 14:17:27 -04:00
Lioncache c409ea78bc OpcodeDispatcher: Remove unnecessary moves from VMOV{A,U}PS/VMOV{A,U}PD
Zero-extension will occur automatically upon storing if necessary.

We can also join the SSE and AVX implementations together.
2023-08-23 14:10:49 -04:00
Ryan Houdek a62ba75ede Merge pull request #2975 from lioncash/scalar
Arm64/ConversionOps: Add scalar support to Vector_FToI
2023-08-23 10:57:38 -07:00
Lioncache 990b70dcd6 OpcodeDispatcher: Use scalar rounding for scalar round instructions 2023-08-23 13:34:01 -04:00
Lioncache d99bcbf01b OpcodeDispatcher: Remove unnecessary moves from AVXExtendVectorElements
Zero-extension will already occur if necessary upon storing.

Also we can join the AVX and SSE implementations together and get
rid of some template instantiations, now that the only differing
behavior is removed.
2023-08-23 12:47:31 -04:00
Lioncache e5f5629ffc OpcodeDispatcher: Remove unnecessary moves from AVX conversion operations
These zero-extensions will already happen automatically if necessary.
2023-08-22 22:20:13 -04:00
Lioncache fa17d9fae9 OpcodeDispatcher: Remove unnecessary move from VPHMINPOSUW
We already do a zero-extend if necessary in StoreResult.

This also lets us unify both the SSE and AVX handling code.
2023-08-22 20:59:56 -04:00
Ryan Houdek a523858f66 Merge pull request #2923 from Sonicadvance1/nonnull_legacy_segment_telemetry
FEXCore: Adds telemetry around legacy segment register setting
2023-08-20 10:27:56 -07:00
Ryan Houdek 92c3014aaa OpcodeDispatcher: Minor optimization around clearing flags
When clearing multiple flags it is more optimal to load the mask
constant in to a register and then clear with a single and/bic.

Back to back bfi is actually less optimal due to dependency tracking.

With #2911, this is a total win since this hits an edge case with
constant loading that #2911 fixes.
2023-08-19 20:14:37 -07:00
Lioncache c31329609f OpcodeDispatcher: Unify handling code for MOVSD and MOVSS
These have the same behavior and only differ based on element size,
so we can join the implementations together instead of duplicating
them across both functions.
2023-08-19 17:44:45 -04:00
Ryan Houdek d19e2507e5 FEXCore: Adds telemetry around legacy segment register setting
Due to Intel dropping support for legacy segment registers[1] there is a
concern that this will break legacy 32-bit software that is doing some
magic segment register handling.

Adds some simple telemetry for 32-bit applications that when they
encounter an instruction that sets the segment register or uses a
segment register that the JIT will do a /relatively/ quick four
instruction check to see if it is not a null segment.

It's not enough to just check if the segment index is 0 or not, 32-bit
Linux software starts with non-zero segment register indexes but the LDT
for each segment index is a null-descriptor.

Once the segment address is loaded, the IR operation will do a quick
check against zero and if it /isn't/ zero then set the telemetry value.

A very minor optimization that segment registers only get checked once
per block to ensure overhead stays low.

[1] https://www.intel.com/content/www/us/en/developer/articles/technical/envisioning-future-simplified-architecture.html
   - 3.6 - Restricted Subset of Segmentation
      - `Bases are supported for FS, GS, GDT, IDT, LDT, and TSS
        registers; the base for CS, DS, ES, and SS is ignored for 32-bit
        mode, same as 64-bit mode (treated as zero).`
   - 4.2.17 - MOV to Segment Register
      - Will fault if SS is written (Breaking anything that writes to
        SS).
      - Will not fault if CS, DS, ES are written (Thus it sets the
        segment but gets ignored due to 3.6).
2023-08-17 17:00:41 -07:00
Lioncache 764c844225 OpcodeDispatcher: Improve output of {V}MOVSHDUP 2023-08-17 17:01:50 -04:00
Lioncache 31719aac6a OpcodeDispatcher: Improve output of {V}MOVSLDUP 2023-08-17 17:01:46 -04:00
Alyssa Rosenzweig af21b8f3c7 Move External/FEXCore/ to FEXCore/
It is not an external component, and it makes paths needlessly long.
Ryan seemed amenable to this when we discussed on IRC earlier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2023-08-17 16:32:16 -04:00