Commit Graph
2128 Commits
Author SHA1 Message Date
Mai afaff9293b Merge pull request #2316 from Sonicadvance1/fix_negative_ficomi_f64
X87_F64: Fixes FICOM
2023-01-24 17:31:05 +00:00
Mai a28039f7cd Merge pull request #2350 from Sonicadvance1/optimize_dispatcher_slightly
Arm64: Merge two loads in to an LDP
2023-01-23 08:35:13 +00:00
Ryan Houdek a823d918c2 ARMEmitter: Support helper for long address generation
The current separated adr and adrp handlers are difficult to use if you
don't know if the resulting address is going to be within 1MB or 4GB.

Adds a `LongAddressGen` helper that will generate the various pieces of
code that will need to be emitted.

Backward labels:
 - Can generate three different code segments depending on distance to
   label
   - adr if label is within 1MB
   - adrp if label is 4K page aligned and within 4GB
   - adrp+add if label is within 4GB

Forward labels:
- Can generate three different code segments depending on distance to
  label
  - nop+adr if label is within 1MB
  - nop+adrp if label is 4K page aligned and within 4GB
  - adrp+add if label is within 4GB

There is still the limitation that this can't generate addresses to
labels that are >4GB away. Which is fine.
2023-01-22 16:03:17 -08:00
Ryan Houdek 7bf1742434 Arm64: Merge two loads in to an LDP
We can do a single LDP upfront when loading from the code cache, which
saves an instruction and one LDP costs the same as a single LDR.

Itty bitty optimization in the hot dispatcher.
2023-01-20 19:19:47 -08:00
Ryan Houdek 7897803753 Dispatcher: Fixes x86-64 SA_SIGINFO generation
Pulled from #2176.

On x86-64 the SA_SIGINFO sa_flag is actually a no-op. It is always used
even if not set.

Ensure that we setup siginfo_t regardless of flag being set.

On 32-bit x86 this still needs to be adhered to.

Little side bits that don't change anything
- EFLAGS is passed in signfo correctly.
- User provided restorer usage locations is documented but not
  implemented.
2023-01-20 11:08:57 -08:00
Ryan Houdek 40e5690e3a Dispatcher: Encode eflags in uc_mcontext
We were missing this.
2023-01-20 11:08:57 -08:00
Ryan Houdek bf4c5797db ARMEmitter: Removes some warnings that cropped up 2023-01-18 17:48:44 -08:00
Mai 4aa984aed9 Merge pull request #2322 from Sonicadvance1/opdispatcher_helpers
OpDispatcher: Fixes a few missing GPR/XMM helper usages
2023-01-19 00:05:11 +00:00
Mai 95e544c840 Merge pull request #2342 from Sonicadvance1/more_asimd_ops_pt2
ArmEmitter: Adds two more classes of ASIMD instructions
2023-01-18 20:33:40 +00:00
Mai 81e0ac7e0b Merge pull request #2331 from Sonicadvance1/more_asimd_ops
ArmEmitter: Adds three more classes of ASIMD instructions
2023-01-18 20:32:46 +00:00
Mai f8d92aa121 Merge pull request #2329 from Sonicadvance1/fix_cache_invalidation
Arm64: Fixes incorrect operation for CacheLineClear
2023-01-18 20:29:44 +00:00
Mai ee58c5de1d Merge pull request #2315 from Sonicadvance1/add_negative_unittests
unittests: Adds negative integer x87 tests
2023-01-18 20:27:59 +00:00
Mai 565ed450aa Merge pull request #2310 from Sonicadvance1/aarch64_move_to_switch
Arm64: Use switch statement for op handlers instead of jump table
2023-01-18 20:26:16 +00:00
Mai 90bcb8c70b Merge pull request #2309 from Sonicadvance1/remove_header
Emitter: Remove unused header
2023-01-18 20:25:07 +00:00
Ryan Houdek 9c93c6ffcd Merge pull request #2317 from Sonicadvance1/fix_spill_register
Arm64: Fix SpillRegister C&P error
2023-01-17 12:56:14 -08:00
Ryan Houdek 0ef8574a56 ARMEmitter: Adds two instruction classes 2023-01-15 20:16:10 -08:00
Ryan Houdek f1e1eaa8e5 ArmEmitter: Adds four missing ASIMD Shift by Imm ops 2023-01-15 20:16:10 -08:00
Ryan Houdek f614fc6fac Merge pull request #2338 from Sonicadvance1/optimize_cpuid
CPUID: Optimize initialization
2023-01-14 14:15:35 -08:00
Ryan Houdek d9a1bb9c35 CPUID: Optimize initialization
Map lookup was quite expensive, switched over to three small vectors
that are constexpr instead.

Some file querying and parsing was fairly slow as well. Optimized to
make that CPU time to go away.

This improves initialization time of CPUIDEmu by 33%
2023-01-14 13:37:26 -08:00
Ryan Houdek d69afaf925 MContext: Insert a stack cookie with assertions enabled
Pulled from #2176.
Ensures that when we are handling signals we are actually restoring a
stack state that is what we expect..

While this could randomly intersect with other stack data, it is highly
unlikely and will still capture incorrect stack frames otherwise.

Keeps it out of release build to ensure we aren't sticking random data
in the stack when it wouldn't have even been checked.
2023-01-14 11:59:10 -08:00
Ryan Houdek d787a38744 ArmEmitter: Adds three more classes of ASIMD instructions
Adds three classes:
- Advanced SIMD three same (FP16)
- Advanced SIMD two-register miscellaneous (FP16)
- Advanced SIMD three-register extension

A handful of the three-register extension unit tests are disabled
because the vixl disassembler doesn't support them.

Only six more classes of ASIMD operations remaining once this is merged.
2023-01-12 13:38:48 -08:00
Ryan Houdek b2f7f526f8 Arm64: Fixes incorrect operation for CacheLineClear
CIVAU does Clean+Invalidate to `Point Of Unification`
CIVAC does Clean+Invalidate to `Point of Coherency`

`Point of Unification` means to L2/L3, so unification of core
visibility.

`Point of Coherency` means SLC/RAM, All cores, DNA engines, etc must be
coherent.
2023-01-11 19:53:33 -08:00
Ryan Houdek bd55ed51b0 OpDispatcher: Moves a few missing XMM loadstores to helper usage
These were missed initially, these need to all be using the helper for
future optimizations.
2023-01-09 08:23:26 -08:00
Ryan Houdek 6d912be31e OpDispatcher: Moves a few missing GPR loadstores to helper usage
These were missed initially, these need to all be using the helper for
future optimizations.
2023-01-09 08:23:26 -08:00
CallumDev 806587d6ae Fix FPREM flags calculation in F64 2023-01-09 22:21:23 +10:30
Ryan Houdek 4daf2f0793 Jit64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:29:38 -08:00
Ryan Houdek 676cf59198 Jit64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:28:17 -08:00
Ryan Houdek 5aacdd744c Arm64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:25:25 -08:00
Ryan Houdek c3c68afc3c Arm64: Fixes incorrect sign extension
We were accidentally zero extending.
2023-01-08 13:23:40 -08:00
Ryan Houdek c7262120a6 Arm64: Fix SpillRegister C&P error
Was using the wrong sized registers in spill which was breaking Steam.
Oops.
2023-01-08 12:47:02 -08:00
Ryan Houdek 6977ae6b79 X87_F64: Fixes FICOM
This was not correctly converting both 32-bit and 16-bit integers over
to 64-bit double.
2023-01-08 11:02:52 -08:00
Ryan Houdek c2325e1772 Merge pull request #2314 from CallumDev/f64-integer-fix
F64: Fix integer immediates for add,mul,div,sub
2023-01-08 10:42:21 -08:00
CallumDev 9373fa0c06 F64: Fix integer immediates for add,mul,div,sub 2023-01-09 01:29:17 +10:30
Ryan Houdek ca9400ba52 Arm64: Fixes large offset spill slots
Found an application today (hashtree tests) that causes us to spill a
large amount of values on to the stack.

We were encoding larger offsets than what unsigned offset load and store
can handle.

If the offset is too large for the loadstore, use a temporary to put the
offset in to first.
2023-01-07 18:03:24 -08:00
Ryan Houdek 9a748c020d Arm64: Use switch statement for op handlers instead of jump table
Removes some startup time where we are copying nearly a page worth of
16byte vtable pointers at startup.

Also allows the compiler to choose to inline functions if it wants to.
2023-01-06 17:44:28 -08:00
Ryan Houdek 842e36e9b2 Emitter: Remove unused header 2023-01-06 10:34:41 -08:00
Ryan Houdek ffd9bb547d Arm64: Convert ARM Emitter over to new emitter
Not yet complete. Missing a full SVE implementation and needs
testing/validation.
2023-01-04 05:30:01 -08:00
Ryan Houdek 7a6ef8821f Arm64: Adds new ARM emitter
Still needs more work.
Missing operations, cleanup, validation
Notably SVE is missing large chunks.
2023-01-04 05:30:01 -08:00
Ryan Houdek 3904a5264f Merge pull request #2306 from lioncash/perm
OpcodeDispatcher: Handle immediate variants of VPERMILPD/VPERMILPS
2022-12-31 21:31:59 -08:00
lioncash b95c1719c3 OpcodeDispatcher: Handle VPERMILPS (immediate) 2023-01-01 05:15:25 +00:00
lioncash dfb3f31453 OpcodeDispatcher: Handle VPERMILPD (immediate) 2023-01-01 04:59:48 +00:00
lioncash 8031f76642 OpcodeDispatcher: Handle VMASKMOVDQU 2023-01-01 04:18:46 +00:00
lioncash ae8a5fa98d OpcodeDispatcher: Handle VPHSUBD 2023-01-01 03:10:21 +00:00
lioncash 6914598f9a OpcodeDispatcher: Handle VPHSUBW 2023-01-01 02:42:05 +00:00
lioncash 450aedc8b6 x86_64: Fix 256-bit UnZip/UnZip2
We weren't swapping the elements so that the operation acts like the two
vectors are concatenated.
2023-01-01 02:42:05 +00:00
lioncash 1e221210a2 OpcodeDispatcher: Move PHSUB impl to helper function
This will be used for the AVX variants.
2023-01-01 02:42:03 +00:00
Ryan Houdek 58ec2b2d7f Merge pull request #2303 from lioncash/swizz
OpcodeDispatcher: Zip elements instead of for loop insertion in PHSUB
2022-12-31 16:09:13 -08:00
lioncash 438adf2f45 x86_64/VectorOps: Handle OpSize==8 case in VUnZip/VUnZip2
Allows the x86 side of things to execute the new codepath in PHSUB
2022-12-31 23:30:59 +00:00
lioncash 9707e9a4df OpcodeDispatcher: Zip elements instead of for loop in PHSUB
Makes this much nicer for 128-bit and soon-to-be 256-bit vectors.
2022-12-31 23:30:38 +00:00
Ryan Houdek 9b8c92e275 Merge pull request #2302 from lioncash/dpp
OpcodeDispatcher: Handle VDPPD/VDPPS
2022-12-31 15:01:52 -08:00