Alyssa Rosenzweig
d9c779289c
OpcodeDispatcher: simplify RDRAND
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 15:18:00 -04:00
Alyssa Rosenzweig
5823e74bcd
OpcodeDispatcher: use increment carry for atomics
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
f138d7d9b8
OpcodeDispatcher: optimize JA/JNA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
256d166126
OpcodeDispatcher: optimize DAA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
34ee3bb8aa
OpcodeDispatcher: optimize clc/stc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Ryan Houdek
83fedd6c8f
Merge pull request #3912 from bylaws/addroverride
...
Don't apply the address-size flag to segment addresses
2024-08-01 18:35:06 -07:00
Billy Laws
be4777110c
OpcodeDispatcher: Don't apply the address-size flag to segment addresses
...
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Billy Laws
2c4fd79304
FEXCore: Add a generic spill/fill-all syscall ABI and use for Windows
...
Also drop the legacy hangover ABI as it has no users.
2024-07-31 17:25:59 +00:00
Paulo Matos
a1378f94ce
X87 Code Refactoring and Optimization Pass
2024-07-22 08:44:45 +02:00
Alyssa Rosenzweig
1e709d1150
OpcodeDispatcher: add RecordX87 helper
...
calls will be generated.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-16 09:07:35 +02:00
Ryan Houdek
d79b7fcc49
Merge pull request #3808 from alyssarosenzweig/rclse/3
...
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek
7e8d734e43
AVX256: Initial fixes just to get my unittest working
...
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Alyssa Rosenzweig
294f10fdd0
OpcodeDispatcher: reg cache mmx
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-11 13:21:14 -04:00
Tony Wasserka
b9829ed316
OpcodeDispatcher: Replace even more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
4ccec17676
OpcodeDispatcher: Replace more hand-written wrapper templates
2024-07-11 16:19:15 +02:00
Tony Wasserka
f45082043b
OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility
2024-07-11 16:19:14 +02:00
Alyssa Rosenzweig
1f01dd53f7
OpcodeDispatcher: reg cache fprs
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
72d41d70b6
OpcodeDispatcher: introduce GPR-only reg cache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig
2949bc211d
OpcodeDispatcher: thunk through FlushRegisterCache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-10 11:36:18 -04:00
Tony Wasserka
59fd13cc2f
OpcodeDispatcher: Avoid monomorphization of even more functions
2024-07-10 17:01:30 +02:00
Ryan Houdek
72d6c8ebd6
Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
...
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00
Alyssa Rosenzweig
05e4678e65
OpcodeDispatcher: fix missing masking on smaller RCR
...
I probably broke this when working on eliminating crossblock liveness.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 18:34:18 -04:00
Alyssa Rosenzweig
adc709db2f
OpcodeDispatcher: drop remnants of deferred flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 17:22:41 -04:00
Alyssa Rosenzweig
395573720d
OpcodeDispatcher: drop pointless flag defers for shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
0e62759d24
OpcodeDispatcher: stop deferring logical
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
926b6c3117
OpcodeDispatcher: don't defer mul flags
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
c9f9304ba5
OpcodeDispatcher: stop deferring obscure bitwise
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig
1bf31d20b6
OpcodeDispatcher: switch to CalculateFlags_SUB
...
most of these are deferred only to be calculated immediately anyway.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 16:24:53 -04:00
Ryan Houdek
653bf04db0
Merge pull request #3819 from alyssarosenzweig/bug/rcr-smol
...
Fix 8/16-bit RCR
2024-07-05 12:49:23 -07:00
Alyssa Rosenzweig
94bd79b2bf
OpcodeDispatcher: fix 8/16-bit RCR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-07-05 10:49:02 -04:00
Ryan Houdek
f6ec99bede
OpcodeDispatcher: Fixes rotates with zero not zero extending 32-bit result
...
For all the 32-bit rotates (except for RORX) we were failing to zero
extend the 32-bit result to the destination register when the rotate was
masked to zero.
Ensure we do this.
2024-07-04 14:35:42 -07:00
Ryan Houdek
0d06e3e47d
Revert "OpcodeDispatcher: add cache"
...
This reverts commit 46676ca376 .
2024-07-02 20:24:57 -07:00
Ryan Houdek
7d05610da7
OpcodeDispatcher: Optimize x86 canonical vector zero register
...
The canonical way to generate a zero register vector in x86 is to xor
itself. Capture this can convert it to canonical zero register instead.
Can get zero-cycle renamed on latest CPUs.
2024-06-29 22:21:53 -07:00
Ryan Houdek
f2f90eeb82
FEXCore: Make more distinctions between host register size and guest vector register size
...
We can support a few combinations of guest and host vector sizes
Host: 128-bit or 256-bit
Guest: 128-bit or 256-bit
The typical case is Host = 128-bit and Guest = 256-bit now that AVX is
implemented.
On 32-bit this changes to Host=128-bit and Guest=128-bit because we
disable AVX.
In the vixl simulator 32-bit turns in to Host=256-bit and Guest=128-bit.
And then in the vixl sim 64-bit turns in to Host=256-bit and
Guest=256-bit.
We cover all four combinations of guest and host vector register sizes!
Fixes a few assumptions that SVE256 = AVX256 basically.
2024-06-28 13:05:52 -07:00
Ryan Houdek
b0eb63ab9a
FEXCore: Fixes address size override on GPR sources and destinations
...
When the source or destination is a register, the address size override
doesn't apply. We were accidentally applying it on all sources
regardless of type which was causing us to zero extend on operations
that aren't affected by address size override.
This fixes the OpenSSL cert error in every application, but most
importantly Steam.
2024-06-27 14:12:01 -07:00
Ryan Houdek
8181552b16
AVX128: Actually install AVX helpers per thread.
...
How this didn't break the world in my testing I don't know.
2024-06-26 16:49:00 -07:00
Ryan Houdek
a4fa3a460e
OpcodeDispatcher: Implement AVX gathers with SVE256
...
Just to ensure we still have feature parity.
2024-06-26 16:00:53 -04:00
Alyssa Rosenzweig
d1d41f5645
Merge pull request #3763 from alyssarosenzweig/rclse/less-aggressive
...
Remove RCLSE
2024-06-26 15:14:14 -04:00
Ryan Houdek
94fd100fc7
Merge pull request #3719 from lioncash/f16c
...
OpcodeDispatcher: Handle F16C operations
2024-06-26 12:12:13 -07:00
Lioncache
cd5a809ec9
OpcodeDispatcher: Handle VCVTPS2PH
2024-06-26 15:05:03 -04:00
Lioncache
045a8efbeb
OpcodeDispatcher: Handle VCVTPH2PS
...
Fairly straightforward, since we already have handling for half-float conversions.
2024-06-26 15:05:00 -04:00
Ryan Houdek
54a1f7d833
Merge pull request #3764 from Sonicadvance1/rorx_masking
...
BMI2: Ensure rorx immediate masks by operation size correctly.
2024-06-26 11:52:47 -07:00
Alyssa Rosenzweig
46676ca376
OpcodeDispatcher: add cache
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-26 14:49:05 -04:00
Ryan Houdek
a515061465
BMI2: Ensure rorx immediate masks by operation size correctly.
2024-06-26 11:11:37 -07:00
Ryan Houdek
832b247fc1
SVE258: Implement support for FMA3
2024-06-25 11:24:46 -07:00
Ryan Houdek
7069643ae6
AVX128: Implement support for VPCLMULQDQ
...
This is just the 128-bit version twice.
2024-06-25 10:03:33 -04:00
Alyssa Rosenzweig
25f8a87429
OpcodeDispatcher: use Literal() helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-06-21 14:58:49 -04:00