Billy Laws
890e5e1f0f
FEXCore: Support disabling host cacheline clean/clear operations
2024-08-13 13:25:34 +00:00
Alyssa Rosenzweig
40812efaae
OpcodeDispatcher: better handle SIB indexing
...
if we have shift and a constant, we can save an instruction
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-13 08:41:16 -04:00
Alyssa Rosenzweig
3429321d59
OpcodeDispatcher: allow upper garbage for MOVGPR
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-13 08:41:16 -04:00
Alyssa Rosenzweig
6cddd6cbe7
OpcodeDispatcher: allow upper garbage for a2/a3
...
stores mask.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-13 08:41:16 -04:00
Alyssa Rosenzweig
23d07d7d0c
OpcodeDispatcher: fix folding negative offsets for 32-bit
...
I don't know what I was thinking when I wrote that code. Drop the silly logic
and let ConstProp inline the immediates. This fixes a lot of silly code
generated for 32-bit.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-13 08:41:16 -04:00
Alyssa Rosenzweig
91f4c54768
OpcodeDispatcher: optimize RDRAND on flagm
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 15:21:08 -04:00
Alyssa Rosenzweig
d9c779289c
OpcodeDispatcher: simplify RDRAND
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 15:18:00 -04:00
Alyssa Rosenzweig
5631ff4fd5
OpcodeDispatcher: optimize variable shifts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
34301319bf
OpcodeDispatcher: optimize IncrementByCarry
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
8eac3198b6
OpcodeDispatcher: use carry increment for ADC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
832edd4da3
OpcodeDispatcher: use carry increment for SBC
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
5823e74bcd
OpcodeDispatcher: use increment carry for atomics
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
a4545f493e
OpcodeDispatcher: add IncrementByCarry helper
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
3b8cd44ca4
IR: add NZCVSelectIncrement
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:43 -04:00
Alyssa Rosenzweig
f138d7d9b8
OpcodeDispatcher: optimize JA/JNA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
3bb9d44bf5
OpcodeDispatcher: optimize SetCFDirect
...
generally same # of instructions, but potentially fewer cycles:
old:
"rmif x4, #63 , #nzCv",
"cfinv"
new:
"xor x20, x4, #1 ",
"rmif x20, #63 , #nzCv"
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
06e6a1b19e
OpcodeDispatcher: optimize logic ops
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
256d166126
OpcodeDispatcher: optimize DAA
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
630285c589
OpcodeDispatcher: optimize SetPackedRFLAG
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
9621eca677
OpcodeDispatcher: fix setrflag masking
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
7eee50d929
OpcodeDispatcher: optimize mul/umul
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
1dc22e33ae
OpcodeDispatcher: inline some flag calculations
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
0ef8aaebeb
OpcodeDispatcher: optimize BZHI
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
d17c427e47
OpcodeDispatcher: optimize non-flagm2 comiss
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
f73fb62c6e
OpcodeDispatcher: optimize V(P)TEST
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
5976b712ca
OpcodeDispatcher: optimize bl*
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
63f5e64adb
OpcodeDispatcher: optimize add/sub
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
34ee3bb8aa
OpcodeDispatcher: optimize clc/stc
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
79c745929a
FEXCore: invert CF internally
...
Flag day change to the ABI.
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:06:42 -04:00
Alyssa Rosenzweig
8330cc6876
OpcodeDispatcher: defer carry inverts
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 13:05:21 -04:00
Alyssa Rosenzweig
16e6163677
OpcodeDispatcher: drop unused NZCVIndexMask
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Alyssa Rosenzweig
bbbc0dc9dc
OpcodeDispatcher: fix weird formatting
...
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io >
2024-08-10 11:26:12 -04:00
Ryan Houdek
0ecfc651b6
Merge pull request #3931 from Sonicadvance1/move_hostfeatures_init
...
FEXCore: Pass HostFeatures in to CreateNewContext directly
2024-08-09 20:30:17 -07:00
Alyssa Rosenzweig
633f624a69
Merge pull request #3930 from Sonicadvance1/hostfeatures_only_harnessrunner
...
HostFeatures: Removes feature flags always supported by FEX
2024-08-09 15:12:04 -04:00
Ryan Houdek
85d1b573ef
Merge pull request #3927 from bylaws/winafp
...
ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit
2024-08-08 22:21:23 -07:00
Ryan Houdek
2f8c5b4820
FEXCore: Pass HostFeatures in to CreateNewContext directly
...
The class constructor for ContextImpl::CPUID requires HostFeatures to be
available at construction time. Pass the host features struct directly
through during construction time instead, which cleans up the interface
slightly and fixes that issue.
2024-08-08 21:02:41 -07:00
Ryan Houdek
a1f55f0b0b
HostFeatures: Removes feature flags always supported by FEX
...
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.
Just check the feature bit directly in the HostRunner frontend for these
bits.
2024-08-08 19:05:55 -07:00
Ryan Houdek
1007f874bf
Merge pull request #3926 from bylaws/windef
...
FEXCore: Drop deferred signal handling on Windows
2024-08-08 17:33:44 -07:00
Mai
4882f10536
Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
...
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws
fe43a2bcb2
ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit
2024-08-07 18:34:35 +00:00
Billy Laws
6700511cdf
FEXCore: Drop deferred signal handling on Windows
...
The async signal issues this handles do not exist on Windows.
2024-08-07 18:31:48 +00:00
Ryan Houdek
e84848b16b
FEX: Moves HostFeatures querying to the frontend
...
This moves the CPU feature querying to the frontend. The primary purpose
here is for the wow64 frontend to not require linux-isms for querying
these features. This is required since non-Linux environments don't have
the "CPUID" feature for reading EL1 MSRs in EL0.
Wiring up the remaining wow64 registry querying is left for a future
exercise.
This also technically removes an xbyak requirement from FEXCore for when
building the x86 Test harness runner, but that doesn't really matter for
regular use cases.
2024-08-07 05:26:02 -07:00
Ryan Houdek
e613876e9d
AVX128: Optimize all cases of vpermq
...
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.
Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.
Fixes #3795
2024-08-06 09:08:30 -07:00
Alyssa Rosenzweig
a7424416d9
Merge pull request #3921 from bylaws/reloadf
...
Arm64Emitter: Reload STATE before SRA fill on ARM64EC
2024-08-06 09:28:23 -04:00
Billy Laws
ccf332d48e
Arm64Emitter: Reload STATE before SRA fill on ARM64EC
...
While ARM64EC code cannot use x28, it can be cleared by the kernel
when performing syscalls etc so restore it from the TEB to be safe.
2024-08-05 17:31:01 +00:00
Ryan Houdek
70c02d5c58
ARM64Emitter: Removes unused vixl CPU object
2024-08-03 22:26:00 -07:00
Ryan Houdek
2e4fb47848
HostFeatures: Read VL ourselves
...
Instead of calling out to vixl
2024-08-03 22:26:00 -07:00
Ryan Houdek
a4d5302369
Arm64: Adds Int helpers
...
One more vixl step removed.
2024-08-03 21:40:28 -07:00
Ryan Houdek
6ff3c90af3
CodeEmitter: Removes vestigial vixl usage
...
- IsImmLogical already existed in our CodeEmitter. We just forgot to
allow nullptr arguments and to use it.
- Adds an equivalent IsImmAddSub helper and uses it
This gets us closer to removing vixl's global initializers from FEXCore.
2024-08-03 21:04:56 -07:00
Ryan Houdek
201fe6ee23
Merge pull request #3909 from bylaws/ec-bitmap
...
Directly use the EC code bitmap for determining page arch
2024-08-02 10:55:43 -07:00