Compare commits

..
969 Commits
Author SHA1 Message Date
Ryan Houdek fb60a8a032 Docs: Update for release FEX-2408 2024-08-12 15:05:37 -07:00
Alyssa Rosenzweig f40bc134df Merge pull request #3945 from Sonicadvance1/config_nonnullable
Config: Little assume non-null check
2024-08-11 15:23:37 -04:00
Ryan Houdek f3811f04bd Config: Little assume non-null check
Removes a simple runtime nullcheck in Config::Layer::Set. Since we never pass a
nullptr to this.
2024-08-11 10:25:53 -07:00
Alyssa Rosenzweig 94bb7eb311 Merge pull request #3937 from Sonicadvance1/fix_script
Scripts: Fix issue in aarch64_fit_native
2024-08-10 20:07:43 -04:00
Ryan Houdek 03ca3e7e68 Merge pull request #3934 from bylaws/wow64-b
WOW64: Support the JIT API as used by Windows
2024-08-10 09:07:58 -07:00
Ryan Houdek 2c3e6cbe65 Merge pull request #3932 from bylaws/arm64-callchk
ARM64EC: Install a custom call checker to bypass NTDLL function patches
2024-08-10 09:07:31 -07:00
Alyssa Rosenzweig f9bdf0bd01 Merge pull request #3938 from Sonicadvance1/fix_vpblend_test
unittests: Fixes vpblend unittest
2024-08-10 11:08:03 -04:00
Ryan Houdek 4afc7adb05 unittests: Fixes vpblend unittest
This typo was causing undefined data to be used in the unittest, showed
up in debug builds.
2024-08-10 07:43:25 -07:00
Ryan Houdek 2152d1b2e9 Scripts: Fix issue in aarch64_fit_native
Apparently I messed this up in testing, is now fixed.
2024-08-09 21:04:26 -07:00
Ryan Houdek 0ecfc651b6 Merge pull request #3931 from Sonicadvance1/move_hostfeatures_init
FEXCore: Pass HostFeatures in to CreateNewContext directly
2024-08-09 20:30:17 -07:00
Ryan Houdek 4a3250ddea Merge pull request #3928 from bylaws/winval
InvalidationTracker: Better match Windows code invalidation behaviour
2024-08-09 20:29:57 -07:00
Alyssa Rosenzweig 633f624a69 Merge pull request #3930 from Sonicadvance1/hostfeatures_only_harnessrunner
HostFeatures: Removes feature flags always supported by FEX
2024-08-09 15:12:04 -04:00
Billy Laws 26a8a2717c WOW64: Mark the CPU area context as dirty initially
After thread creation, the WOW64 CPU area context needs to be flushed
into the FEX state before entering the JIT. Wine explicitly calls
BTCpuSetContext to trigger this but Windows doesn't.
2024-08-09 11:57:09 +00:00
Billy Laws 4c0e6d5779 WOW64: Shift down used TLS slots
Fixes a crash on native Windows.
2024-08-09 11:57:09 +00:00
Billy Laws a350ef5d1b WOW64: Match the Windows function protoypes 2024-08-09 11:57:09 +00:00
Billy Laws 9d9bd750e2 ARM64EC: Install a custom call checker to bypass NTDLL function patches
Some programs will hook the NTDLL exports that FEX depends on, the
regular ARM64EC call checker will detect such patches and invoke the
JIT to run them, which leads to infinite recursion if those same
exports are used during code compilation. Fix this by resolving all
patchable FFSs to their native ARM implementations for all indirect
calls performed by FEX, skipping any x86 patches.
2024-08-09 11:48:18 +00:00
Ryan Houdek 85d1b573ef Merge pull request #3927 from bylaws/winafp
ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit
2024-08-08 22:21:23 -07:00
Ryan Houdek 2f8c5b4820 FEXCore: Pass HostFeatures in to CreateNewContext directly
The class constructor for ContextImpl::CPUID requires HostFeatures to be
available at construction time. Pass the host features struct directly
through during construction time instead, which cleans up the interface
slightly and fixes that issue.
2024-08-08 21:02:41 -07:00
Ryan Houdek a1f55f0b0b HostFeatures: Removes feature flags always supported by FEX
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.

Just check the feature bit directly in the HostRunner frontend for these
bits.
2024-08-08 19:05:55 -07:00
Ryan Houdek 7b1d9540b7 Merge pull request #3925 from bylaws/arm64ecrt
ARM64EC: Introduce FEX-side CRT and Windows API replacements
2024-08-08 17:57:25 -07:00
Ryan Houdek 1007f874bf Merge pull request #3926 from bylaws/windef
FEXCore: Drop deferred signal handling on Windows
2024-08-08 17:33:44 -07:00
Ryan Houdek 24ea4b7537 Merge pull request #3924 from bylaws/svc
ARM64EC: Handle direct syscall instructions
2024-08-08 17:32:56 -07:00
Billy Laws 93ab57454a InvalidationTracker: Better match Windows code invalidation behaviour
When given a NULL base address, Windows invalidation callbacks will
ignore the given size and invalidate all code.
2024-08-08 12:46:49 +00:00
Mai 4882f10536 Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws fe43a2bcb2 ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit 2024-08-07 18:34:35 +00:00
Billy Laws 6700511cdf FEXCore: Drop deferred signal handling on Windows
The async signal issues this handles do not exist on Windows.
2024-08-07 18:31:48 +00:00
Billy Laws e836639427 ARM64EC: Manually define the ARM64EC linker structures 2024-08-07 15:49:41 +00:00
Billy Laws 7eb3f33162 Update jemalloc submodule 2024-08-07 15:49:41 +00:00
Billy Laws b35a514ea6 CMake: Don't link CMake's set of extra system libraries on Windows
These are already linked in by default with clang, sp having these set
here only served to prevent the -nostdlib compiler option from having
any effect, as CMake will always explicitly include the libs in the
compiler cmdline.
2024-08-07 15:49:41 +00:00
Billy Laws 8f372821f8 Windows: Use the FEX CRT 2024-08-07 15:49:41 +00:00
Billy Laws 6376d06bed Windows: Introduce a minimal Windows API replacement 2024-08-07 15:49:41 +00:00
Billy Laws 1bb293d4be Windows: Pull in some math and formatting functions from Musl 2024-08-07 15:49:41 +00:00
Billy Laws 0e069f1e97 Windows: Introduce a minimal CRT replacement
It is dangerous for FEX to rely on the system CRT as calls can have side
effects that are also visible to the running application, and if patches
are applied to any CRT exports used during compilation the call checker
would trigger a reentry into the JIT to compile the patch and hence
deadlock. Only functions that FEX actively uses are implemented, with
the rest triggering an abort.
2024-08-07 15:49:41 +00:00
Billy Laws 9074b810a9 CMake: Always enable jemalloc for MinGW builds 2024-08-07 15:49:41 +00:00
Mai c5fe8723d3 Merge pull request #3918 from Sonicadvance1/constexpr_config_maps
Config: Converts two LUT maps over linear scan arrays
2024-08-07 10:40:22 -04:00
Ryan Houdek 434bffac33 Merge pull request #3916 from Sonicadvance1/frontend_hostfeatures
FEX: Moves HostFeatures querying to the frontend
2024-08-07 05:33:55 -07:00
Ryan Houdek e84848b16b FEX: Moves HostFeatures querying to the frontend
This moves the CPU feature querying to the frontend. The primary purpose
here is for the wow64 frontend to not require linux-isms for querying
these features. This is required since non-Linux environments don't have
the "CPUID" feature for reading EL1 MSRs in EL0.

Wiring up the remaining wow64 registry querying is left for a future
exercise.

This also technically removes an xbyak requirement from FEXCore for when
building the x86 Test harness runner, but that doesn't really matter for
regular use cases.
2024-08-07 05:26:02 -07:00
Ryan Houdek 69ed39d49e Merge pull request #3892 from Sonicadvance1/optimize_vpermq
AVX128: Optimize all cases of vpermq
2024-08-06 20:07:28 -07:00
Ryan Houdek 230bde6aef InstcountCI: Adds vpermq coverage 2024-08-06 09:08:30 -07:00
Ryan Houdek c24d7aacba unittests/ASM: Adds vpermq test that covers all immediate encodings
To ensure we cover all tests when optimizing.
2024-08-06 09:08:30 -07:00
Ryan Houdek e613876e9d AVX128: Optimize all cases of vpermq
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.

Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.

Fixes #3795
2024-08-06 09:08:30 -07:00
Mai 1473129a8f Merge pull request #3920 from Sonicadvance1/fix_newline_asm
SpinWaitLock: Fixes missing newline in asm
2024-08-06 12:08:14 -04:00
Alyssa Rosenzweig c42808cb70 Merge pull request #3904 from Sonicadvance1/packaging
Scripts: Workaround deprecated parse_version
2024-08-06 11:52:59 -04:00
Ryan Houdek 054c119e2e Config: Converts two LUT maps over linear scan arrays
These two maps used for environment lookup translations were getting
globally initialized and then registers with atexit handlers.

Switch over to a constexpr array and just do linear scans. This plus
short-circuiting the environment loader so it skips all entries that
don't start with `FEX_` has the side benefit of cutting the CPU time to
1/10th the time.

This plus #3917 removes the global static initializers entirely from
this file.
2024-08-06 07:44:08 -07:00
Alyssa Rosenzweig f75bd2f09b Merge pull request #3922 from bylaws/structs
Windows: Pull in additional method and structure definitions from wine
2024-08-06 09:29:03 -04:00
Alyssa Rosenzweig a7424416d9 Merge pull request #3921 from bylaws/reloadf
Arm64Emitter: Reload STATE before SRA fill on ARM64EC
2024-08-06 09:28:23 -04:00
Alyssa Rosenzweig cadb0a2ddb Merge pull request #3923 from bylaws/except
ARM64EC: Improvements to exception flag handling
2024-08-06 09:27:50 -04:00
Alyssa Rosenzweig 2da819c0f3 Merge pull request #3919 from Sonicadvance1/remove_vestigial_vixl_usage
CodeEmitter: Removes vestigial vixl usage
2024-08-06 09:26:24 -04:00
Alyssa Rosenzweig e0c783de74 Merge pull request #3917 from Sonicadvance1/remove_static_vector
Config: Removes a static vector initializer
2024-08-06 09:26:03 -04:00
Billy Laws 6c003fcb9a Windows: Pull in more method/structure definitions from wine 2024-08-05 19:23:18 +00:00
Billy Laws c4faffc0e2 Windows: Add complete NTDLL export definitions
Generated from wine's ntdll.spec
2024-08-05 19:22:12 +00:00
Billy Laws cc2d21f411 WOW64: Resolve the wine unix call dispatcher at runtime 2024-08-05 19:22:12 +00:00
Billy Laws 59686a6c60 ARM64EC: Clear TF in the exception resumption context after a trap
Matches Windows behaviour.
2024-08-05 17:38:45 +00:00
Billy Laws 0ab864da17 ARM64EC: Reset the CPU area JIT state before handling exceptions
An exception in JIT code acts as a transition to ARM64EC code (in
NTDLL for exception handling etc) as such, much like ExitFunction,
InSimulation must be unset. InSyscallCallback is unset for robustness
against exception in the JIT itself.
2024-08-05 17:38:45 +00:00
Billy Laws 3c32271dd0 ARM64EC: Merge EFlags with the current JIT flags state on a ctx sync
Only NZCV and TF are passed through to BeginSimulation as the rest are
lost when converting to a native context and back on the ntdll side. To
prevent thread suspension from wiping out the rest of the flags, only
copy these specific flags into the current JIT EFlags state.
2024-08-05 17:38:45 +00:00
Billy Laws 507a95b817 ARM64EC: Map TF to PSTATE.SS when reconstructing a native context 2024-08-05 17:38:45 +00:00
Billy Laws 9bb9e954c2 ARM64EC: Spill EFlags when reconstructing state from in the JIT 2024-08-05 17:38:45 +00:00
Billy Laws 7f3582bb23 ARM64EC: Only clear the trap flag when handling an exception
Better matches Windows emulator behaviour.
2024-08-05 17:38:45 +00:00
Billy Laws efe15ce336 ARM64EC: Handle direct syscall instructions
Most syscalls on Windows are done by calling into their NTDLL thunks,
however some DRMs parse out their numbers from NTDLL and directly call
them. Support this by redirecting to their entry thunks in the FEX
syscall handler.
2024-08-05 17:35:32 +00:00
Billy Laws 21b0f35ef4 ARM64EC: Populate a LUT mapping NTDLL FFS exports to their native impls
To prevent FEX from redirecting to x86 code when NTDLL exports it calls
into are patched, a custom call checker will be used that checks this LUT
to redirect calls rather than the FFS itself.
2024-08-05 17:35:32 +00:00
Billy Laws ccf332d48e Arm64Emitter: Reload STATE before SRA fill on ARM64EC
While ARM64EC code cannot use x28, it can be cleared by the kernel
when performing syscalls etc so restore it from the TEB to be safe.
2024-08-05 17:31:01 +00:00
Ryan Houdek 802a32ce8a SpinWaitLock: Fixes missing newline in asm
This would cause the atomic load after the wfe to be dropped,
effectively returning stale data.
2024-08-04 06:57:05 -07:00
Ryan Houdek 70c02d5c58 ARM64Emitter: Removes unused vixl CPU object 2024-08-03 22:26:00 -07:00
Ryan Houdek 2e4fb47848 HostFeatures: Read VL ourselves
Instead of calling out to vixl
2024-08-03 22:26:00 -07:00
Ryan Houdek a4d5302369 Arm64: Adds Int helpers
One more vixl step removed.
2024-08-03 21:40:28 -07:00
Ryan Houdek 6ff3c90af3 CodeEmitter: Removes vestigial vixl usage
- IsImmLogical already existed in our CodeEmitter. We just forgot to
  allow nullptr arguments and to use it.
- Adds an equivalent IsImmAddSub helper and uses it

This gets us closer to removing vixl's global initializers from FEXCore.
2024-08-03 21:04:56 -07:00
Ryan Houdek c114279118 Config: Removes a static vector initializer
Saw this vector was getting initialized at runtime, sticking around, and
installing an atexit handler. This is completely unnecessary, just use
the OPT_BASE handler directly to walk the environment variable names.
2024-08-03 19:13:29 -07:00
Ryan Houdek 7ffd3e55d5 Merge pull request #3915 from bylaws/winbase
ARM64EC: Support the JIT API as is used by Windows
2024-08-02 10:56:46 -07:00
Ryan Houdek 201fe6ee23 Merge pull request #3909 from bylaws/ec-bitmap
Directly use the EC code bitmap for determining page arch
2024-08-02 10:55:43 -07:00
Ryan Houdek 83fedd6c8f Merge pull request #3912 from bylaws/addroverride
Don't apply the address-size flag to segment addresses
2024-08-01 18:35:06 -07:00
Ryan Houdek dedf4a93d0 Merge pull request #3913 from bylaws/logcommon
Commonise logging and fallback to a log file for debug output on Windows
2024-08-01 12:06:35 -07:00
Ryan Houdek 1f59f0e226 Merge pull request #3914 from bylaws/wincfg
Config: Search more locations for the config directory on Windows
2024-08-01 12:05:50 -07:00
Ryan Houdek c3c2b6115d Merge pull request #3910 from bylaws/f80
F80: Drop dependency on state stored in TLS
2024-08-01 12:03:41 -07:00
Billy Laws 115fbb5039 ARM64EC: Implement remaining notification callbacks 2024-08-01 12:06:25 +00:00
Billy Laws 27973d5637 ARM64EC: Fix the exception dispatcher stack layout 2024-08-01 12:06:24 +00:00
Billy Laws f121be649d ARM64EC: Match the Windows BT API function prototypes 2024-08-01 12:06:24 +00:00
Billy Laws af9bcb3efd ARM64EC: Avoid syncing uninitialized context members to the JIT state
Windows can sometimes pass in incomplete contexts to BeginSimulation. So
only sync the valid parts specified in ContextFlags.
2024-08-01 12:06:05 +00:00
Billy Laws 4877bb3f19 ARM64EC: Support directly issuing the NtContinue syscall
This is required for handling SMC with the ResetToConsistentState
arguments as used in Windows, as using the NTDLL exported NtContinue
would wipe out any reserved registers in the ARM64EC ABI.

For Windows the syscall numbers are somewhat stable, and the SVC
instruction can be called directly. Since wine doesn't handle that on
ARM64, hardcode the system call number and manually call into wine
dispatcher. Once wine gains proper syscall thunks, those can be
parsed to get the number and the hardcoding dropped.
2024-08-01 12:06:05 +00:00
Billy Laws e2583249c6 ARM64EC: Allocate the emulator stack ourselves
Actual Windows does not allocate it for us.
2024-08-01 12:06:05 +00:00
Billy Laws f48071dcbe ARM64EC: Switch to the emulator stack in BeginSimulation
Windows calls this function on the guest stack for some reason.
2024-08-01 12:06:05 +00:00
Billy Laws fa72be2ec5 Windows: Don't warn for unknown CPU features
This happens regularly as wine/games will scan all features from 0 to 64.
2024-08-01 12:06:05 +00:00
Billy Laws b7ff6dd9c6 Windows: Disable logging if SilentLog is enabled 2024-08-01 12:04:59 +00:00
Billy Laws cb6d60aa87 Config: Search more locations for the config directory on Windows 2024-08-01 11:48:40 +00:00
Billy Laws 80c9a43eef Windows: Fallback to a log file for debug output on Windows
OutputDebugString etc are exception based and thus don't really work for
FEX's needs as often times logs can happen in places where exceptions
cannot be thrown.
2024-08-01 11:45:07 +00:00
Billy Laws 05155778d4 Windows: Commonise logging code 2024-08-01 11:45:07 +00:00
Tony Wasserka 49b8dae189 Merge pull request #3908 from bylaws/pdb
CMake: Add option to build PDB debug info instead of DWARF
2024-08-01 10:01:04 +02:00
Ryan Houdek 3e59fc0a8c Merge pull request #3911 from bylaws/x80bug
x87StackOptimizationPass: Default initialise StackMemberInfo members
2024-07-31 22:57:16 -07:00
Ryan Houdek f98c010854 Merge pull request #3907 from bylaws/ec-mema
AllocatorHooks: Correct memory API usage on Windows
2024-07-31 17:48:13 -07:00
Ryan Houdek 10ee963f52 Merge pull request #3906 from bylaws/ec-sysbi
FEXCore: Add a generic spill/fill-all syscall ABI and use for Windows
2024-07-31 17:46:55 -07:00
Billy Laws af7462ee6a unittests: Add test using the address-override flag with segment addressing 2024-07-31 20:04:30 +01:00
Billy Laws be4777110c OpcodeDispatcher: Don't apply the address-size flag to segment addresses
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Billy Laws 696503680a F80: Drop dependency on state stored in TLS
Windows cannot support the implicit TLS as was used prior, so introduce
a state structure and pass it in to functions where necessary.
2024-07-31 18:51:42 +01:00
Billy Laws dd4d3bcf38 AllocatorHooks: Correct memory API usage on Windows
These issues end up being tolerated by wine but not actual windows.
2024-07-31 18:36:50 +01:00
Billy Laws f26bb6bf53 x87StackOptimizationPass: Default initialise StackMemberInfo members
Not doing so is UB.
2024-07-31 17:30:29 +00:00
Billy Laws 2c4fd79304 FEXCore: Add a generic spill/fill-all syscall ABI and use for Windows
Also drop the legacy hangover ABI as it has no users.
2024-07-31 17:25:59 +00:00
Billy Laws 910ec4aadd Dispatcher: Directly use the EC code bitmap for determining page arch
The prior approach using the L2 cache was flawed as it assumed L2
page entries had a 1-1 correspondence with actual pages. While the L2
cache could be extended to handle aliases with EC, this could lead to
thrashing etc. The cost of a lookup in the actual EC code bitmap is
cheap enough to perform every time considering the infrequency of calls
to ARM64EC code when compare to X86 L2 hits.
2024-07-31 17:24:50 +00:00
Billy Laws fb7275b3d8 Revert: "LookupCache: Track ARM64EC page state in the code cache"
This reverts the commit 526e3e654f.
2024-07-31 17:24:50 +00:00
Billy Laws 51b4bfc6a6 FEXCore: Move ARM64EC TEB offset constants to Arm64Emitter
These need to be used from outside the dispatcher, and there are already
similar defines for EC registers in the emitter header.
2024-07-31 17:24:50 +00:00
Billy Laws 9f7bc94f9d CMake: Add option to build PDB debug info instead of DWARF 2024-07-31 17:23:24 +00:00
Ryan Houdek 069e2ce62a Merge pull request #3905 from pmatos/FixWarn
Fix nasm warning in Rounding.asm
2024-07-31 08:06:12 -07:00
Paulo Matos 5bbbced1bd Fix nasm warning in Rounding.asm 2024-07-31 16:23:39 +02:00
Alyssa Rosenzweig 941fd9c6ea Merge pull request #3901 from pmatos/TopUsage
Reuse Top in ReconstructFSW_Helper
2024-07-31 08:01:05 -04:00
Paulo Matos 3332220d06 instcountci: Intersperse flag retrieval and FSW insertion 2024-07-31 12:05:17 +02:00
Paulo Matos 5933a59c09 Intersperse flag retrieval and FSW insertion 2024-07-31 12:04:54 +02:00
Paulo Matos aee8c9def2 instcountci: Reuse Top in ReconstructFSW_Helper 2024-07-31 11:57:19 +02:00
Paulo Matos c2136272bf Reuse Top in ReconstructFSW_Helper
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek dd26b0c879 Merge pull request #3903 from Sonicadvance1/v6.10_syscalls
Syscalls: Updates for v6.10
2024-07-31 02:22:26 -07:00
Ryan Houdek 3d9114b74b Merge pull request #3902 from Sonicadvance1/man_page_fix
man: Fixes newline issue with strenum
2024-07-31 02:21:47 -07:00
Ryan Houdek 3f3e937967 man: Fixes newline issue with strenum
In the environment section this was causing the next environment
variable line to be merged with the strenum options

Also makes it so strenum options doesn't have a spurious comma at the
end of the list.
2024-07-31 02:02:05 -07:00
Ryan Houdek 4beb29141b Scripts: Workaround deprecated parse_version
Different approach from #3579

Instead of completely ddropping the deprecated path, support the new
path and the old path using python try-except import exceptions.

This allows us to continue using old packages in CI, while supporting
the future API once pkg_resources gets deprecated and removed. Best of
both worlds.
2024-07-31 01:57:31 -07:00
Paulo Matos 9a8e7eaace instcountci: Add instcountci for fnstsw from fast path 2024-07-31 09:59:32 +02:00
Paulo Matos 4227e012aa Add instcountci for fnstsw from fast path 2024-07-31 09:59:28 +02:00
Ryan Houdek 40bbcb7061 Syscalls: Updates for v6.10
Only mseal was added and can be a simple passthrough for us. Which is
nice.
2024-07-30 19:08:27 -07:00
Ryan Houdek 67663b812e Linux: Update syscall defines for v6.10 2024-07-30 19:01:52 -07:00
Ryan Houdek d24d0a95a0 Merge pull request #3894 from pmatos/RoundingModeTests
ASM Tests: X87 Rounding modes
2024-07-29 23:26:52 -07:00
Ryan Houdek 87fbcf754d Vector: Optimize pblendw
Using a brute force solver to add in more optimized code paths

- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations

Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek 403fd62b34 Merge pull request #3890 from Sonicadvance1/refactor_frontend_threadmanager
FEXCore: Removes ThreadManager
2024-07-26 13:27:43 -07:00
Ryan Houdek d92b6a9ac4 Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek c2092bfed0 Merge pull request #3893 from pmatos/FNINITFix
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Ryan Houdek a6cf7fa508 Merge pull request #3896 from pmatos/CTestSkip
Test running scripts tell ctest of skipped tests
2024-07-26 13:25:28 -07:00
Alyssa Rosenzweig 5ff09f5091 OpcodeDispatcher/X87: use less creative Refs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-26 14:30:42 -04:00
Paulo Matos 9af7ee6bd2 instcountci: Fix call to FNINITF64 and refactor 2024-07-26 14:56:10 +02:00
Paulo Matos d1e36f264f Fix call to FNINITF64 and refactor 2024-07-26 14:56:06 +02:00
Paulo Matos b1ec50c7c2 Test running scripts tell ctest of skipped tests
CMake sets 125 as the skipped test exit code that the scripts use.
2024-07-26 14:04:54 +02:00
Mai 93eead243f Merge pull request #3864 from Sonicadvance1/threads_atexit_remove
Threads: Setup the stack tracker to not need global initialization
2024-07-26 06:38:29 -04:00
Paulo Matos ceac38a6ac ASM Tests: X87 Rounding modes 2024-07-26 10:07:58 +02:00
Ryan Houdek 3b2e657fd4 FEXCore: Removes ThreadManager
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.

Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
2024-07-25 14:54:10 -07:00
Ryan Houdek 380ba0a014 Merge pull request #3889 from Sonicadvance1/refactor_frontend_exithandler
FEXCore: Refactor ExitHandler slightly
2024-07-25 14:53:25 -07:00
Mai 1fe497d1dd Merge pull request #3891 from Sonicadvance1/remove_cpubackendfeatures
FEXCore: Removes CPUBackendFeatures
2024-07-24 20:53:10 -04:00
Ryan Houdek 7816b150d0 FEXCore: Removes CPUBackendFeatures
We were only ever hardcoding true for TBL2 and Flags now. Get rid of it.
2024-07-24 17:19:30 -07:00
Ryan Houdek ce8bc9d25c FEXCore: Refactor ExitHandler slightly
Instead of passing the TID back to the exit handler, just pass the whole
thread object. This will allow some cleanups with the frontend thread
tracking soon

NFC
2024-07-24 14:39:56 -07:00
Ryan Houdek 4634688aca InstcountCI: Update for AVX128 blends 2024-07-23 19:24:19 -07:00
Ryan Houdek dd3e3ed189 unittests/ASM: Implements a vpblendw test
Runs through all immediate encodings for vpblendw and crcs the results
to ensure correct behaviour. This was just a concern because of the typo
in documentation. But it is also good to have.
2024-07-23 19:24:19 -07:00
Ryan Houdek f8ef6feff9 AVX128: Optimize blends
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.

One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.

Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek 9201ac5a6b Merge pull request #3882 from Sonicadvance1/scalar_afp_fma
AVX128: Implement support for scalar FMA with AFP
2024-07-22 13:19:59 -07:00
Ryan Houdek 8ebf049fb9 InstcountCI: Update for Scalar FMA with AFP 2024-07-22 12:58:20 -07:00
Ryan Houdek 3c5b59d985 AVX128: Implement support for scalar FMA with AFP
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.

Fixes #3793
2024-07-22 12:58:19 -07:00
Ryan Houdek 6b91e0cb0e Merge pull request #3887 from alyssarosenzweig/ir/prefix
json_ir_generator: stop prefixing arguments
2024-07-22 12:57:09 -07:00
Alyssa Rosenzweig 587b924de9 json_ir_generator: stop prefixing arguments
stop prefixing the arguments when we generate allocate ops (in particular), this
is more convenient and simpler. in exchange we need to prefix Op to avoid a
collision on fcmpscalarinsert which has an argument named Op, but that's a local
change at least.

came up when experimenting with new IR, but I think this is probably a win by
itself.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-22 13:50:21 -04:00
Tony Wasserka d507f4c9b1 Merge pull request #3547 from pmatos/wip_x87_stack
x87 Stack Optimization
2024-07-22 14:42:02 +02:00
Paulo Matos 39bc2a82c1 instcountci: X87 Pass and refactoring 2024-07-22 08:50:01 +02:00
Paulo Matos 774325dcf2 Tests: X87 Refactoring and Pass 2024-07-22 08:44:45 +02:00
Paulo Matos a1378f94ce X87 Code Refactoring and Optimization Pass 2024-07-22 08:44:45 +02:00
Ryan Houdek 77ec950ff2 Merge pull request #3885 from alyssarosenzweig/opt/zero-flag
Optimize zero x87 flags
2024-07-21 13:06:25 -07:00
Alyssa Rosenzweig 592d6cc43f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:50:10 -04:00
Alyssa Rosenzweig 610caf8529 ConstProp: treat StoreContext as zeroable
todo: FPR equivalent.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:09 -04:00
Alyssa Rosenzweig d20b46e46f IR: drop LoadFlag/StoreFlag ops
pointless, we can just load/store the context now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:09 -04:00
Alyssa Rosenzweig 4094aa1b9a DeadStoreElimination: drop flag handling
now that we do everything via NZCV, this is mostly vestigial. DF/x87 flags are
sufficiently rare to be "don't care"s here, and we don't even have multiblock
enabled yet!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:08 -04:00
Ryan Houdek f8c6baae97 Merge pull request #3883 from Sonicadvance1/implement_daz
Arm64: Implements support for DAZ using AFP.FIZ
2024-07-21 10:03:34 -07:00
Ryan Houdek 5c9bb6594c Merge pull request #3884 from Sonicadvance1/remove_vex_telem
Telemetry: Remove VEX flag
2024-07-20 19:44:28 -07:00
Ryan Houdek 56df57e980 InstcountCI: Update 2024-07-20 17:26:27 -07:00
Ryan Houdek f7b4d25803 Telemetry: Remove VEX flag
This is no longer necessary and it also no longer provides us any useful
information. Since we expose the AVX CPUID flag, basically everything
uses VEX encoding now, so it is basically always set.
2024-07-20 17:24:00 -07:00
Ryan Houdek 4fffe68f81 InstcountCI: Update 2024-07-20 15:57:01 -07:00
Ryan Houdek 95b15d788b Arm64: Fix filling static registers
Some locations could end up with SRA registers that only spilled one
register.
Allow passing in temporaries from the call site.
Fixes rpid and syscalls asserting.
2024-07-20 15:57:01 -07:00
Ryan Houdek ae9312bdab unittests: Implements a DAZ test
Specifically does a vector add with and without DAZ enabled and ensures
the value is different when the source values contain a denormal.
2024-07-20 15:34:54 -07:00
Ryan Houdek b78da2e5ad Arm64: Implements support for DAZ using AFP.FIZ
When AFP is supported then we can actually support DAZ. This might also
fix the audio corruption in Animal Well but I can't test it until Steam
is running on Oryon. Requires a bit of plumbing for MXCSR which we were
hacking around before but now we actually want to store the value.

Fixes #3856
2024-07-20 15:34:54 -07:00
Ryan Houdek 54fc8cb0bd TestHarnessRunner: Support querying AFP for features
Also fixes desync of flags
2024-07-20 15:34:54 -07:00
Ryan Houdek 228009c283 Merge pull request #3881 from bylaws/race
FixedSizePooledAllocation: Fix a race when unclaiming disowned buffers
2024-07-20 12:35:35 -07:00
Billy Laws 1f878ce4cd FixedSizePooledAllocation: Fix a race when unclaiming disowned buffers
A disowned buffer could be unclaimed or claimed by a different thread in
the time between the !IsFree check and locking the allocation mutex.
Fix this and prevent such errors in the future by always checking
ownership with the allocator locked before attempting to unclaim
buffers.
2024-07-20 00:12:52 +00:00
Alyssa Rosenzweig e4b7a65a49 Merge pull request #3880 from pmatos/InstCountMemcpy
Add x87 memcpy instcountci tests
2024-07-19 08:53:23 -04:00
Paulo Matos c77a707dbe Add x87 memcpy instcountci tests 2024-07-19 09:09:34 +02:00
Ryan Houdek f81fc4e4f0 Merge pull request #3866 from Sonicadvance1/ArgumentLoader_atexit_remove
ArgumentLoader: Removes static fextl::vector usage
2024-07-18 13:12:03 -07:00
Ryan Houdek d385e496d3 Merge pull request #3879 from Sonicadvance1/cpuid_leafs
EmulatedFiles: Adds a few leaf CPUID flags
2024-07-18 13:10:21 -07:00
Mai f8c4c543e3 Merge pull request #3871 from Sonicadvance1/improve_vpshufd_vpermilps
AVX128: Improve VPERMILPS/PD and VPSHUFD
2024-07-18 15:58:47 -04:00
Ryan Houdek d2f903ae55 EmulatedFiles: Adds a few leaf CPUID flags
We support leaf functions now, so add the few that were calling for it.
We will be gaining support for the xsave ones relatively soon, so its
good to have them supported.

Also deletes a couple of cdt/cqm things that aren't exposed and we won't
be supporting.
2024-07-18 07:02:49 -07:00
Ryan Houdek d1249ec5cf Merge pull request #3878 from neobrain/refactor_fix_format_oops
EmulatedFiles: Fix bad formatting
2024-07-18 06:44:44 -07:00
Tony Wasserka bf9a6d763c EmulatedFiles: Fix bad formatting 2024-07-18 15:06:57 +02:00
Ryan Houdek 0b829d2c46 unittests: Adds a test for full pshufd imm coverage 2024-07-18 04:13:03 -07:00
Ryan Houdek bddb533fa0 InstcountCI: Add some more of the cases 2024-07-18 04:13:03 -07:00
Ryan Houdek 1c35eeffeb Vector: Optimize PSHUFD with brute force search
With a brute force search of methods between 1-3 instructions we cover a
lot more cases more optimally.

There's definitely still more cases (and probably some that can reduce
from 3 instruction to 2), but covering 44 cases is a pretty good margin
already.
2024-07-18 04:10:58 -07:00
Ryan Houdek c7254e31ed InstcountCI: Update for VPERM/VPSHUFD improvements 2024-07-18 04:10:58 -07:00
Ryan Houdek b0bd8a62a2 AVX128: Improve VPERMILPS/PD and VPSHUFD
VPSHUFD and VPERMILPS are aliases of each other.

Reuses the implementation path from the PSHUFD implementation which has
a few swizzles and then a table lookup.

VPERMILPD is a very simple swizzle per 128-bit lane.

Fixes #3797
Fixes #3784
2024-07-18 04:10:58 -07:00
Ryan Houdek 17a55fbb39 Merge pull request #3876 from alyssarosenzweig/json/x87
Autogenerate LoweredX87() query, misc json_ir_generator cleanup in the area
2024-07-18 04:10:21 -07:00
Alyssa Rosenzweig dedec83881 json_ir_generator: autoderive array names
these are purely internal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:32:51 -04:00
Alyssa Rosenzweig 9fd5c73633 json_ir_generator: generate IsLoweredX87 helper
X87 pass will use this query.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:31:46 -04:00
Alyssa Rosenzweig bdb890a8b0 json_ir_generator: rename X87 -> LoweredX87
to reflect its actual meaning

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:31:14 -04:00
Alyssa Rosenzweig 5043d09771 json_ir_generator: use textwrap.dedent, f-string
for size

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:28:59 -04:00
Ryan Houdek da51169ba9 Merge pull request #3875 from alyssarosenzweig/ir/gethostflag
IR: garbage collect premature F80Cmp optimizations
2024-07-17 03:05:48 -07:00
Ryan Houdek f72cee480f Merge pull request #3874 from alyssarosenzweig/opt/reconstructftw
X87: save uop in ReconstructFTW
2024-07-17 03:05:37 -07:00
Alyssa Rosenzweig 7546160811 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Alyssa Rosenzweig e7d5a01c5f IR: remove F80Cmp flags
nothing is optimizing around this, it's just adding pointless complexity. if we
want to actually optimize F80Cmp, the right way would be to lift the
implementation into the OpcodeDispatcher or JIT. it wouldn't be terribly
difficult. This kludge doesn't get us closer there.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Alyssa Rosenzweig 0c3a8d0bc8 IR: remove GetHostFlag
it doesn't get host flags, it's just an extra Bfe used in x87. pointless and
confusing!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:44:34 -04:00
Alyssa Rosenzweig 19e58cac62 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 13:54:28 -04:00
Alyssa Rosenzweig c4ba7eee87 X87: save uop in ReconstructFTW
noticed while reviewing Paulo's work

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 13:54:09 -04:00
Ryan Houdek 09c4a5594a Merge pull request #3870 from Sonicadvance1/enable_more_tests
github: Vixl simulator enable more asm tests
2024-07-16 07:23:22 -07:00
Alyssa Rosenzweig d204155661 Merge pull request #3872 from pmatos/X87AutoMarking
X87 Stack Ops Auto-marking
2024-07-16 09:45:42 -04:00
Tony Wasserka 924b8c10a9 Merge pull request #3873 from pmatos/UnusedFunction
Remove unused function MmapOverride
2024-07-16 12:30:04 +02:00
Paulo Matos 9017cd14c8 Remove unused function MmapOverride 2024-07-16 11:16:07 +02:00
Paulo Matos 8d89adef2e Add IR stack operations
These IR operations deal implicitly with the x87 stack and are removed
by the x87 stack optimization pass.
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 6615b55c12 json_ir_generator: call RecordX87Use when generating ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 66865dd177 json_ir_generator: alias X87 to !JITDispatch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 1e709d1150 OpcodeDispatcher: add RecordX87 helper
calls will be generated.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 476ee0cd7d IR: track whether x87 is used in header
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Ryan Houdek 6df51a57b3 Merge pull request #3868 from bylaws/arm64ec-oldnew
ARM64EC frontend
2024-07-15 09:54:56 -07:00
Ryan Houdek e65545a537 Merge pull request #3867 from pmatos/NoDisableTests
Remove Disabled_Tests file
2024-07-15 09:54:08 -07:00
Ryan Houdek b8e864ffdf Merge pull request #3865 from Sonicadvance1/telemetry_atexit
Telemetry: Change how visibility of telemetry values work
2024-07-15 09:53:37 -07:00
Paulo Matos ed87c01470 Simplify Disabled_Tests and remove pr57275 from failures
Disabled_Tests was mostly a copy of Known_Failures. Leave only the race
condition on SIGPROF test (mcount_pic.c).

Also remove pr57275.c from known failures. It passes now that we support
 AVX.

Also if a test is disabled, just skip it
2024-07-13 07:08:02 +02:00
Ryan Houdek 6c21b86a8f github: Vixl simulator enable more asm tests
We were only running SVE256 and SVE128 with AVX disabled.

Enable asm tests with SVE256, SVE128, and ASIMD, all running with AVX
enabled to hit all the tests.
2024-07-12 20:39:01 -07:00
Ryan Houdek d79b7fcc49 Merge pull request #3808 from alyssarosenzweig/rclse/3
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek b9a6caea8d Merge pull request #3844 from Sonicadvance1/fix_vmovq
AVX128: Fixes vmovq loading too much data
2024-07-12 17:07:32 -07:00
Ryan Houdek 9688f5e17d Merge pull request #3863 from Sonicadvance1/remove_static_ioctl_handlers
Ioctl32: Removes static fextl::vector in ioctlemulation
2024-07-12 17:06:16 -07:00
Billy Laws f6f8d26426 Update jemalloc submodule 2024-07-12 19:24:13 +00:00
Billy Laws dba0a1d09e ARM64EC: Initialize x86 control registers on thread start 2024-07-12 18:51:31 +00:00
Billy Laws af3145674e ARM64EC: Fixup exception information for faulting x86 instructions
FEX emulates faulting instructions (e.g. ud2 or int 2d) by jumping to
the dispatcher and filling out a structure with fault details in the
thread context. Parse this out into a windows exception record structure
so the correct fault information can be seen by the guest.
2024-07-12 18:51:31 +00:00
Billy Laws 3c19e634b3 ARM64EC: Rethrow exceptions from within the JIT
As the exception dispatcher is initially invoked on the emulator stack,
control needs to be transferred to the dispatcher on the guest stack
after recovering the x86 RSP to allow for invoking x86 exception
handlers.
2024-07-12 18:41:20 +00:00
Billy Laws f964a5187e ARM64EC: Implement BeginSimulation
This is used by the kernel (or UNIX side of ntdll in wine) to jump into
x86 code with the given context as is necessary when e.g. returning from
an exception.
2024-07-12 18:41:13 +00:00
Billy Laws 8e0fdfc325 ARM64EC: Add a helper to lookup the redirected address of an export
FEX is unable to deal with reentrant compilation of any x64 hotpatches
so they need to be ignored by bypassing FFSs and calling directly into
the native target.
2024-07-12 18:41:08 +00:00
Billy Laws 839f9ecd3b Windows: Add ARM64EC image structures 2024-07-12 18:41:06 +00:00
Billy Laws 95fc69b628 ARM64EC: Handle SMC 2024-07-12 18:41:02 +00:00
Billy Laws b9da95838a ARM64EC: Handle unaligned atomic accesses 2024-07-12 18:40:43 +00:00
Billy Laws 1059279d5d ARM64EC: Handle calls into ARM64EC code with an 8-byte-aligned SP
ARM64 requires that SP is always 16-byte aligned for memory accesses,
but ARM64EC shares the SP between x64 code and ARM64 code, the former
of which doesn't enforce such a restriction. This causes crashes in
programs such as HITMAN 3 that don't correctly follow the Windows ABI
and call into system library functions with SP only 8-byte-aligned.
Fixup stack alignment in such cases by leaving the 8-byte return
address on the stack and returning to a lone 'ret' instruction instead.
2024-07-12 18:30:04 +00:00
Billy Laws 5dc85307a6 Windows: Introduce an initial ARM64EC frontend
This allows for running x64 applications under wine without having to run all
of wine under FEX. The JIT is invoked when ARM64EC code performs an indirect
branch to x64 code, and left whenever the x64 code calls into ARM64EC
code.
2024-07-12 18:07:50 +00:00
Billy Laws 549e06aade CMake: Enable assembly source file support 2024-07-12 18:01:22 +00:00
Billy Laws 3b189f6d7d WOW64: Install into lib
This convention is used by most other projects.
2024-07-12 18:01:22 +00:00
Ryan Houdek a5d3692b53 ArgumentLoader: Removes static fextl::vector usage
Removes a global initializer and atexit registration

Ownership of this data has always been the frontend and the config
system, we just used these static vectors as a side-channel.
2024-07-12 04:48:22 -07:00
Ryan Houdek 97a68cb643 Telemetry: Change how visibility of telemetry values work
Removes global initializer for telemetry values since their address is
visible and PIC relative code loading handles the address fetching for
us.
2024-07-12 03:18:23 -07:00
Ryan Houdek d1b5dfd4b1 Threads: Setup the stack tracker to not need global initialization
Also removes the atexit handler installation
This now gets tracked by an object owned by FEXLoader (and shared with
the pthreads interface)
2024-07-12 03:00:20 -07:00
Ryan Houdek 6cdaea680d Ioctl32: Removes static fextl::vector in ioctlemulation
Removes a global static initializer for the vector and its atexit
handler.

This handler array can be consteval similar to the x86 tables so it can
be generated entirely at compile time.
2024-07-12 02:05:32 -07:00
Ryan Houdek 870e395ac4 Merge pull request #3862 from Sonicadvance1/remove_atexit_logman
LogManager: Removes fextl::vector usage
2024-07-12 02:05:02 -07:00
Ryan Houdek 04592f82f5 Merge pull request #3861 from Sonicadvance1/remove_atexit_vdso
VDSO: Stop using a vector for a static
2024-07-12 02:04:25 -07:00
Ryan Houdek 19e849283f Merge pull request #3860 from Sonicadvance1/force_noinline
OpcodeDispatcher: Force noinline for the function call in the Bind helper
2024-07-12 00:14:04 -07:00
Ryan Houdek b6e1469cd4 Merge pull request #3847 from pmatos/CoverageSupport
Enable coverage configuration for FEX
2024-07-12 00:02:08 -07:00
Ryan Houdek 5ef0db994d VDSO: Stop using a vector for a static
This causes a global initializer that registers an atexit handler.

Be smarter, use an std::array and pass its data around using a span
instead.

Removes the global initializer and removes the atexit installation
2024-07-11 23:53:57 -07:00
Ryan Houdek b523407a3e LogManager: Removes fextl::vector usage
We never use more than one logging method at a time so this was
overengineered for what it is doing.

Instead only allow one handler for messages and throw messages each
which just is a pointer.

Removes a global initializer and an atexit handler being installed
2024-07-11 22:51:56 -07:00
Ryan Houdek 8021dc10a1 OpcodeDispatcher: Force noinline for the function call in the Bind helper
Clang was inlining a few of the functions it was calling. So force it
never to inline since this is supports to be a little shim trampoline
only.
2024-07-11 19:00:42 -07:00
Ryan Houdek 7e8d734e43 AVX256: Initial fixes just to get my unittest working
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Ryan Houdek 3d90d1ab4f InstcountCI: Update for vmovq fix 2024-07-11 18:34:06 -07:00
Ryan Houdek 3c7318d7c8 AVX128: Fixes vmovq loading too much data
This was doing a 128-bit load from memory and then a 64-bit zero extend
which looked like a spurious move but it was trying to match the
behaviour of vmovq where it needed the zero extend.

Also adds a unit test to ensure that we aren't loading too much data by
loading right up against a page boundary.

Fixes #3787
2024-07-11 18:34:05 -07:00
Ryan Houdek fc0b233046 Merge pull request #3859 from neobrain/refactor_opdispatch_templates
OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility
2024-07-11 18:18:23 -07:00
Mai e25918d846 Merge pull request #3858 from Sonicadvance1/implement_nt_load
Implement support for SSE4.1/AVX NT loads
2024-07-11 14:22:41 -04:00
Alyssa Rosenzweig d78b0ea435 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 3a334c4585 Reapply "IR: drop RCLSE"
This reverts commit 78aee4d96e.
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 8dae4bcd44 OpcodeDispatcher: drop stale comment
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 294f10fdd0 OpcodeDispatcher: reg cache mmx
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Tony Wasserka b9829ed316 OpcodeDispatcher: Replace even more hand-written wrapper templates 2024-07-11 16:19:15 +02:00
Tony Wasserka 4ccec17676 OpcodeDispatcher: Replace more hand-written wrapper templates 2024-07-11 16:19:15 +02:00
Tony Wasserka f45082043b OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility 2024-07-11 16:19:14 +02:00
Tony Wasserka 3222f13dde Fix comment formatting 2024-07-11 16:19:14 +02:00
Mai b282620a48 Merge pull request #3857 from Sonicadvance1/sve_bitperm
Arm64: Implement support for SVE bitperm
2024-07-11 05:05:41 -04:00
Ryan Houdek 3ff1ff8f74 InstcountCI: Update for svebitperm 2024-07-11 01:46:35 -07:00
Ryan Houdek e24b01b6cb Arm64: Implement support for SVE bitperm 2024-07-11 01:46:35 -07:00
Tony Wasserka 9a8694c2f3 Merge pull request #3853 from neobrain/refactor_warn_fixes
Fix all the warnings
2024-07-11 10:12:41 +02:00
Tony Wasserka 070a9148aa Merge pull request #3852 from neobrain/refactor_opdispatch_codesize
OpcodeDispatcher: Avoid template monomorphization to reduce FEXLoader binary size
2024-07-11 09:58:49 +02:00
Tony Wasserka f19fe3b6f3 Fix warning about an expression with side effects being passed to __builtin_assume
LOGMAN_THROW_AA_FMT has no benefit over LOGMAN_THROW_A_FMT here, so just use
the latter.
2024-07-11 09:54:31 +02:00
Tony Wasserka 8d2b15665d Fix unused-variable warnings 2024-07-11 09:54:30 +02:00
Tony Wasserka 4dec8f22f8 Fix packed-non-pod warnings 2024-07-11 09:54:30 +02:00
Tony Wasserka a39b3aca78 Fix invalid-offsetof warnings due to JsonAllocator not being standard layout
Inheritance can be used here instead, which allows the JsonAllocator to be
reconstructed using a downcast.
2024-07-11 09:54:30 +02:00
Tony Wasserka 5dc4ab062d Fix invalid-offsetof warnings due to InternalThreadState not being standard layout
See https://github.com/llvm/llvm-project/issues/53021 for more information
about unique_ptr turning non-standard-layout.
2024-07-11 09:54:30 +02:00
Ryan Houdek 31f82c1d96 InstcountCI: Update for SVE NT load support 2024-07-10 23:07:58 -07:00
Ryan Houdek 548fd9daf8 OpcodeDispatcher: Implement support for SSE4.1 NT load 2024-07-10 23:07:37 -07:00
Ryan Houdek f831f5a0e1 AVX128: Implement support for NT Load 2024-07-10 23:07:14 -07:00
Ryan Houdek 4c21aa2604 Arm64: Implement support for NT Loads with ASIMD fallback 2024-07-10 23:06:46 -07:00
Ryan Houdek c9efb75714 CodeEmitter: Implement support for SVE NT loads 2024-07-10 23:06:19 -07:00
Ryan Houdek 5e56bdc0fd InstcountCI: Add support for SVE bitperm 2024-07-10 21:48:37 -07:00
Ryan Houdek 3554d5c2f7 HostFeatures: Check for SVE bit permute extension 2024-07-10 21:45:07 -07:00
Mai 5fe405e1fb Merge pull request #3855 from neobrain/fix_aotir_uniqueptr
AOTIR: Change std::unique_ptr to fextl::unique_ptr
2024-07-10 17:04:12 -04:00
Tony Wasserka 8381d44bbd Merge pull request #3854 from neobrain/fix_default_delete
fextl: Properly handle nullptr arguments in fextl::default_delete
2024-07-10 23:00:52 +02:00
Tony Wasserka 56bb3744a5 AOTIR: Change std::unique_ptr to fextl::unique_ptr 2024-07-10 19:34:24 +02:00
Tony Wasserka 470b435afd fextl: Properly handle nullptr arguments in fextl::default_delete
This reflects behavior of std::default_delete.
2024-07-10 19:17:50 +02:00
Alyssa Rosenzweig a4f8bbff02 OpcodeDispatcher: reg cache avx high
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig cf5ab05b90 OpcodeDispatcher: reg cache AbridgedFTW
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 3a2ce240f9 OpcodeDispatcher: reg cache DF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 1f01dd53f7 OpcodeDispatcher: reg cache fprs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 72d41d70b6 OpcodeDispatcher: introduce GPR-only reg cache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 42b5b1f64c Core: partially flush register cache per instruction
This will mitigate problems later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 2949bc211d OpcodeDispatcher: thunk through FlushRegisterCache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 5e0952159d unittests: add test for a MMX register cache bug
this failed on an earlier version of the register cache.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:34:24 -04:00
Tony Wasserka 441187470e OpcodeDispatcher: Avoid monomorphization of some AVX functions 2024-07-10 17:01:30 +02:00
Tony Wasserka 59fd13cc2f OpcodeDispatcher: Avoid monomorphization of even more functions 2024-07-10 17:01:30 +02:00
Tony Wasserka c9e7bfdf16 OpcodeDispatcher: Avoid monomorphization of more functions 2024-07-10 17:01:30 +02:00
Tony Wasserka 2d700c381e OpcodeDispatcher: Avoid monomorphization of large functions 2024-07-10 17:01:30 +02:00
Ryan Houdek 72d6c8ebd6 Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00
Ryan Houdek 991c6941c1 Merge pull request #3849 from alyssarosenzweig/ir/drop-parser-2
Scripts: drop remnant of IR parser
2024-07-09 16:48:36 -07:00
Alyssa Rosenzweig f974696e34 Scripts: drop remnant of IR parser
unused.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-09 16:08:38 -04:00
Alyssa Rosenzweig 3ef9ea94e5 Merge pull request #3848 from pmatos/FTSTX87Tests
Tests for X87 FTST
2024-07-09 09:10:29 -04:00
Paulo Matos 381ce23fd7 Tests for X87 FTST 2024-07-09 13:36:16 +02:00
Mai af6a0be832 Merge pull request #3842 from Sonicadvance1/fix_f64_to_i32
VCVT{T,}PD2DQ fixes and optimization
2024-07-09 03:49:31 -04:00
Ryan Houdek 287fe5beac InstcountCI: Update 2024-07-09 00:38:48 -07:00
Ryan Houdek b9c214e6e8 OpcodeDispatcher: Use new IR op for vcvt{t,}pd2dq
Also fixes a bug where it was failing to zero the upper bits of the
destination register in the AVX128 implementation. Which the updated
unit tests now check against.

Fixes a minor precision issue that was reported in #2995. We still don't
return correct values for overflow. x86 always returns maximum negative
int32_t on overflow, ARM will return maximum negative or positive
depending on sign of the double.
2024-07-09 00:38:47 -07:00
Ryan Houdek d3d76aa8ce IR: Adds new F64 -> I32 operation that changes behaviour depending on SVE
SVE added the ability to do F64 -> I32 conversions directly without an
fcvtn inbetween. So maybe sure to support them.
2024-07-09 00:38:47 -07:00
Ryan Houdek 3bea08da5f Merge pull request #3843 from Sonicadvance1/remove_half_moves_fma3
Arm64: Remove one move if possible in FMA operations
2024-07-09 00:25:07 -07:00
Paulo Matos 3d5cacbdc3 Enable coverage configuration for FEX 2024-07-09 08:03:41 +02:00
Mai 7ccb252069 Merge pull request #3837 from Sonicadvance1/optimize_sve_vpgatherdq
AVX128: Extends 32-bit indexes path for 128-bit operations
2024-07-08 22:01:02 -04:00
Ryan Houdek 31547462bb InstcountCI: Update for final SVE AVX128 improvements. 2024-07-08 18:44:07 -07:00
Ryan Houdek b3a7a973a1 AVX128: Extends 32-bit indexes path for 128-bit operations
The codepath from #3826 was only targeting 256-bit sized operations.
This missed the vpgatherdq/vgatherdpd 128-bit operations. By extending
the codepath to understand 128-bit operations, we now hit these
instruction variants.

With this PR, we now have SVE128 codepaths that handle ALL variants of
x86 gather instructions! There are zero ASIMD fallbacks used in this
case!

Of course depending on the instruction, the performance still leaves a
lot to be desired, and there is no way to emulate x86 TSO behaviour
without an ASIMD fallback, which we will likely need to add as a
fallback at some point.

Based on #3836 until that is merged.
2024-07-08 18:44:07 -07:00
Mai 22b26696ba Merge pull request #3836 from Sonicadvance1/optimize_sve_vpgatherdd
AVX128: Optimize the vpgatherdd/vgatherdps cases that would fall back to ASIMD
2024-07-08 21:43:36 -04:00
Ryan Houdek 495241f8ca InstcountCI: Update for wide gather vpgatherdd SVE usage 2024-07-08 18:12:28 -07:00
Ryan Houdek 4afbfcae17 AVX128: Optimize the vpgatherdd/vgatherdps cases that would fall back to ASIMD
With the introduction of the wide gathers in #3828 this has opened new
avenues for optimizing these cases that would typically fall back to
ASIMD. In the cases that 32-bit SVE scaling doesn't fit, we can instead
sign extend the elements in to double-width address registers.

This then feeds naturally in to the SVE path even though we end up
needing to allocate 512-bits worth of address registers. This ends up
being significantly better than the ASIMD path still.

Relies on #3828 to be merged first
Fixes #3829
2024-07-08 18:12:28 -07:00
Mai 3627de4cbc Merge pull request #3828 from Sonicadvance1/optimize_wide_gathers
AVX128: Optimize QPS/QD variant of gather loads!
2024-07-08 21:11:36 -04:00
Ryan Houdek 007c07e612 InstcountCI: Update for wide gathers 2024-07-08 17:19:18 -07:00
Ryan Houdek ec7c8fd922 AVX128: Optimize QPS/QD variant of gather loads!
SVE has a special version of their gather instruction that gets similar
behaviour to x86's VGATHERQPS/VPGATHERQD instructions.

The quirk of these instructions that the previous SVE implementation
didn't handle and required ASIMD fallback, was that most gather
instructions require the data element size and address element size to
match. This x86 instruction uses a 64-bit address size while loading 32-bit
elements. This matches this specific variant of the SVE instruction, but
the data is zero-extended once loaded, requiring us to shuffle the data
after it is loaded.

This isn't the worst but the implementation is different enough that
stuffing it in to the other gather load will cause headaches.

Basically gets 32 instruction variants to use the SVE version!

Fixes #3827
2024-07-08 17:19:18 -07:00
Ryan Houdek c5a0ae7b34 IR: Adds new QPS gather load variant! 2024-07-08 17:19:18 -07:00
Ryan Houdek 4bd207ebf3 Arm64: Moves 128Bit gather ASIMD emulation to its own helper
It is going to get reused.
2024-07-08 17:19:18 -07:00
Tony Wasserka 45011234d9 Merge pull request #3845 from pmatos/TESTJOBCOUNTFix
Use nproc only if TEST_JOB_COUNT not specified
2024-07-08 22:31:23 +02:00
Paulo Matos 24017f379e Use nproc only if TEST_JOB_COUNT not specified 2024-07-08 21:38:56 +02:00
Mai aad7656b38 Merge pull request #3826 from Sonicadvance1/scale_32bit_gather
AVX128: Extend 32-bit address indices when possible
2024-07-08 15:29:44 -04:00
Ryan Houdek 80de890f05 InstcountCI: Update for removed FMA moves 2024-07-08 04:50:49 -07:00
Ryan Houdek 62cec7b6b2 Arm64: Remove one move if possible in FMA operations
If the destination isn't any of the incoming sources then we can avoid
one of the moves at the end. This half works around the problem proposed
in #3794, but doesn't solve the entire problem.

To solve the other half of the moving problem means we need to solve the
SRA allocation problem for this temporary register with addsub/subadd, so it gets allocated
for both the FMA operation and the XOR operation.
2024-07-08 04:44:40 -07:00
Ryan Houdek c9c163cd7b unittests: Update vcv{t,tt}pd2dq tests to ensure upper bits of destination are cleared 2024-07-08 03:30:10 -07:00
Mai 95a9f32bf0 Merge pull request #3840 from Sonicadvance1/extend_vinsert128_tests
unittests: Extends vinsert{i,f}128 tests for garbage data
2024-07-07 13:39:20 -04:00
Mai c4ae761a0e Merge pull request #3841 from Sonicadvance1/add_missing_cpu_names
CPUID: Adds a few missing CPU names for new CPU cores
2024-07-07 13:38:27 -04:00
Ryan Houdek 0653b346e0 CPUID: Adds a few missing CPU names for new CPU cores
These should be making their way to the market sooner rather than later
so make sure we have the descriptor text for them.
2024-07-07 02:40:19 -07:00
Ryan Houdek fa587398bd unittests: Extends vinsert{i,f}128 tests for garbage data
Just to ensure we don't hit an issue with masking the immediate bits.

Fixes #3753
2024-07-07 02:16:21 -07:00
Ryan Houdek 6b67857151 InstcountCI: Adds a missing gather instruction invariant
Oops, must have accidentally deleted this while copying things around.
2024-07-06 18:32:36 -07:00
Ryan Houdek 81165f0c40 InstcountCI: Update for 32-bit gather sign extend optimization 2024-07-06 18:32:35 -07:00
Ryan Houdek df40515087 AVX128: Extend 32-bit address indices when possible
When loading 256-bits of data with only 128-bits of address indices, we
can sign extend the source indices to be 64-bit. Thus falling down the
ideal path for SVE where each 128-bit lane is loading the data to
addresses in a 1:1 element ratio.

This means we use the SVE path more often because of this.

Based on top of #3825 because the prescaling behaviour was introduced
there. This implements its own prescaling when the sign extension occurs
because ARM's SSHLL{,2} instruction gives us that for free.

This additionally fixes a bug where we were accidentally loading the top
128-bit half of the addresses for gathers when it was unnecessary, and
on the AVX256 side it was duplicating and doing some additional work
when it shouldn't have.

It'll be good to walk the commits when looking at this one, as there are
a couple of incremental changes that are easier to follow that way.

Fixes #3806
2024-07-06 18:32:35 -07:00
Ryan Houdek c77922e3e5 InstcountCI: Update for previous fix 2024-07-06 18:32:35 -07:00
Ryan Houdek 0f9abe68b9 AVX128: Fixes accidentally loading high addr register when unnnecessary
Was missing a clamp on the high half when encounting a 128-bit gather
instruction. Was causing us to unconditionally load the top half when it
was unncessary.
2024-07-06 18:32:35 -07:00
Ryan Houdek c168ee6940 Arm64: Implements VSSHLL{,2} IR ops 2024-07-06 18:32:35 -07:00
Ryan Houdek 0d4414fdd0 AVX128: Removes templated AddrElementSize and add as argument
NFC
2024-07-06 18:32:35 -07:00
Ryan Houdek 968d5e0d8f Merge pull request #3774 from bylaws/win-ci
FEXCore ARM64EC CI support
2024-07-06 18:22:57 -07:00
Ryan Houdek 635182b57c Merge pull request #3832 from bylaws/wow64-wine
WOW64: Mark the FEX dll as a wine builtin
2024-07-06 17:58:00 -07:00
Ryan Houdek 9d0b6ce75e Merge pull request #3835 from bylaws/ec-topdown
AllocatorHooks: Allocate from the top down on windows
2024-07-06 17:40:36 -07:00
Ryan Houdek 2fdd80fe3a Merge pull request #3833 from bylaws/common-tso
Windows: Commonise TSOHandlerConfig
2024-07-06 17:38:45 -07:00
Ryan Houdek dbac23b749 Merge pull request #3834 from bylaws/ec-amd64
Windows: Report as an AMD64 processor when targeting ARM64EC
2024-07-06 17:38:13 -07:00
Billy Laws 7fa7061aa5 Windows: Report as an AMD64 processor when targeting ARM64EC 2024-07-06 20:37:15 +00:00
Billy Laws e45e631199 AllocatorHooks: Allocate from the top down on windows
FEX allocations can get in the way of allocations that are 4gb-limited
even in 65-bit mode (i.e. those from LuaJIT), so allocate starting from
the top of the AS to prevent conflicts.
2024-07-06 20:35:38 +00:00
Billy Laws b21e77c1e0 Windows: Commonise TSOHandlerConfig 2024-07-06 19:20:49 +00:00
Billy Laws ba33294225 WOW64: Mark the FEX dll as a wine builtin
Allows it to be automatically picked up by wine during prefix setup,
without a manual dll override.

Thanks to AndreRH for pointing me to this.
2024-07-06 19:19:36 +00:00
Billy Laws 97c21cc3a7 CI: Add ARM64EC build CI 2024-07-06 17:27:41 +01:00
Billy Laws 7d7e6f5326 CMake: Disable WOW64 module for ARM64EC 2024-07-06 17:27:41 +01:00
Billy Laws 5e15bd935e CMake: Disable glibc jemalloc for MinGW builds 2024-07-06 17:27:41 +01:00
Ryan Houdek 9bad09c45f Merge pull request #3823 from alyssarosenzweig/bug/shl-var-small
Fix CF with small shifts
2024-07-06 01:33:57 -07:00
Ryan Houdek 47d077ff22 Merge pull request #3825 from Sonicadvance1/scale_64bit_gather
AVX128: Prescale addresses in gathers if possible
2024-07-05 19:10:43 -07:00
Ryan Houdek bbf8dde3ca Merge pull request #3824 from alyssarosenzweig/bug/rc2
OpcodeDispatcher: Fix 8/16-bit rcr masking
2024-07-05 17:01:16 -07:00
Ryan Houdek 6e8ca3bc6c InstcountCI: Update for gather prescaling 2024-07-05 16:47:11 -07:00
Ryan Houdek 11a494d7b3 AVX128: Prescale addresses in gathers if possible
If the host supports SVE128, if the address element size and data size is 64-bit, and the scale is not one of the two that is supported by SVE; Then prescale the addresses.
64-bit address overflow masks the top bits so is well defined that we
can scale the vector elements and still execute the SVE code path in
that case. Removing the ASIMD code paths from a lot of gathers.

Fixes #3805
2024-07-05 16:47:11 -07:00
Alyssa Rosenzweig 9b570de33f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:44:21 -04:00
Ryan Houdek b67343fc5a unittests: Adds a test for small shift flags calculation
Currently we calculate CF incorrectly in the case of small shifts with
large offsets.
2024-07-05 18:38:12 -04:00
Alyssa Rosenzweig 5a3c0eb83c OpcodeDispatcher: fix shl with 8/16-bit variable
the special case here lines up with the special case of using a larger shift for
a smaller result, so we can just grab CF from the larger result.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:38:12 -04:00
Alyssa Rosenzweig 10391608a0 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:34:18 -04:00
Ryan Houdek 51c57cc5ae unittests: More rotate with carry unit tests
Looks like we missed some edge cases with small carry rotate. Adds even
more unit tests.
2024-07-05 18:34:18 -04:00
Alyssa Rosenzweig 05e4678e65 OpcodeDispatcher: fix missing masking on smaller RCR
I probably broke this when working on eliminating crossblock liveness.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:34:18 -04:00
Alyssa Rosenzweig 0f0e402db4 OpcodeDispatcher: fix CF with 8/16-bit immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:24:34 -04:00
Alyssa Rosenzweig 837bccb1d8 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 17:24:51 -04:00
Alyssa Rosenzweig adc709db2f OpcodeDispatcher: drop remnants of deferred flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 17:22:41 -04:00
Alyssa Rosenzweig 395573720d OpcodeDispatcher: drop pointless flag defers for shifts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 0e62759d24 OpcodeDispatcher: stop deferring logical
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 926b6c3117 OpcodeDispatcher: don't defer mul flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig c9f9304ba5 OpcodeDispatcher: stop deferring obscure bitwise
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig fabd6be5af OpcodeDispatcher: drop SUB defer
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 1bf31d20b6 OpcodeDispatcher: switch to CalculateFlags_SUB
most of these are deferred only to be calculated immediately anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:53 -04:00
Ryan Houdek 653bf04db0 Merge pull request #3819 from alyssarosenzweig/bug/rcr-smol
Fix 8/16-bit RCR
2024-07-05 12:49:23 -07:00
Ryan Houdek b77a25b21a Merge pull request #3818 from alyssarosenzweig/jit/shiftbymaskstozero
JIT: fix ShiftFlags masking
2024-07-05 12:49:16 -07:00
Alyssa Rosenzweig 9db6931cea InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 10:49:12 -04:00
Ryan Houdek bad5cef52b unittests: Adds rotate with carry test for large rotates
FEX-Emu currently doesn't do large rotates for small data sources
correctly. This will fail CI until fixed in OpcodeDispatcher
2024-07-05 10:49:02 -04:00
Alyssa Rosenzweig 94bd79b2bf OpcodeDispatcher: fix 8/16-bit RCR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 10:49:02 -04:00
Alyssa Rosenzweig b746146f4e InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 09:57:42 -04:00
Ryan Houdek 8ac9bb5c72 unittests: Adds test for flags when shifting by zero 2024-07-05 09:57:42 -04:00
Alyssa Rosenzweig 1b552a6f62 JIT: fix ShiftFlags masking
we don't update flags for a nonzero shift that masks to zero.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 09:57:42 -04:00
Alyssa Rosenzweig 97329ccc7a Merge pull request #3812 from Sonicadvance1/fix_rotates_with_zero
OpcodeDispatcher: Fixes rotates with zero not zero extending 32-bit result
2024-07-05 09:48:01 -04:00
Mai f2d1f2de56 Merge pull request #3817 from Sonicadvance1/fix_x87_integer_indefinite
Softfloat: Fixes Integer indefinite return for 16-bit signed values
2024-07-04 23:11:44 -04:00
Ryan Houdek 692c2fae96 Merge pull request #3813 from alyssarosenzweig/bug/fix-sbb
Fix 16-bit SBB
2024-07-04 19:52:37 -07:00
Mai 3d65b701a2 Merge pull request #3816 from Sonicadvance1/fix_long_signed_divide
Arm64: Fixes long signed divide
2024-07-04 21:43:11 -04:00
Ryan Houdek ecaca0fe15 unittests: Adds x87 integer indefinite test
Tests 16-bit, 32-bit, and 64-bit integer conversions
2024-07-04 17:53:28 -07:00
Ryan Houdek 8955f83ef6 Softfloat: Fixes Integer indefinite return for 16-bit signed values
Regardless of positive or negative value, if the converted integer
doesn't fit in to the converted int16_t then it returns INT16_MIN.
2024-07-04 17:43:28 -07:00
Ryan Houdek 1a8aaebd79 unittests: Adds long signed divide test 2024-07-04 16:43:21 -07:00
Ryan Houdek 38a823cc54 Arm64: Fixes long signed divide
The two halves are provided as two uint64_t values that shouldn't be
sign extended between them. Treat them as uint64_t until combined in to
a single int128_t. Fixes long signed divide.
2024-07-04 16:42:23 -07:00
Ryan Houdek 25306cb373 InstcountCI: Update 2024-07-04 14:35:43 -07:00
Ryan Houdek 1084a031e7 unittests: Adds test for previous fix
All of these results would have failed except for the rorx result.
2024-07-04 14:35:43 -07:00
Ryan Houdek f6ec99bede OpcodeDispatcher: Fixes rotates with zero not zero extending 32-bit result
For all the 32-bit rotates (except for RORX) we were failing to zero
extend the 32-bit result to the destination register when the rotate was
masked to zero.

Ensure we do this.
2024-07-04 14:35:42 -07:00
Ryan Houdek 90a6647fa4 Merge pull request #3811 from alyssarosenzweig/ra/fix-lsp
RA: fix interaction between SRA & shuffles
2024-07-04 14:20:46 -07:00
Alyssa Rosenzweig a926bb81a9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig fbf41e3149 unittests: add test for small sbc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig a38205069b OpcodeDispatcher: fix SBB carry flag
do it the naive way, just applying the x86 definitions of SBB.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig 2d75801024 unittests: add tricky RA test
this fails on current main with blocksize=500 due to mentioned RA bug. passes
with blocksize=1.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 13:37:13 -04:00
Alyssa Rosenzweig 504511fe7e RA: fix interaction between SRA & shuffles
missed a Map. tricky case hit by the unit test added in the next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 13:37:13 -04:00
Ryan Houdek d3399a261b Docs: Update for release FEX-2407 2024-07-03 17:59:42 -07:00
Ryan Houdek d2437e6a21 Merge pull request #3810 from Sonicadvance1/x87_mmx_unittest
unittests: Adds MMX and x87 conflating unit test
2024-07-03 14:39:05 -07:00
Ryan Houdek 95dd6ceba8 unittests: Adds MMX and x87 conflating unit test
This failed with prior RCLSE deletion caching.
2024-07-03 13:54:07 -07:00
Alyssa Rosenzweig 1a0d135201 Merge pull request #3809 from alyssarosenzweig/rm/old-md
FEXCore: remove very out-of-date optimizer docs
2024-07-03 15:46:27 -04:00
Ryan Houdek f453e1523e Merge pull request #3803 from pmatos/NinjaCore
Use number of jobs as defined by TEST_JOB_COUNT
2024-07-03 12:42:14 -07:00
Alyssa Rosenzweig 622b0bfbc9 FEXCore: remove very out-of-date optimizer docs
most of this doesn't exist and won't exist. nothing lost here but hopes &
dreams.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-03 11:36:48 -04:00
Paulo Matos ad52514b97 Use number of jobs as defined by TEST_JOB_COUNT
At the moment we always run ctest with max number of cpus. If
undefined, it will keep current behaviour, otherwise it will
honour TEST_JOB_COUNT.

Therefore to run ctest one test at a time, use
`cmake ... -DTEST_JOB_COUNT=1`
2024-07-03 14:09:39 +02:00
Alyssa Rosenzweig 02a218c6e3 Merge pull request #3804 from Sonicadvance1/revert_rclse_drop
Revert removing RCLSE
2024-07-03 07:37:02 -04:00
Ryan Houdek 2d617ad173 InstcountCI: Update 2024-07-02 20:24:58 -07:00
Ryan Houdek 0d06e3e47d Revert "OpcodeDispatcher: add cache"
This reverts commit 46676ca376.
2024-07-02 20:24:57 -07:00
Ryan Houdek 78aee4d96e Revert "IR: drop RCLSE"
This reverts commit a5b24bfe4c.
2024-07-02 20:21:59 -07:00
Ryan Houdek ba04da87e5 Merge pull request #3780 from Sonicadvance1/optimize_gathers
Optimize gathers slightly
2024-07-02 10:58:38 -07:00
Ryan Houdek 2e6b08cbcb Merge pull request #3798 from Sonicadvance1/minor_128bit_vbsl_opt
Arm64: Minor VBSL optimization with SVE128
2024-07-01 18:57:46 -07:00
Ryan Houdek 472a373861 Merge pull request #3786 from Sonicadvance1/non_temporal_stores
OpcodeDispatcher: Implement support for non-temporal vector stores
2024-07-01 18:57:38 -07:00
Ryan Houdek a451420911 Merge pull request #3783 from Sonicadvance1/optimize_vector_zeroregister
OpcodeDispatcher: Optimize x86 canonical vector zero register
2024-07-01 18:57:31 -07:00
Mai 2e84f21c18 Merge pull request #3802 from Sonicadvance1/fix_sse41_helper
CodeEmitter: Fixes vector {ldr,str}{b,h} with reg-reg source
2024-07-01 20:42:49 -04:00
Ryan Houdek fb7167c2d2 CodeEmitter: Fixes vector {ldr,str}{b,h} with reg-reg source
We had failed to enable these implementations for the
`ExtendedMemOperand` helpers. We had already implemented the non-helper
forms, which are already tested in CI. These helpers just weren't
updated?

Noticed this when running libaom's SSE4.1 tests, where it managed to
execute a pmovzxbq instruction with reg+reg memory source and was
breaking the test results.

There are /very/ few vector register operations that access only 8-bit
or 16-bit in vectors so this flew under the radar for quite a while.

Fixes their unit tests.

Also adds a unittest using sse4.1 pmovzxbq to ensure we support the
reg+reg case, and also a few other instructions to test 8-bit and 16-bit
vector loads and stores.
2024-07-01 17:03:47 -07:00
Mai d884eb9287 Merge pull request #3801 from Sonicadvance1/fix_vpcmpgtw_typo
unittests: Fixes typo in vpcmpgtw test
2024-07-01 18:16:54 -04:00
Ryan Houdek 8b9b1a90e4 unittests: Fixes typo in vpcmpgtw test 2024-07-01 14:42:23 -07:00
Ryan Houdek e2d4010b59 Merge pull request #3800 from Sonicadvance1/fix_vmovlhps
AVX128: Fixes vmovlhps
2024-07-01 14:41:43 -07:00
Ryan Houdek babde31bf0 AVX128: Fixes vmovlhps
We didn't have a unit test for this and we weren't implementing it at
all.
We treated it as vmovhps/vmovhpd accidentally. Once again caught by the
libaom Intrinsics unit tests.
2024-07-01 13:54:11 -07:00
Ryan Houdek c282239077 InstcountCI: Add SVE128 VEX_map3 2024-06-30 16:27:58 -07:00
Ryan Houdek 8d28a441ab Arm64: Minor VBSL optimization with SVE128
This is a very minor performance change. On Cortex CPUs that support
SVE, they do movprfx+<instruction> fusion to remove two cycles and a
dependency from the backend.

This is a minor win to convert from ASIMD mov+bsl to SVE movprfx+bsl
because of this, saving two cycles and a dependency on Cortex A710 and
A715. This is slightly less of a win on Cortex-A720/A725 because it supports
zero-cycle vector register renames, but it is still a win on Cortex-X925
because that is an older core design that doesn't support zero-cycle
vector register renames.

Very silly little thing.
2024-06-30 16:22:29 -07:00
Ryan Houdek 5821054d91 Merge pull request #3789 from Sonicadvance1/avx128_minor_pshufb_opt
AVX128: Minor optimization to 256-bit vpshufb
2024-06-30 15:45:11 -07:00
Ryan Houdek 4626145374 Merge pull request #3792 from Sonicadvance1/avx128_fix_scalar_fma
AVX128: Fixes scalar FMA accidentally using vector wide
2024-06-30 15:36:09 -07:00
Ryan Houdek a786d3621d InstcountCI: Update for Scalar FMA 2024-06-30 14:36:56 -07:00
Ryan Houdek 1393dc2a5b AVX128: Fixes scalar FMA accidentally using vector wide 2024-06-30 14:36:33 -07:00
Ryan Houdek c4604465ba InstcountCI: Update 2024-06-30 13:41:14 -07:00
Ryan Houdek cffae9cb0f AVX128: Minor optimization to 256-bit vpshufb 2024-06-30 13:41:03 -07:00
Ryan Houdek cf24d3c33f Merge pull request #3781 from Sonicadvance1/optimize_vmovlh
AVX128: Minor optimization to vmov{l,h}{ps,pd}
2024-06-29 23:15:53 -07:00
Ryan Houdek 672e885e40 InstcountCI: Adds canonical zero register tests 2024-06-29 22:21:53 -07:00
Ryan Houdek 7d05610da7 OpcodeDispatcher: Optimize x86 canonical vector zero register
The canonical way to generate a zero register vector in x86 is to xor
itself. Capture this can convert it to canonical zero register instead.

Can get zero-cycle renamed on latest CPUs.
2024-06-29 22:21:53 -07:00
Ryan Houdek a843ecf4c8 InstcountCI: Update for non-temporal stores 2024-06-29 22:05:56 -07:00
Ryan Houdek f4ff1b0688 OpcodeDispatcher: Implement support for non-temporal vector stores
x86 doesn't have a lot of non-temporal vector stores but we do have a
few of them.

- MMX: MOVNTQ
- SSE2: MOVNTDQ, MOVNTPS, MOVNTPD
- AVX: VMOVNTDQ (128-bit & 256-bit), VMOVNTPD

Additionally SSE4a adds 32-bit and 64-bit scalar vector non-temporal
stores, which we keep as regular stores. Since ARM doesn't have matching
semantics for those.

Additionally SSE4.1 adds non-temporal vector LOADS which this doesn't
touch.
- SSE4.1: MOVNTDQA
- AVX: VMOVNTDQA (128-bit)
- AVX2: VMOVNTDQA (256-bit)

Fixes #3364
2024-06-29 22:05:56 -07:00
Ryan Houdek 2b4cec8385 Arm64: Implement support for non-temporal vector stores 2024-06-29 22:03:17 -07:00
Ryan Houdek 8ab4ab29f8 CodeEmitter: Add SVE contiguous non-temporal instructions 2024-06-29 21:51:58 -07:00
Ryan Houdek cc0509c0f3 InstcountCI: Update 2024-06-29 19:27:39 -07:00
Ryan Houdek ebfa65fedc AVX128: Minor optimization to vmov{l,h}{ps,pd} 2024-06-29 19:27:16 -07:00
Ryan Houdek a34ae24b3f InstcountCI: Update for SVE non-base address reg 2024-06-29 13:16:02 -07:00
Ryan Houdek 58ea76eb24 Arm64: Minor optimization to gather loads with no base addr register and SVE path
Arm64's SVE load instruction can be minorly optimized in the case that a
base GPR register isn't provided, as it has a version of the instruction
that doesn't require one.

The limitation of this instruction is that it doesn't support scaling at
all so it only works if the offset scale is 1.
2024-06-29 13:14:35 -07:00
Ryan Houdek e9a17b19c5 InstcountCI: Add SVE gathers without base addr 2024-06-29 13:07:32 -07:00
Ryan Houdek ce8d111453 InstcountCI: Update 2024-06-29 13:04:21 -07:00
Ryan Houdek 47fd73f6cf Arm64: Optimize non-SVE gather load
When FEX hits the optimal case that the destination isn't one of the
incoming sources (other than the incomingDest source) then we can
optimize out two moves per 128-bit lane.

Cuts 256-bit non-SVE gather loads from 50 instructions down to 46.
2024-06-29 13:02:10 -07:00
Ryan Houdek 76f3391ebc Merge pull request #3779 from Sonicadvance1/cpuinfo_cyclecounter
Linux: Calculate cycle counter frequency for cpuinfo
2024-06-29 11:58:32 -07:00
Ryan Houdek be6ff52709 Linux: Calculate cycle counter frequency for cpuinfo
Some applications don't measure rdtsc correctly and instead use cpuinfo
to get the CPU core's base clock speed. Which for most x86 CPUs is the
base clock speed which also matches their cycle counter speed.

Did this as a quick test to see if this would help `Unbound: Worlds
Apart` stuttering while BinaryNinja was disassembling the binary.

Turns out the game doesn't use cpuinfo for its cycle counter speed
determination, but it is good to implement this regardless.
2024-06-28 16:38:49 -07:00
Ryan Houdek e99e252188 Merge pull request #3731 from Sonicadvance1/avx_5
HostFeatures: Always disable AVX in 32-bit mode to protect from stack overflows
2024-06-28 13:37:55 -07:00
Ryan Houdek 98b980f7e3 TestHarnessRunner: Ensure we are still reconstructing XMM registers if we don't support AVX
Also fixes a bug where we were destroying the thread context before
reading the data from it, spooky.
2024-06-28 13:05:52 -07:00
Ryan Houdek f2f90eeb82 FEXCore: Make more distinctions between host register size and guest vector register size
We can support a few combinations of guest and host vector sizes
Host: 128-bit or 256-bit
Guest: 128-bit or 256-bit

The typical case is Host = 128-bit and Guest = 256-bit now that AVX is
implemented.
On 32-bit this changes to Host=128-bit and Guest=128-bit because we
disable AVX.

In the vixl simulator 32-bit turns in to Host=256-bit and Guest=128-bit.
And then in the vixl sim 64-bit turns in to Host=256-bit and
Guest=256-bit.

We cover all four combinations of guest and host vector register sizes!

Fixes a few assumptions that SVE256 = AVX256 basically.
2024-06-28 13:05:52 -07:00
Ryan Houdek f267fd2250 HostFeatures: Always disable AVX in 32-bit mode to protect from stack overflows 2024-06-28 13:05:52 -07:00
Ryan Houdek 500ad34769 Merge pull request #3778 from pmatos/LargeX87Blocks
Largest x87 blocks of code from games
2024-06-28 09:40:19 -07:00
Ryan Houdek 1700d54012 Merge pull request #3776 from Sonicadvance1/fix_vsib_invalid_index 2024-06-28 08:43:24 -07:00
Paulo Matos 70d8a10484 Largest x87 blocks of code from games 2024-06-28 16:50:58 +02:00
Ryan Houdek 9e94784e26 unittests: Adds test for xmm4 VSIB bug 2024-06-27 20:55:30 -07:00
Ryan Houdek 4060f4018e Frontend: Fixes invalid VSIB Index problem
In regular SIB land the index register encoding of 0b100 encodes to "no
register", this feature lets you get SIB encodings without an index
register for flexibility.

In VSIB encoding this isn't expected behaviour and instead there are no
encodings where an index register is missing. Allowing you to encode all
sixteen registers as an index register.

This was causing an abort in `AVX128_LoadVSIB` because the index turned
in to an invalid register.

Working instruction:
`vgatherdps ymm2, dword [eax+ymm5*4], ymm7`

Broken instruction:
`vgatherdps ymm0, dword [eax+ymm4*4], ymm7`

This fixes a crash in libfmod where it is using gathers in the wild.
Fixing a crash in Ender Lilies.
2024-06-27 20:55:30 -07:00
Ryan Houdek 739ac0f18f Merge pull request #3775 from Sonicadvance1/avx_bugfixes
AVX128: Some quick bugfixes
2024-06-27 17:44:12 -07:00
Ryan Houdek 98d62a7eb1 InstcountCI: Update 2024-06-27 17:21:12 -07:00
Ryan Houdek aba7a3a830 AVX128: Fixes vblendps lower and upper selector 2024-06-27 17:20:39 -07:00
Ryan Houdek 9027d1eee7 AVX128: Fixes bug in vector immediate shift 2024-06-27 16:22:14 -07:00
Ryan Houdek 4e5da4946d Merge pull request #3773 from bylaws/win-fixes
Windows: Small fixes for compat with newer toolchains/wine versions
2024-06-27 15:14:20 -07:00
Billy Laws a70e3e42b2 FEXCore: Drop unneeded MinGW library naming workaround
It's generally expected for libraries to use the .a suffix with MinGW,
and DLLs are still correctly named without the prior special handling.
2024-06-27 23:01:21 +01:00
Billy Laws 09f476924f FEXCore: Fix missing return in win32 SetSignalMask path 2024-06-27 23:01:21 +01:00
Billy Laws 230e3245fd FileLoading: Fix compilation with newer libc++ 2024-06-27 23:01:21 +01:00
Billy Laws 8de876daf2 Windows: Use newer wine unixcall API
__wine_unix_call is no longer exported in recent wine versions.
2024-06-27 23:01:19 +01:00
Ryan Houdek 53b1d155cc Merge pull request #3772 from Sonicadvance1/fix_addrsize_override
FEXCore: Fixes address size override on GPR sources and destinations
2024-06-27 15:01:08 -07:00
Ryan Houdek b0eb63ab9a FEXCore: Fixes address size override on GPR sources and destinations
When the source or destination is a register, the address size override
doesn't apply. We were accidentally applying it on all sources
regardless of type which was causing us to zero extend on operations
that aren't affected by address size override.

This fixes the OpenSSL cert error in every application, but most
importantly Steam.
2024-06-27 14:12:01 -07:00
Ryan Houdek 2e3242682d Merge pull request #3771 from alyssarosenzweig/opt/asimd-masked
OpcodeDispatcher: optimize nzcv with asimd masked load/store
2024-06-27 10:27:10 -07:00
Ryan Houdek ad4d4c9e67 Merge pull request #3770 from alyssarosenzweig/opt/vzeroall
Tiny opt for vzeroall
2024-06-27 10:25:35 -07:00
Alyssa Rosenzweig 3250d4e405 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-27 10:37:11 -04:00
Alyssa Rosenzweig 196a0531e0 OpcodeDispatcher: optimize nzcv with asimd masked load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-27 10:37:06 -04:00
Alyssa Rosenzweig e61cb5b2c3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-27 10:30:45 -04:00
Alyssa Rosenzweig f9b53c6b51 AVX_128: save a move in vzeroall
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-27 10:30:25 -04:00
Mai 58e949e148 Merge pull request #3769 from Sonicadvance1/avx2_cpuid
CPUID: Oops, forgot to enable AVX2
2024-06-26 21:17:44 -04:00
Ryan Houdek dad47b7bda CPUID: Oops, forgot to enable AVX2 2024-06-26 17:43:56 -07:00
Ryan Houdek e519bf5978 Merge pull request #3768 from Sonicadvance1/avx128_letsgo
AVX128: Enable all the things
2024-06-26 17:40:21 -07:00
Ryan Houdek fc50e52157 InstCountCI: Adds AVX128 tests 2024-06-26 16:49:00 -07:00
Ryan Houdek 7669df0e16 InstCountCI: SVE256: Fixes behaviour change 2024-06-26 16:49:00 -07:00
Ryan Houdek 4d56fec5f1 AVX128: Work around glibc fault testing 2024-06-26 16:49:00 -07:00
Ryan Houdek 8181552b16 AVX128: Actually install AVX helpers per thread.
How this didn't break the world in my testing I don't know.
2024-06-26 16:49:00 -07:00
Ryan Houdek c6c147daf6 unittests: Updates vcvtps2ph test for failure case of writing too much memory. 2024-06-26 16:49:00 -07:00
Ryan Houdek 975069825e AVX128: Fix a real bug with VCVTPS2PH 2024-06-26 16:49:00 -07:00
Ryan Houdek 5133f480d1 InstcountCI: Update for xsave/xrstor behaviour changes with AVX 2024-06-26 16:49:00 -07:00
Ryan Houdek ce4b252e5c InstCountCI: Stop disabling AVX if SVE256 is disabled. 2024-06-26 15:06:03 -07:00
Ryan Houdek 031d56de35 HostFeatures: Enables AVX unconditionally 2024-06-26 15:03:21 -07:00
Ryan Houdek 3cdaf6736b InstcountCI: Update for SVE256 FMA implementation 2024-06-26 14:56:01 -07:00
Ryan Houdek b5e696b3cb CPUID: Implement support for XCR0 when AVX is enabled
This enables AVX, AVX2, FMA3 for the entire CPUID!

```bash
$ FEX_HOSTFEATURES=enableavx,enableavx2 ./Bin/FEXInterpreter /usr/bin/cat /proc/cpuinfo
processor       : 0
vendor_id       : GenuineIntel
cpu family      : 6
model           : 23
model name      : Cortex-A78AE
stepping        : 0
microcode       : 0x0
cpu MHz         : 3000
cache size      : 512 KB
physical id     : 0
siblings        : 12
core id         : 0
cpu cores       : 12
apicid          : 0
initial apicid  : 0
fpu             : yes
fpu_exception   : yes
cpuid level     : 22
wp              : yes
flags           : fpu vme tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht tm syscall nx mmxext fxsr_opt rdtscp lm 3dnow 3dnowext constant_tsc art rep_good nopl xtoplogy nonstop_tsc cpuid tsc_known_freq pni pclmulqdq dtes64 monitor tm2 ssse3 fma cx16 sse4_1 sse4_2 movbe popcnt aes xsave avx hypervisor lahf_lm cmp_legacy extapic abm 3dnowprefetc
h tce fsgsbase bmi1 avx2 smep bmi2 erms invpcid adx clflushopt clwb sha_ni clzero arat vpclmulqdq rdpid fsrm
bugs            :
bogomips        : 8000.0
TLB size        : 2560 4K pages
clflush size    : 64
cache_alignment  : 64
address sizes   : 40 bits physical, 48 bits virtual
power management:
```

Notice avx, avx2, and fma
2024-06-26 14:56:01 -07:00
Ryan Houdek 43aef377d7 HostFeatures: Allow enabling AVX without SVE256 2024-06-26 14:56:01 -07:00
Ryan Houdek add0e7a8db HostFeatures: Removes distinction between AVX and AVX2
We now no longer care about AVX versions, consolidate them in to a
single config option which enables both.
2024-06-26 14:56:01 -07:00
Ryan Houdek 52e541d453 Unittests: Stop using AVX2 flag 2024-06-26 14:56:01 -07:00
Mai a031a49546 Merge pull request #3767 from Sonicadvance1/avx128_fix_wide_shift
AVX128: Fixes wide shifts
2024-06-26 17:29:09 -04:00
Alyssa Rosenzweig 4d821b8dd8 Merge pull request #3765 from Sonicadvance1/avx128_f16c
AVX128: F16C support
2024-06-26 17:25:05 -04:00
Ryan Houdek f277025c9a AVX128: Fixes wide shifts
During refactoring this was missed and rerunning unittests locally
caught it. 256-bit operations get their shift only from the lower half
of the vector register.
2024-06-26 14:16:39 -07:00
Ryan Houdek ba28e6f82e unittests: Adds vcvtps2ph tests that use mxcsr 2024-06-26 14:08:20 -07:00
Ryan Houdek 3a89df9bed AVX128: Implement support for F16C 2024-06-26 14:05:12 -07:00
Ryan Houdek f6a0866fbb IR: Split Vector_FToF2 in to VFCVTL2 and VCVTFN2
I forgot in the narrowing case we need to be careful about insert. No IR
op used Vector_FToF2 with narrowing.
2024-06-26 14:03:41 -07:00
Ryan Houdek 756fa2ecc5 Merge pull request #3766 from alyssarosenzweig/opt/f16c-round
Optimize vcvtps2ph
2024-06-26 14:03:24 -07:00
Alyssa Rosenzweig cf834aa6da InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 16:46:21 -04:00
Alyssa Rosenzweig d2324f4a93 OpcodeDispatcher: optimize vcvtps2ph
We can avoid a LOT of pointless work with some dedicated IR ops for specifically
overriding the round mode.

Small behaviour change here: we no longer reset FTZ. I think this is a bug fix?
But if it's not it's not hard to fix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 16:46:21 -04:00
Ryan Houdek 6226c7f4f3 Merge pull request #3757 from Sonicadvance1/avx_16
AVX128: Implement support for gathers
2024-06-26 13:29:58 -07:00
Ryan Houdek 991ecd558e InstcountCI: Update for SVE256 gathers! 2024-06-26 16:00:53 -04:00
Ryan Houdek a4fa3a460e OpcodeDispatcher: Implement AVX gathers with SVE256
Just to ensure we still have feature parity.
2024-06-26 16:00:53 -04:00
Ryan Houdek 77ba708933 AVX128: Implement support for gather load instructions
This is the last family of instructions that we needed to implement for
AVX2 to be properly advertised!
2024-06-26 16:00:53 -04:00
Ryan Houdek 662d50a966 X86Tables: Describe VPGather in the VEX tables 2024-06-26 16:00:53 -04:00
Ryan Houdek 5472d1cc04 Arm64: Implement VLoadVectorGatherMasked operation
This does a gather load three ways, SVE256, SVE128, and ASIMD.

This operation is a bit special since it it can't quite handle all
gather loadstores in the 256-bit case and requires the frontend to
decompose the operation in the case that the striding hits a mode that
SVE doesn't support!

The 128-bit case is a lot simpler since both support all the cases where
stride doesn't match. I find this to be a nice compromise while there
aren't any SVE256 products on the market.

In the 128-bit case there is an SVE path which is utilized if the passed
in stride supports what SVE understands, otherwise it falls back to an
ASIMD implementation which manually emulates everything that is
necessary.

This instruction is very explicitly doing basically exactly what AVX
gather instructions want, because it's complex enough that we don't want
to try and make this a generic solution.
2024-06-26 16:00:53 -04:00
Alyssa Rosenzweig d1d41f5645 Merge pull request #3763 from alyssarosenzweig/rclse/less-aggressive
Remove RCLSE
2024-06-26 15:14:14 -04:00
Ryan Houdek 94fd100fc7 Merge pull request #3719 from lioncash/f16c
OpcodeDispatcher: Handle F16C operations
2024-06-26 12:12:13 -07:00
Lioncache b9ff36b5d9 CPUID: Signify F16C support if AVX is available
On Aarch64 hardware, if we have SVE2 available (which we use in the AVX implementation),
then we can also enable F16C support.
2024-06-26 15:05:03 -04:00
Lioncache cd5a809ec9 OpcodeDispatcher: Handle VCVTPS2PH 2024-06-26 15:05:03 -04:00
Lioncache 045a8efbeb OpcodeDispatcher: Handle VCVTPH2PS
Fairly straightforward, since we already have handling for half-float conversions.
2024-06-26 15:05:00 -04:00
Ryan Houdek 54a1f7d833 Merge pull request #3764 from Sonicadvance1/rorx_masking
BMI2: Ensure rorx immediate masks by operation size correctly.
2024-06-26 11:52:47 -07:00
Alyssa Rosenzweig 1b496cda8f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 14:49:58 -04:00
Alyssa Rosenzweig a5b24bfe4c IR: drop RCLSE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 14:49:05 -04:00
Alyssa Rosenzweig 46676ca376 OpcodeDispatcher: add cache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 14:49:05 -04:00
Alyssa Rosenzweig 7d939a3b3d Merge pull request #3758 from Sonicadvance1/avx_17
AVX128: FMA3
2024-06-26 14:18:32 -04:00
Ryan Houdek a515061465 BMI2: Ensure rorx immediate masks by operation size correctly. 2024-06-26 11:11:37 -07:00
Ryan Houdek 1c24d63f73 Merge pull request #3762 from alyssarosenzweig/bug/constprop-bextr 2024-06-26 09:28:18 -07:00
Alyssa Rosenzweig 7e10dba5e2 unittests: add test for a BEXTR bug
Ryan reduced this test while debugging openssl. This fails without the constprop
fix.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 12:06:47 -04:00
Alyssa Rosenzweig e2d73014f1 ConstProp: fix LSHR constant prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-26 12:06:47 -04:00
Ryan Houdek 53aa30596e InstcountCI: Update 2024-06-25 11:37:18 -07:00
Ryan Houdek 122ae5b710 unittests: Adds FMA3 unittests 2024-06-25 11:37:18 -07:00
Ryan Houdek 45c27b2965 CPUID: Enable support for FMA3 when AVX is enabled 2024-06-25 11:24:53 -07:00
Ryan Houdek 832b247fc1 SVE258: Implement support for FMA3 2024-06-25 11:24:46 -07:00
Ryan Houdek 0e8b53d566 AVX128: Implement FMA3 instructions 2024-06-25 11:23:50 -07:00
Ryan Houdek d03d69273b X86Tables: Describe FMA3 instructions 2024-06-25 11:22:27 -07:00
Ryan Houdek efa05ba19d IR: Adds support for new SUBADD FMA constants
ADDSUB didn't cover this new variant.
2024-06-25 11:22:22 -07:00
Ryan Houdek 5da205d91a Merge pull request #3760 from alyssarosenzweig/avx/vpclmulqdql
AVX128: fix VPCLMULQDQl
2024-06-25 10:31:52 -07:00
Ryan Houdek 41923bac99 OpcodeDispatcher: Fixes PCMUL with weird selectors and zero-extend
We had a bug where we weren't correctly ignoring the non-used bits in
the selector. This was causing an assert in the ARM backend.
2024-06-25 12:54:03 -04:00
Alyssa Rosenzweig c6148f6bf1 AVX128: fix VPCLMULQDQl
use the helper. I assumed the lack of zero extension here was intentional.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 12:51:21 -04:00
Alyssa Rosenzweig 77aaa9af4d Merge pull request #3748 from Sonicadvance1/avx_15
AVX128: More instructions Part 4
2024-06-25 12:39:48 -04:00
Ryan Houdek 00cf8d530c Merge pull request #3752 from Sonicadvance1/fma_ir_operations
ARM64: Adds new FMA vector instructions
2024-06-25 09:07:06 -07:00
Alyssa Rosenzweig 98aa58e9f5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 10:03:33 -04:00
Ryan Houdek 6911917819 Disable vpclmulqdq_256 on simulator 2024-06-25 10:03:33 -04:00
Ryan Houdek a8255aa475 CPUID: Expose support for VPCLMULQDQ
Wasn't exposed before since we couldn't unit test the SVE256
implementation.
2024-06-25 10:03:33 -04:00
Ryan Houdek 48e7aae38f unittests: Adds support for 256-bit vpclmulqdq
It's easy because the test was already written for this in mind.
2024-06-25 10:03:33 -04:00
Ryan Houdek 7069643ae6 AVX128: Implement support for VPCLMULQDQ
This is just the 128-bit version twice.
2024-06-25 10:03:33 -04:00
Ryan Houdek 34272fc134 AVX128: Implement support for vperm{d,ps}! 2024-06-25 10:03:33 -04:00
Ryan Houdek 1d41002dfe AVX128: Implement support for variable vpermil{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek 563bf342d5 AVX128: Implement support for vptest 2024-06-25 10:03:33 -04:00
Ryan Houdek efd5fabb95 AVX128: Implement support for vtest{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek c1da525110 AVX128: Implement support for vperm2{f128,i128} 2024-06-25 10:03:33 -04:00
Ryan Houdek 5ce6c88a88 AVX128: Reenable {ldm,stm}mxcsr. Can use the regular implementation. 2024-06-25 10:03:33 -04:00
Ryan Houdek 64cce7c6fa AVX128: Implement support for xsave/xrstor 2024-06-25 10:03:33 -04:00
Ryan Houdek 4544e5b51f AVX128: Implement support for vblend{ps,pd}/vpblendvb 2024-06-25 10:03:33 -04:00
Ryan Houdek eb3e314946 AVX128: Implement support for vmaskmovdqu 2024-06-25 10:03:33 -04:00
Ryan Houdek 8b65c3de10 AVX128: Implement vmaskmov{ps,pd}, vpmaskmov{d,q} using SVE2 gather loadstores. 2024-06-25 10:03:33 -04:00
Ryan Houdek c2beb27a9d AVX128: Implement support for vpalignr 2024-06-25 10:03:33 -04:00
Ryan Houdek 05fdec9e72 AVX128: Implement support for vmpsadbw 2024-06-25 10:03:33 -04:00
Ryan Houdek e8e3c95349 AVX128: Implement support for vpsadbw 2024-06-25 10:03:33 -04:00
Ryan Houdek a87fa3f246 AVX128: Implement support for vpshufb 2024-06-25 10:03:33 -04:00
Ryan Houdek b31ad523f5 AVX128: Implement support for hsub{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek 34bce540ff AVX128: Implement support for vpblendw/vpblendd/vblendps/vblendpd 2024-06-25 10:03:33 -04:00
Ryan Houdek 8ea38e1d80 AVX128: Implement support for vpmaddwd 2024-06-25 10:03:33 -04:00
Ryan Houdek ce591a9541 AVX128: Implement support for vpmaddubsw 2024-06-25 10:03:33 -04:00
Ryan Houdek a48c65cd65 AVX128: Implement support for vphaddsw 2024-06-25 10:03:33 -04:00
Ryan Houdek c283f80f48 AVX128: Implement support for vhaddpd/vphadd{w,d} 2024-06-25 10:03:33 -04:00
Ryan Houdek d6bf276b5a AVX128: Implement support for imm vpermil{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek 96a51650b1 AVX128: Implement support for vshuf{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek f35a9c74a2 AVX128: Implement support for vpshuf{lw,hw,d} 2024-06-25 10:03:33 -04:00
Ryan Houdek e2457943f5 AVX128: Implement support for vperm{q,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek a05644172a AVX128: Implement support for vdd{ps,pd} 2024-06-25 10:03:33 -04:00
Ryan Houdek cc168ce0fb VectorOps: Restructure DPPOpImpl. This will get reused by AVX128 2024-06-25 10:03:33 -04:00
Alyssa Rosenzweig 76bd22d279 OpcodeDispatcher: rm gratuitous lambda
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 10:03:33 -04:00
Alyssa Rosenzweig 18574f3cf1 OpcodeDispatcher: extract VPERMILRegOpImpl
for avx

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 10:03:33 -04:00
Alyssa Rosenzweig 665215ab47 OpcodeDispatcher: extract PTestOpImpl
for avx128

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 09:52:48 -04:00
Alyssa Rosenzweig 6009f36403 OpcodeDispatcher: extract VPERMDIndices
and rename things accordingly.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 09:46:17 -04:00
Alyssa Rosenzweig 2580efda0d OpcodeDispatcher: tweak VTestOp signature
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-25 09:46:17 -04:00
Ryan Houdek 3a310b8815 Merge pull request #3756 from Sonicadvance1/fix_vmovhlps
Fix VMOVLHPS instruction
2024-06-24 19:14:56 -07:00
Ryan Houdek 7ff96227c0 Merge pull request #3755 from Sonicadvance1/fix_avx128_vmovntdqa
AVX128: Fix vmovntdqa failing to zero upper 128-bits
2024-06-24 19:14:48 -07:00
Ryan Houdek 3e8d78051c InstcountCI: Update 2024-06-24 17:26:18 -07:00
Ryan Houdek bd24ebc96a unittests: Adds VMOVHLPS unit test
A bit confusing because the instruction encoding is the same between
VMOVHLPS and VMOVLPS so this unittest was missed.

Implement the test to ensure it stays working
2024-06-24 17:22:55 -07:00
Ryan Houdek 6d3745b8f1 AVX128: Fixes VMOVLHPS instruction
We didn't have unit tests for this
2024-06-24 17:22:51 -07:00
Ryan Houdek d0f0b975be SVE256: Fixes VMOVLHPS instruction
We didn't have unit tests for this
2024-06-24 17:22:47 -07:00
Ryan Houdek ff2e6ed59f X86Tables: Fixes instruction encoding for VMOVLP{S,D}
These can have both register and memory modrm encoding
2024-06-24 17:22:43 -07:00
Ryan Houdek 99b2018d0e unittests: Extend vmovntpd test 2024-06-24 16:32:13 -07:00
Ryan Houdek f0d9c8c10a AVX128: Fix vmovntdqa failing to zero upper 128-bits 2024-06-24 16:32:09 -07:00
Ryan Houdek dce1b24c00 Merge pull request #3754 from Sonicadvance1/fix_avx128_stringops
AVX128: Fixes SSE4.2 string compare instructions
2024-06-24 16:30:42 -07:00
Ryan Houdek b47e981932 AVX128: Fixes SSE4.2 string compare instructions 2024-06-24 15:54:06 -07:00
Ryan Houdek dc44eb4caf Merge pull request #3749 from Sonicadvance1/contigous_mask_optimization_removal
Arm64: Remove contiguous masked element optimization
2024-06-24 15:22:23 -07:00
Ryan Houdek dfda6733f0 Merge pull request #3750 from Sonicadvance1/pshuf_bug
OpcodeDispatcher: Fixes bug in pshuf{lw,hw}
2024-06-24 15:22:07 -07:00
Alyssa Rosenzweig 21c6986dc7 OpcodeDispatcher: tweak HSUBPOpImpl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 18:21:08 -04:00
Alyssa Rosenzweig 635720fe12 OpcodeDispatcher: tweak PHADDSOpImpl signature
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 18:21:08 -04:00
Ryan Houdek 8c751d7423 Merge pull request #3747 from Sonicadvance1/avx_14
AVX128: More instructions Part 3
2024-06-24 12:35:22 -07:00
Ryan Houdek 702ecf7637 AVX128: Implement support for round{ss,sd} 2024-06-24 15:19:08 -04:00
Ryan Houdek 0595f1e044 AVX128: Implement support for vround{ps,pd} 2024-06-24 15:19:08 -04:00
Ryan Houdek cebb032bd3 AVX128: Implement support for vphminposuw
Reuses the non-AVX implementation since it only operates on 128-bits.
2024-06-24 15:19:08 -04:00
Ryan Houdek 8e32763ada AVX128: Implements support for AVX string ops
Reuses the implementation from the SSE4.2 implementation, just
explicitly zeroes the hardcoded YMM0's upper 128-bits.
2024-06-24 15:19:08 -04:00
Ryan Houdek 7532337231 AVX128: Implements support for vector AES instructions 2024-06-24 15:19:08 -04:00
Ryan Houdek 4a66d4570e AVX128: Implement support for a trinary operation with a passed in vector
Will be used for AES operations
2024-06-24 15:19:08 -04:00
Alyssa Rosenzweig 6f5e99d47d OpcodeDispatcher: factor out TranslateRoundType
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 15:19:08 -04:00
Alyssa Rosenzweig 9ee9f5bddd OpcodeDispatcher: tweak VectorRoundImpl signature
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 15:14:55 -04:00
Ryan Houdek ddb9f6d3ad Merge pull request #3746 from Sonicadvance1/avx_13
AVX128: More instructions
2024-06-24 11:40:52 -07:00
Ryan Houdek d29139d88a AVX128: Implement support for vextract{i,f}128 2024-06-24 14:27:19 -04:00
Ryan Houdek 317575ba99 AVX128: Implement support for cvtdq2{ps,pd} 2024-06-24 14:27:19 -04:00
Ryan Houdek d4f2638a2e AVX128: Implement support for cvt{t,}pd2pq 2024-06-24 14:27:19 -04:00
Ryan Houdek b67d9be227 AVX128: Implement support for vcvt{pd2ps,ps2pd}
Fairly complex set of instructions due to the edge cases.
2024-06-24 14:27:19 -04:00
Ryan Houdek d52add8fad AVX128: Implement support for vcvt{ss2sd,sd2ss} 2024-06-24 14:27:19 -04:00
Ryan Houdek aa9159d25c AVX128: Implement support for vpmulh{u,}w 2024-06-24 14:27:19 -04:00
Ryan Houdek 94c777259e AVX128: Implements support for vpmulhrsw 2024-06-24 14:27:19 -04:00
Ryan Houdek c9f8fa5662 AVX128: Implement support for vpmul{u,}dq 2024-06-24 14:27:19 -04:00
Ryan Houdek 64ee6b119e AVX128: Implement support for vaddsubp{s,d} 2024-06-24 14:27:19 -04:00
Ryan Houdek d2ec9a8936 AVX128: Implement support for vpsubsw 2024-06-24 14:27:19 -04:00
Ryan Houdek 2a927453f7 AVX128: Implement support for vphsub{w,d} 2024-06-24 14:27:19 -04:00
Ryan Houdek c19d489c9a AVX128: Implement support for vinsertps
This one actually reuses the core base implementation which is nice.
2024-06-24 14:27:19 -04:00
Ryan Houdek 6012eb051b AVX128: Implement support for vinsert{f128,i128} 2024-06-24 14:27:19 -04:00
Alyssa Rosenzweig 3974746473 OpcodeDispatcher: tweak PHSUBOpImpl
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig e1bcdcf387 OpcodeDispatcher: tweak PHSUBSOpImpl signature
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig fd5fbddae9 OpcodeDispatcher: tweak PMULLOpImpl for avx128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig 8ff72beddb OpcodeDispatcher: tweak PMULHRSWOpImpl signature for avx128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig cba5f7877b OpcodeDispatcher: tweak ADDSUBPOpImpl signature for AVX128
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Alyssa Rosenzweig 9d7e9fd9fc OpcodeDispatcher: add AVX128_Zext helper
should let us clean up a lot.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-24 14:14:23 -04:00
Ryan Houdek 082a0baff3 JIT: Implement missing Vector_FToF2 2024-06-24 14:14:23 -04:00
Ryan Houdek 3a4914315b Arm64: Remove contiguous masked element optimization
This was a premature optimization and currently breaks. Just remove it
for now.
2024-06-24 07:49:00 -07:00
Ryan Houdek 448b5a338a ARM64: Adds new FMA vector instructions 2024-06-24 07:48:05 -07:00
Ryan Houdek 9b68617fa8 InstCountCI: Update for pshuf fixes 2024-06-24 07:44:21 -07:00
Ryan Houdek 4c9890d7f8 OpcodeDispatcher: Fixes bug in pshuf{lw,hw}
This optimization was incorrect. Updates unittests to ensure it keeps
working.
2024-06-24 07:43:48 -07:00
Ryan Houdek b2db04f5d7 Merge pull request #3745 from Sonicadvance1/add_x86_cmake_assert_back
Adds back cmake error on x86-64 hosts
2024-06-24 06:44:33 -07:00
Alyssa Rosenzweig be8ff9ccb9 Merge pull request #3740 from Sonicadvance1/avx_12
AVX128: More various instructions
2024-06-24 09:28:40 -04:00
Ryan Houdek 9c531d97b0 AVX128: Implements the various vector shift instructions
These are very closely related to each other so it makes sense to
implement the roughly three different families in one commit.
2024-06-24 09:20:19 -04:00
Ryan Houdek 055d8d75a2 Revert "CI: Drop use of obsolete ENABLE_X86_HOST_DEBUG setting"
This reverts commit a054b998c5.
2024-06-24 06:05:19 -07:00
Ryan Houdek 9fcf79ce0e Adds back cmake error on x86-64 hosts 2024-06-24 06:05:19 -07:00
Ryan Houdek 6edf4619d4 Merge pull request #3742 from Sonicadvance1/export_avx_reg_helpers
FEXCore: Implement AVX reconstruction helpers
2024-06-24 05:57:44 -07:00
Ryan Houdek 8f769ce5a3 Merge pull request #3743 from alyssarosenzweig/cleanup/literal
X86Tables: add Literal() helper
2024-06-23 13:45:42 -07:00
Ryan Houdek 96ac71750a Wow64: Use SSE register reconstruction helpers
It doesn't support AVX today but it should do in the future.
2024-06-21 17:13:56 -04:00
Ryan Houdek d0852cf1bb TestHarnessRunner: Reconverge YMM registers if AVX is supported
The TestHarness infrastructure doesn't understand the difference between
converged versus split view.

So fetch the split view immediately and reconverge the view manually
inside of the state object so it continues working with the split ymm
view.
2024-06-21 17:13:56 -04:00
Ryan Houdek f5fea8af96 SignalDelegator: Use new YMM register reconstruction helpers
Otherwise we would be setting up signal handlers with incorrect register
state.
2024-06-21 17:13:56 -04:00
Ryan Houdek d52a1da501 FEXCore: Implement support for fetching/setting YMM registers
Because we have two views of the YMM registers depending on if the host
supports SVE256 or not, add helper functions to fetch them correctly.

We fetch them in the way that Linux desires them in signal handlers, if
we want to return the converged view directly, that is easy to add
support for. It's unnecessary for now.
2024-06-21 17:13:56 -04:00
Ryan Houdek abdcaa7c86 AVX128: Implement support for vpinsr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek ad122cf463 AVX128: Implement support for vpmovmskb 2024-06-21 15:53:52 -04:00
Ryan Houdek b58a57d225 AVX128: Implement support for vmovmskp{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek 28d679de98 AVX128: Implement support for vpmov{s,z}{b,w,d}{w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek d1dd055e6a AVX128: Implement support for vpextr{b,w,d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 3045578da4 AVX128: Implement vmov{d,q} 2024-06-21 15:53:52 -04:00
Ryan Houdek 9566dda73e AVX128: Implement support for vcmps{s,d} 2024-06-21 15:53:52 -04:00
Ryan Houdek a0ced2b685 AVX128: Implement support for vcmpp{s,d} 2024-06-21 15:50:26 -04:00
Ryan Houdek df232f567b AVX128: Implement support for v{add,sub,mul,fmin,fmax,fdiv,sqrt,rsqrt,rcp}s{s,d} 2024-06-21 15:50:05 -04:00
Ryan Houdek 2a6d6a9d13 AVX128: Implement support for v{u,}comis{s,d} 2024-06-21 15:50:05 -04:00
Alyssa Rosenzweig cd03932bd1 OpcodeDispatcher: tweak InsertScalarFCMPOpImpl signature
so AVX128 can reuse it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 15:50:05 -04:00
Ryan Houdek 2e5fa1ef1b Merge pull request #3739 from Sonicadvance1/avx_11
Frontend: Expose AVX W flag
2024-06-21 12:15:14 -07:00
Alyssa Rosenzweig 0c6c4cd532 OpcodeDispatcher: make FCMP more compact
I told Ryan to change this for AVX, but it needs to be changed in the original
to match!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 15:02:10 -04:00
Alyssa Rosenzweig 25f8a87429 OpcodeDispatcher: use Literal() helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:58:49 -04:00
Alyssa Rosenzweig edf1a7970d X86Tables: add Literal() helper
Any time we get the value of Literal, we want to assert that it's actually a
literal. We've been open coding this pattern sporadically throughout the
opcodedispatcher. Let's add an ergonomic helper to fetch the value of literal,
asserting that the value is indeed literal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:46:46 -04:00
Ryan Houdek fac9972bad Merge pull request #3741 from alyssarosenzweig/cleanup/comiss
OpcodeDispatcher: refactor Comiss helper
2024-06-21 11:43:05 -07:00
Alyssa Rosenzweig 9ecb960f3a OpcodeDispatcher: refactor Comiss helper
AVX128 will use this, it's not SSE-specific.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-21 14:23:20 -04:00
Ryan Houdek 3d26e23891 Merge pull request #3737 from Sonicadvance1/avx_10
Arm64: Implement support for emulated masked vector loadstores
2024-06-21 11:04:01 -07:00
Ryan Houdek 7bbbd95775 Merge pull request #3736 from Sonicadvance1/avx_9
AVX128: Some pun pickles, moves and conversions
2024-06-21 10:55:19 -07:00
Ryan Houdek bb308899b9 Frontend: Expose AVX W flag
Previously we could always tell the size of the operation depending on
how this effects the operating size of the instruction. Converting
64-bit down to 32-bit as an example.

AVX gather instructions are the first instruction class that can't infer
this information. The element load size is determined by the W flag but
the operating size of 128-bit or 256-bit is determined by other means.

Expose this flag so we can determine this difference. The FMA
instructions are going to need this flag as well.
2024-06-21 10:54:42 -07:00
Ryan Houdek e95c8d703c Arm64: Implement support for emulated masked vector loadstores
In order to support `vmaskmov{ps,pd}` without SVE128 this is required.
It's pretty gnarly but they aren't often used so that's fine from a
compatibility perspective.

Example SVE128 implementation:
```json
    "vmaskmovps ymm0, ymm1, [rax]": {
      "ExpectedInstructionCount": 9,
      "Comment": [
        "Map 2 0b01 0x2c 256-bit"
      ],
      "ExpectedArm64ASM": [
        "ldr q2, [x28, #32]",
        "mrs x20, nzcv",
        "cmplt p0.s, p6/z, z17.s, #0",
        "ld1w {z16.s}, p0/z, [x4]",
        "add x21, x4, #0x10 (16)",
        "cmplt p0.s, p6/z, z2.s, #0",
        "ld1w {z2.s}, p0/z, [x21]",
        "str q2, [x28, #16]",
        "msr nzcv, x20"
      ]
    },
```

Example ASIMD implementation
```json
    "vmaskmovps ymm0, ymm1, [rax]": {
      "ExpectedInstructionCount": 37,
      "Comment": [
        "Map 2 0b01 0x2c 256-bit"
      ],
      "ExpectedArm64ASM": [
        "ldr q2, [x28, #32]",
        "mrs x20, nzcv",
        "movi v0.2d, #0x0",
        "mov x1, x4",
        "mov x0, v17.d[0]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[0], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[1], [x1]",
        "add x1, x1, #0x4 (4)",
        "mov x0, v17.d[1]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[2], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[3], [x1]",
        "mov v16.16b, v0.16b",
        "add x21, x4, #0x10 (16)",
        "movi v0.2d, #0x0",
        "mov x1, x21",
        "mov x0, v2.d[0]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[0], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[1], [x1]",
        "add x1, x1, #0x4 (4)",
        "mov x0, v2.d[1]",
        "tbz x0, #63, #+0x8",
        "ld1 {v0.s}[2], [x1]",
        "add x1, x1, #0x4 (4)",
        "tbz w0, #31, #+0x8",
        "ld1 {v0.s}[3], [x1]",
        "mov v2.16b, v0.16b",
        "str q2, [x28, #16]",
        "msr nzcv, x20"
      ]
    },
```

There's a little bit of an improvement where nzcv isn't needed to get
touched on the ASIMD implementation, but I'll leave that for a future
improvement.
2024-06-21 08:21:32 -07:00
Ryan Houdek 903d6a742e CPUBackend: Removes SupportsSaturatingRoundingShifts option
This has always been true ever since we removed the x86 JIT and
Interpreter. This was left over and adding more code for no reason.
2024-06-21 08:11:22 -07:00
Ryan Houdek 424218e327 AVX128: Implement support for vpsign{b,w,d} 2024-06-21 08:11:22 -07:00
Ryan Houdek 17dc03d414 AVX128: Implement support for vpack{s,u}{wb,dw} 2024-06-21 08:11:21 -07:00
Ryan Houdek baf699c6e1 AVX128: Implements support for vandnps and vpandn
This can't use the previous binary operator handler since the register
sources need to be swapped.
2024-06-21 08:11:21 -07:00
Ryan Houdek 1431af1ff5 AVX128: Implements support for vcvt{t,}s{s,d}2si 2024-06-21 08:11:21 -07:00
Ryan Houdek 775a41b903 AVX128: Implement support for vcvtsi2s{s,d} 2024-06-21 08:11:21 -07:00
Ryan Houdek 2da1e90dd5 Merge pull request #3738 from Sonicadvance1/cpuid_label
CPUID: Update labeling on some reserved bits
2024-06-21 07:41:57 -07:00
Ryan Houdek e614340c0c CPUID: Update labeling on some reserved bits
These aren't reserved and I was confused that they were missing.
2024-06-21 05:34:44 -07:00
Ryan Houdek 3c293b9aed Arm64: Loosen restrictions on V{Load,Store}VectorMasked to allow 128-bit operation 2024-06-21 04:26:09 -07:00
Ryan Houdek 283c2861c9 AVX128: Implement suppor for vlddqu 2024-06-21 00:56:36 -07:00
Ryan Houdek 757dc95116 AVX128: Implement support for the punpckh instructions 2024-06-21 00:56:32 -07:00
Ryan Houdek 6192250b8a AVX128: Implement support for the punpckl instructions 2024-06-21 00:56:28 -07:00
Ryan Houdek f489135b1d Merge pull request #3734 from Sonicadvance1/avx_8
AVX128: Move moves!
2024-06-21 00:53:41 -07:00
Ryan Houdek 4d00a52761 Merge pull request #3732 from Sonicadvance1/avx_6
unittests: Split up vtestps unittest to accumulate flags in independent registers.
2024-06-21 00:52:02 -07:00
Ryan Houdek 6941a59223 unittests: Split up vtestps unittest to accumulate flags in independent registers.
Makes it easier to see what is failing on the 128-bit side versus
256-bit side.
2024-06-21 00:45:30 -07:00
Ryan Houdek 3f232e631e Merge pull request #3730 from Sonicadvance1/avx_4
Vector: Helper refactorings
2024-06-21 00:31:14 -07:00
Ryan Houdek 6e3643c3ef Merge pull request #3714 from pmatos/FSTstiTagSet
Set tag properly in X87 FST(reg)
2024-06-21 00:27:24 -07:00
Ryan Houdek d7348c8aff Merge pull request #3683 from Sonicadvance1/fix_broken_mprotect
SMCTracking: Fix incorrect mprotect tracking
2024-06-20 22:49:51 -07:00
Ryan Houdek e7bdb8679d Merge pull request #3735 from alyssarosenzweig/instcountci/seg-reg-cases
InstCountCI: add segment register cases
2024-06-20 09:43:42 -07:00
Ryan Houdek c28824f94d AVX128: Implements support for vbroadcast* 2024-06-20 09:43:10 -07:00
Ryan Houdek 664d766b45 AVX128: Implement support for vmovshdup 2024-06-20 09:43:10 -07:00
Ryan Houdek fce694ed92 AVX128: Implement support for vmovsldup 2024-06-20 09:43:10 -07:00
Ryan Houdek 96aafb4f07 AVX128: Implement support for vmovddup
This instruction is a little weird.
When accessing memory, the 128-bit operating size of the instruction
only loads 64-bits.
Meanwhile the 256-bit operating size of the instruction fetches a full
256-bits.

Theoretically the hardware could get away with two 64-bit loads or a
wacky 24-byte load, but it looks like to simplify hardware they just
spec'd it that the 256-bit version will always load the full range.
2024-06-20 09:43:10 -07:00
Alyssa Rosenzweig a474f86ea8 InstCountCI: add segment register cases
add a bit of coverage for this funny addressing corner. We do handle this
optimally but I had to write this to check ;)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-20 11:37:35 -04:00
Ryan Houdek dbaf95a8f3 AVX128: Implement support for vmovhps/d 2024-06-20 06:53:21 -07:00
Ryan Houdek e67df96ad9 AVX128: Implement support for movlps/d 2024-06-20 06:53:17 -07:00
Ryan Houdek 56de94578d AVX128: Implement support for vmovq 2024-06-20 06:53:13 -07:00
Ryan Houdek 06fc2f5ef0 AVX128: Implement support for non-temporal moves. 2024-06-20 06:53:09 -07:00
Ryan Houdek b3ba315cbd AVX128: Implements unary/binary lambda helper 2024-06-20 06:53:05 -07:00
Ryan Houdek e5a531e683 Vector: Refactor MPSADBWOpImpl so AVX128 can use it. 2024-06-20 06:43:57 -07:00
Ryan Houdek e2de57bd04 Vector: Refactor PSADBWOpImpl so AVX128 can use it. 2024-06-20 06:43:57 -07:00
Ryan Houdek 4eebca93e3 Vector: Refactor PSHUFBOpImpl. This will be reused for AVX128 2024-06-20 06:33:27 -07:00
Ryan Houdek 3919ec9692 Vector: Expose VBLENDOpImpl in the OpcodeDispatcher. It will be reused by AVX128 2024-06-20 06:33:21 -07:00
Ryan Houdek 02aeb0ac1a Vector: Restructure PMADDWDOpImpl. It's going to get reused for AVX128 2024-06-20 06:33:15 -07:00
Ryan Houdek 206544ad09 Vector: Reconfigure PMADDUBSWOpImpl, it's going to get reused for AVX128 2024-06-20 06:33:08 -07:00
Ryan Houdek 3854cd2b2f Vector: Restruture SHUFOpImpl. AVX128 is going to reuse it. 2024-06-20 06:32:58 -07:00
Alyssa Rosenzweig b2eb8aaf66 Merge pull request #3718 from Sonicadvance1/avx128_3
OpcodeDispatcher: Adds initial groundwork for decomposed AVX operations
2024-06-20 08:57:35 -04:00
Ryan Houdek acbd920c9a OpcodeDispatcher: Adds initial groundwork for decomposed AVX operations
Only installs the tables if SVE256 isn't supported yet AVX is explicitly
enabled with HostFeatures, to protect accidental enablement early.

- Only implements 85 instructions starting out
- Basic vector moves
- Basic vector unary operations
- Basic vector binary operations
- VZeroUpper/VZeroAll

The bulk of the implementation is currently the handling for loading and
storing the halves of the registers from the context or from memory.

This means the load/store helpers must always return a pair unless only
requesting the bottom half of the register, which occurs with 128-bit
AVX operations. The store side then needing to consume the named zero
register if it occurs since those cases will zero the upper bits.

This implementation approach has a few benefits.
- I can pound this out extremely quickly
- SSE implementations are unaffected and don't need to deal with the
  insert behaviour of SVE256.
- We still keep the SVE256 implementation for the inevitable future when
  hardware vendors actually do implement it (Give it 8 years or
  something).
- We can actually unit test this path in CI once it is complete.
- We can partially optimize some paths with SVE128 (Gathers) and support
  a full ASIMD path if necessary.

One downside is that I can't enable this in CI yet because it can't pass
all unittests. but that's a non-issue since it is going to be in heavy
flux as I'm hammering out the implementation. It'll get switched on at
the end when it's passing all 1265 AVX unittests. Currently at 1001 on
this.
2024-06-20 08:44:14 -04:00
Alyssa Rosenzweig db0bdd48e5 Merge pull request #3729 from alyssarosenzweig/refactor/address-modes
OpcodeDispatcher: Refactor address modes
2024-06-20 08:18:33 -04:00
Ryan Houdek da21ee3cda Merge pull request #3692 from pmatos/AFP_RPRES_fix
Fixes AFP.NEP handling on scalar insertions
2024-06-19 19:23:49 -07:00
Ryan Houdek d2baef2b36 Merge pull request #3727 from Sonicadvance1/vaes
VAES support
2024-06-19 19:22:56 -07:00
Ryan Houdek df96bc83cc Merge pull request #3726 from Sonicadvance1/oryon_errata
HostFeatures: Work around Qualcomm Oryon RNG errata
2024-06-19 19:21:14 -07:00
Alyssa Rosenzweig ec03831a21 OpcodeDispatcher: plumb A.NonTSO deeper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 9ca821316a OpcodeDispatcher: factor out DecodeAddress
this is the common guts of the load/store routines.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 025a060337 OpcodeDispatcher: extract IsNonTSOReg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Alyssa Rosenzweig 371d6f0730 OpcodeDispatcher: extract IsOperandMem
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-19 08:52:07 -04:00
Ryan Houdek 643bc10d52 CPUID: Expose VAES if supported 2024-06-19 05:51:47 -07:00
Ryan Houdek 8fb801069f unittests: Adds new VAES tests 2024-06-19 05:51:47 -07:00
Ryan Houdek 542ed8b6ad Implement support for querying AES256 support
This is a different feature flag than regular AES as the default AES+AVX
only operates on 128-bit wide vectors.

With the newer `VAES` extension this is expanded to 256-bit.
2024-06-19 05:51:47 -07:00
Ryan Houdek 053620f4f5 Merge pull request #3728 from pmatos/PythonIgnore
Ignore python files for clang-format
2024-06-19 05:51:25 -07:00
Paulo Matos 88b01a0ca9 Ignore python files for clang-format 2024-06-19 14:23:27 +02:00
Alyssa Rosenzweig 197140498b Merge pull request #3721 from alyssarosenzweig/scripts/instcountci
Scripts: add update_instcountci.sh script
2024-06-19 06:50:00 -04:00
Paulo Matos 9acd325aa4 instcountci: Fixes AFP.NEP handling on scalar insertions 2024-06-19 10:02:54 +02:00
Paulo Matos 2483329ef6 Fixes AFP.NEP handling on scalar insertions
Fixes #3690

When doing scalar insertions, upper bits come from different arguments
depending on the operation. These are listed in the ARM spec under the
NEP bit documentation.
2024-06-19 10:02:54 +02:00
Paulo Matos f6b58b4219 instcountci: Set tag properly in X87 FST(reg) 2024-06-19 10:02:05 +02:00
Paulo Matos 6c6d86f761 unittests: Set tag properly in X87 FST(reg) 2024-06-19 10:02:05 +02:00
Paulo Matos 359221b379 Set tag properly in X87 FST(reg) 2024-06-19 10:02:05 +02:00
Ryan Houdek 87fe1d672e Merge pull request #3715 from pmatos/FXCHFlag
FXCH should set C1 to zero
2024-06-19 00:20:01 -07:00
Paulo Matos 9257221b3b instcountci: FXCH should set C1 to zero 2024-06-19 09:11:49 +02:00
Paulo Matos f9b38a1de7 FXCH should set C1 to zero 2024-06-19 08:57:48 +02:00
Ryan Houdek 67e1ac0442 Merge pull request #3725 from alyssarosenzweig/ir/vbic
IR: rename _VBic -> _VAndn
2024-06-18 16:34:26 -07:00
Ryan Houdek c57e9e008f Merge pull request #3723 from alyssarosenzweig/fexcore/zero-helper
OpcodeDispatcher: refactor zero vector loads
2024-06-18 16:34:15 -07:00
Ryan Houdek b34c23fe3d HostFeatures: Work around Qualcomm Oryon RNG errata
The Oryon is the first CPU we know of that implemented support for the
RNG extension. It also has an errata where reading the RNDRRS register
never returns success. X86's RDSEED guarantees forward progress with
enough retries.

When an x86 processor messed this up at one point, some Linux systems
would infinite loop (presumably when something in boot was filling an
entropy pool). This required a microcode change to fix that processor.

The rdseed unittest infinite loops on this platform if RNG was exposed.
2024-06-18 16:29:53 -07:00
Ryan Houdek 29f644235d Merge pull request #3724 from alyssarosenzweig/ryan-avx-cut
First few commits from Ryan's AVX branch
2024-06-18 11:39:51 -07:00
Alyssa Rosenzweig 01da5972fc IR: rename _VBic -> _VAndn
to be consistent with the scalar _Andn opcode, which is specifically named _Andn
and not _Bic.

noticed while reviewing AVX patches

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 14:00:01 -04:00
Alyssa Rosenzweig 643e964edd Merge pull request #3694 from Sonicadvance1/fix_3691
FEX: Consolidate JSON allocators and fix 3691
2024-06-18 13:54:10 -04:00
Ryan Houdek 30e3d795da FEX: Consolidate JSON allocators and fix 3691
Fixes #3691

We weren't checking if the file was empty before using its `at` function
member. This was causing an early crash if the config file existed but
was empty.

Consolidates the three locations that copy and pasted the json allocator
tools and adds an empty check for all of them.

Also adds two missing checks to the ThunksDB handler that could have
resulted in the same crash if ThunksDB was an empty file.
2024-06-18 13:31:25 -04:00
Alyssa Rosenzweig 89b05a2ea4 Merge pull request #3706 from Sonicadvance1/threadstateobject_cast
LinuxEmulation: Add a helper for getting the ThreadStateObject from CPU frame
2024-06-18 13:28:46 -04:00
Alyssa Rosenzweig 2e009be27c Merge pull request #3708 from Sonicadvance1/fexgetconfig_tsoemulation_facts
FEXGetConfig: Support the ability to get TSO emulation facts
2024-06-18 13:28:02 -04:00
Alyssa Rosenzweig 32150cf7b5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 12:01:23 -04:00
Ryan Houdek bf812aae8f CoreState: Adds avx_high structure for tracking decoupled AVX halves.
Needed something inbetween the `InlineJITBlockHeader` and `avx_high` in
order to match alignment requirements of 16-byte for avx_high. Chose the
`DeferredSignalRefCount` because we hit it quite frequently and it is
basically the only 64-bit variable that we end up touching
significantly.

In the future the CPUState object is going to need to change its view of
the object depending on if the device supports SVE256 or not, but we
don't need to frontload the work right now. It'll become significantly
easier to support that path once the RCLSE pass gets deleted.
2024-06-18 12:00:45 -04:00
Ryan Houdek 9a71443005 CoreState: Adds a gregs offset check
This is required to be less than the maximum range for LDP and STP in
the Arm64 Dispatcher otherwise it breaks. Necessary to ensure this when
reorganizing the CoreState.
2024-06-18 12:00:45 -04:00
Ryan Houdek ee165249bc Dispatcher: Fix ARM64EC
We don't have CI for this and was missed.
2024-06-18 12:00:45 -04:00
Mai 7c7d767195 Merge pull request #3722 from alyssarosenzweig/instcountci/disable-afp
InstCountCI: explicitly disable AFP everywhere
2024-06-18 11:52:21 -04:00
Alyssa Rosenzweig af8cfb79e5 OpcodeDispatcher: refactor zero vector loads
AVX128 is going to slam this, so make it more ergonomic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 11:44:46 -04:00
Alyssa Rosenzweig 27c8bf3021 InstCountCI: explicitly disable AFP everywhere
(except for when we explicitly enable AFP).

Since AFP gets saved/restored, we get `msr fpcr` garbage in random instructions
when AFP is enabled. Explicitly disable everywhere since it's not worth our time
to triage which files might hit that path. Fixes instcountci on AFP-supporting
hosts now that we have AFP enabled.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 11:40:20 -04:00
Alyssa Rosenzweig b0a09b31bb Scripts: add update_instcountci.sh script
This is helpful for devs working on FEXCore, I've been using this locally but it
might make sense to stick it in tree.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-18 09:07:49 -04:00
Ryan Houdek 13ebfb1a49 Merge pull request #3711 from Sonicadvance1/avx128_2
FEXCore: Disentangle the SVE256 feature from AVX
2024-06-17 17:35:15 -07:00
Ryan Houdek f863b30951 Merge pull request #3716 from alyssarosenzweig/ir-dump/unrecoverable
json_ir_generator: don't print unrecoverable temps
2024-06-17 17:25:27 -07:00
Ryan Houdek 1ce27a5e6b FEXCore: Disentangle the SVE256 feature from AVX
In quite a few locations we are mixing the case that SVE256 == AVX or
that AVX means the guest register size is 256-bit.

While this is true today, this is entanglement is going to change very
quickly and cause confusion in follow-up PRs.

Now we have SVE128, SVE256, and SVE2 HostFeatures to disambiguate the
different features which mean different things.

This PR keeps the alias that `SupportsAVX` = `SupportsSVE256 && SupportsSVE2`
but that alias is going to very quickly change its definition.
2024-06-17 17:20:32 -07:00
Ryan Houdek 933d622860 Merge pull request #3710 from Sonicadvance1/avx128_1
CoreState: Move `InlineJITBlockHeader` to the start of the struct
2024-06-17 17:17:56 -07:00
Ryan Houdek 5d67223236 Merge pull request #3707 from Sonicadvance1/clang_version_check
CMake: Add a clang version check
2024-06-17 17:17:17 -07:00
Ryan Houdek 825d2c948c Merge pull request #3717 from alyssarosenzweig/ra/ood-comment
Arm64Emitter: drop out of date comment
2024-06-17 16:02:35 -07:00
Alyssa Rosenzweig 29390b439a json_ir_generator: don't print unrecoverable temps
this makes the print more noisy for no benefit, don't do it.

before:

    %9(GPRFixed16) i32 = Add OpSize:Tmp:Size, %6(GPRFixed0) i64, %17(Invalid)
    %10(GPR0) i64 = Bfi OpSize:Tmp:Size, #0x10, #0x0, %6(GPRFixed0) i64, %9(GPRFixed16) i32
    (%11 i64) StoreRegister %6(GPRFixed0) i64, #0x11, GPR, u8:Tmp:Size
    (%12 i64) StoreRegister %9(GPRFixed16) i32, #0x10, GPR, u8:Tmp:Size
    (%13 i64) StoreRegister %10(GPR0) i64, #0x0, GPR, u8:Tmp:Size

after:

    %9(GPRFixed16) i32 = Add %6(GPRFixed0) i64, %17(Invalid)
    %10(GPR0) i64 = Bfi #0x10, #0x0, %6(GPRFixed0) i64, %9(GPRFixed16) i32
    (%11 i64) StoreRegister %6(GPRFixed0) i64, #0x11, GPR
    (%12 i64) StoreRegister %9(GPRFixed16) i32, #0x10, GPR
    (%13 i64) StoreRegister %10(GPR0) i64, #0x0, GPR

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-17 14:58:56 -04:00
Alyssa Rosenzweig 799c17eb90 Arm64Emitter: drop out of date comment
I fixed this when we landed the new RA

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-17 14:58:08 -04:00
Alyssa Rosenzweig 5fb84866e0 json_ir_generator: rework argument printing
for next commit

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-17 14:40:29 -04:00
Alyssa Rosenzweig 4965344ef5 Merge pull request #3705 from alyssarosenzweig/pre-rclse
Clean ups from my RCLSE branch
2024-06-17 14:22:01 -04:00
Alyssa Rosenzweig 46ca53ad0d Merge pull request #3704 from alyssarosenzweig/ra/spill-better
RA: priorize remat over spilling
2024-06-17 09:01:50 -04:00
Alyssa Rosenzweig 61ff1b3584 Merge pull request #3712 from alyssarosenzweig/jit/silly-assert
JIT: delete silly assert
2024-06-17 08:59:00 -04:00
Alyssa Rosenzweig 7c0c5de4bd JIT: delete silly assert
noticed in the area.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-17 08:51:22 -04:00
Ryan Houdek 8d134b8df8 InstcountCI: Update 2024-06-17 03:03:47 -07:00
Ryan Houdek a9bacc1b6b CoreState: Move InlineJITBlockHeader to the start of the struct
This currently doesn't do much but soon this will be very important to
ensure the data prefetcher of Cortex keeps the cachelines following this
variable in L1.
2024-06-17 02:59:56 -07:00
Ryan Houdek e4ff3dac86 Merge pull request #3701 from Sonicadvance1/fix_arch
FEXCore: Fixes Call with 32-bit displacement and address size override
2024-06-16 13:40:34 -07:00
Alyssa Rosenzweig 9443b18076 RegisterAllocationPass: optimize spill loop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-16 08:15:15 -04:00
Ryan Houdek bb4e81aa19 FEXGetConfig: Support the ability to get TSO emulation facts
Allows a nice way to get information about how TSO emulation is occuring
on the individual's hardware. In particular, the more subtle details
around the implementation rather than just a hard on or off toggle.
2024-06-15 20:13:03 -07:00
Ryan Houdek a9a9f6782a CMake: Add a clang version check
Currently our minimum clang version requirement is 12.0 but soon will
require at least 13 or 14. Add a new version check in cmake to ensure
minimum version requirements.

Makes it easier to determine why a build is failing due to old compiler.
2024-06-15 18:41:19 -07:00
Ryan Houdek 2fa6c3c918 LinuxEmulation: Add a helper for getting the ThreadStateObject from CPU frame
Pulled from the seccomp WIP PR where it pulls this object more frequently.
Since it is an opaque frontend pointer it needs to be cast and we
already have a few locations that use it.

No functional change.
2024-06-15 18:31:37 -07:00
Alyssa Rosenzweig 4bd84eb523 OpcodeDispatcher: extract PF/AF invalidate helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig e2073dcd30 OpcodeDispatcher: extract safe Thunk
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig fd72669c7e OpcodeDispatcher: extract safe Break
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig 81c144697b OpcodeDispatcher: extract safe ExitFunction
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig aecf180dfe OpcodeDispatcher: extract FlushRegisterCache
The "end the clause" signal. for now just flushes flags.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig 10fa4a4f20 OpcodeDispatcher: remove never-gonna-be-done todo
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:47 -04:00
Alyssa Rosenzweig 534732564b OpcodeDispatcher: drop pointless thunks for packss
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:23:46 -04:00
Alyssa Rosenzweig 6a314bc9cd RegisterAllocationPass: prioritize remat over spilling
No instcountci changes yet, since nothing currently spills in instcountci. This
mitigates spilling later seen with #3703, and should help for certain
pathological blocks even without those changes (maybe we should try to get some
of those blocks in instcountci?).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-15 20:21:57 -04:00
Ryan Houdek 1d4356b97e Change logic 2024-06-14 14:50:35 -07:00
Ryan Houdek a11566012d SMCTracking: Fix incorrect mprotect tracking
Fixes #3675

This was the first time I've ever actually dived in to this code and this
function melted my brain a bit while reading it. It was trying to be too
smart in tracking VMA splits, but if it split right at the end of a VMA
range it would then add a new range at the end where one already
existed. This then caused us to have overlapping VMA ranges and it would
completely break our SMC tracking since all tracking in the map must not
overlap.

Instead of being too smart, just break it down in to 4 merge strategies,
three of which are one shot. This is significantly easier to reason
about and each strategy is mostly self-contained. The fourth strategy in
the list is the most complex since it requires multiple steps since it
needs to walk multiple VMAs.

To help test this I added some sanity checking code that proved
invaluable to ensure everything was correct. That's not getting merged
since the overhead is too much to run, but good to have available. Diff
for this is at
  https://gist.github.com/Sonicadvance1/1ee60101ed0742b971a476fadbb51083
2024-06-14 14:36:16 -07:00
Ryan Houdek 8d929027c8 Merge pull request #3699 from Sonicadvance1/fix_3698
FEXConfig: Clear up TSO emulation string
2024-06-14 14:35:45 -07:00
Ryan Houdek 1d1ed012d8 FEXCore: Fixes Call with 32-bit displacement and address size override
FEX had a bug with this instruction where it was incorrectly using both
the address size override and operand size override to truncate the
immediate offset. This isn't how the instruction should behave as it
should actually ignore the address size override.

This now puts it correctly inline with how the jump instruction works
and adds a unit test to ensure it doesn't break again.

This fixes a crash from the Arch rootfs from the glibc dynamic linker
being compiling in a way where a call instruction was getting aligned
using this prefix (Since the compiler knew it does nothing).
2024-06-14 14:00:35 -07:00
Ryan Houdek b092b7a937 Merge pull request #3700 from lioncash/update
Externals: Update vixl submodule
2024-06-14 13:26:52 -07:00
Lioncache d133fa6dc1 ASIMD Tests: Remove erroneous disassembly tests
The vixl disassembler has gotten more strict about certain instruction types, so these tests
aren't really needed.

Alternatively, we could mark them as unallocated, but we can opt to remove them here.
2024-06-14 16:12:21 -04:00
Lioncache afa7de969e Externals: Update vixl submodule
Updates vixl to track the latest upstream changes that fix erroneous
non-zeroing behavior for 256-bit vectors
2024-06-14 15:57:03 -04:00
Ryan Houdek 41b6a89ffd FEXConfig: Clear up TSO emulation string
Fixes #3698
2024-06-14 12:39:21 -07:00
Alyssa Rosenzweig 9aa82ec5bf Merge pull request #3695 from Sonicadvance1/fix_3686
Revert "OpcodeDispatcher: optimize logical flags"
2024-06-14 07:47:51 -04:00
Ryan Houdek b17a2e9f96 Merge pull request #3697 from pmatos/ENABLEHOSTF
Use FEX_HOSTFEATURES instead of FEX_ENABLEAVX
2024-06-14 01:01:55 -07:00
Paulo Matos af4e9ceeed Remove vestigial options from workflows
FEX_ENABLEAVX was removed. Settings for host features should use
FEX_HOSTFEATURES but actually FEX_HOSTFEATURES=enableavx has different
behaviour, so we don't enable it.
2024-06-14 09:50:02 +02:00
Ryan Houdek 9c62c41f5f InstcountCI: Update 2024-06-13 19:29:55 -07:00
Ryan Houdek 184c9d21bb Revert "OpcodeDispatcher: optimize logical flags"
This reverts commit bb8336fcad.
2024-06-13 19:28:16 -07:00
Ryan Houdek 9744d8de99 Merge pull request #3689 from catfella/fix_ppa_detection
Scripts/InstallFEX: update PPA URL
2024-06-13 18:06:08 -07:00
Mikhail Nitenko f3e6ecb2c3 Scripts/InstallFEX: update PPA URL detection
Installing PPA with the script now installs a different
URL. This means that GetPPAStatus always returns false.
For the sake of backwards compatibility match the end
of the line to check.
2024-06-14 00:56:42 +00:00
Ryan Houdek aa0f2c3975 Docs: Update for release FEX-2406 2024-06-12 18:41:54 -07:00
Tony Wasserka 4dc8648d81 Merge pull request #3630 from neobrain/feature_libfwd_gl32
Library Forwarding: Add support for 32-bit OpenGL
2024-06-11 17:33:29 +02:00
Tony Wasserka 7a703e1176 Library Forwarding/GL: Enable stricter pointer parameter checks 2024-06-11 17:14:23 +02:00
Tony Wasserka 02df1a2924 Library Forwarding/GL: Avoid pointer array repacking for 64-bit guests 2024-06-11 17:14:23 +02:00
Tony Wasserka efac7efc97 Library Forwarding/GL: Remap _XDisplay returned by glXGetCurrentDisplay 2024-06-11 17:14:23 +02:00
Tony Wasserka d99b4a80c8 Library Forwarding/GL: Add glX support for 32-bit 2024-06-11 17:14:23 +02:00
Tony Wasserka 2a76744d30 Library Forwarding/GL: Add 32-bit support for most core GL APIs 2024-06-11 17:13:45 +02:00
Tony Wasserka 843b2d1969 Library Forwarding/GL: Assume void* always points to compatible data 2024-06-11 16:58:37 +02:00
Tony Wasserka b275c96889 Library Forwarding: Use the fixed-size guest type for passthrough parameters 2024-06-11 16:58:36 +02:00
Ryan Houdek 033b1ce449 Merge pull request #3688 from catfella/extend_bextr_tests
unittests/bextr: add SrcSize tests
2024-06-09 23:18:57 -07:00
Mikhail Nitenko 99a43283be unittests/bextr: add SrcSize tests
dougallj mentioned that adding these tests might expose
a bug in bextr. Since bextr implementation was changed
apparently it now works correctly, that's good.
2024-06-10 05:45:12 +00:00
Mai 55bfd6394b Merge pull request #3640 from Sonicadvance1/cleanup_execve_envp
LinuxSyscalls: Cleanup envp copying in execve
2024-06-07 21:40:13 -04:00
Ryan Houdek 14bfe6016e Merge pull request #3684 from alyssarosenzweig/constprop/cleanup
Constprop: clean up
2024-06-04 13:00:32 -07:00
Alyssa Rosenzweig 0d4ad70875 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig a8bf3859ea ConstProp: rm pointless constant folding
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig aa7dcffcea ConstProp: drop const pool heuristic
slightly worse for compile time, slightly better output, honestly I'll take the
win because this is easier to reason about.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig be1a5cea8e ConstProp: drop addressgen const pool stuff
I don't get the point, it should be handled by a combination of existing
passes/techniques just fine. no instcountci changes.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 402ea84aa0 RedundantFlagCalculationElimination: cleanup DCE
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 19a7b06b91 ConstProp: swallow up LongDivideElimination
as usual.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 96bd643e5b ConstProp: always inline constants
x86/interpreter leftover, I think.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 6b9293979c ConstProp: swallow up InlineCallOptimization
No reason to have a separate pass for this, merging should be a bit faster since
it eliminates an IR walk.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig 7d5cee4384 InlineCallOptimization: rm x86 leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-04 10:09:51 -04:00
Alyssa Rosenzweig c0bab70161 Merge pull request #3682 from alyssarosenzweig/ir/ref
Find-and-replace OrderedNode* with Ref
2024-06-04 10:09:36 -04:00
Alyssa Rosenzweig 32f5a28433 IR: use Ref instead of OrderedNode
find-and-replace across the tree, excluding IR.h itself.

also excluded IRValidation because its treatment of blocks blows up and will be
reformed in the new IR anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig ce30179ed1 IR: add Ref typedef
To put new IR lipstick on the old IR pig.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-03 12:19:34 -04:00
Alyssa Rosenzweig a515b707f3 Merge pull request #3679 from Sonicadvance1/memory_model_emulation_programmer_documentation
FEXCore/docs: Adds programmer documentation about memory model emulation
2024-06-03 09:24:37 -04:00
Ryan Houdek 9ab0fa01bd Merge pull request #3681 from alyssarosenzweig/opt/shld
optimize shld
2024-06-01 11:53:42 -07:00
Alyssa Rosenzweig c3bffa2929 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 14:44:24 -04:00
Alyssa Rosenzweig 951fee361f OpcodeDispatcher: optimize shld
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 14:44:24 -04:00
Ryan Houdek 8c4860b9a7 Merge pull request #3680 from alyssarosenzweig/opt/sib 2024-06-01 11:17:54 -07:00
Ryan Houdek ee221e6a8c Syscalls: Removes unnecessary lambda that was only called once.
Local refactor.
2024-06-01 11:08:34 -07:00
Alyssa Rosenzweig f5625093bb InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:50 -04:00
Alyssa Rosenzweig abfd974d70 OpcodeDispatcher: select hardware addressing modes
Now that we have a framework to do this in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:50 -04:00
Alyssa Rosenzweig 97966930e9 OpcodeDispatcher/x87f64: fuse addr calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig a52a2e3ae4 OpcodeDispatcher/x87: fuse addr
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig c49b30f105 OpcodeDispatcher/Vector: fuse addr calc
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig b0b4ad2083 OpcodeDispatcher: fuse xlat address
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig ee4bee4fef OpcodeDispatcher: fuse BT address
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig c3a0f5a2f6 OpcodeDispatcher: fuse sgdt
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig 0413a6bf68 OpcodeDispatcher: improve bmi2 shift
allow upper garbage, use simpler clean.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:33 -04:00
Alyssa Rosenzweig 7bd036d1ae OpcodeDispatcher: refactor address modes
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-06-01 09:42:32 -04:00
Alyssa Rosenzweig 112c49a348 ConstProp: fix inlining shifted imm to mem instructions
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 80878ae611 ConstProp: rework mem immediate inlining
deduplicate all the things.

functional change:
hit by sse4_1-pmaxuw.c.gcc-target-test-64.jit.gcc-target-64

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:42:48 -04:00
Alyssa Rosenzweig 85a69be5b6 ConstProp: drop address fusion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 17:38:03 -04:00
Ryan Houdek 8dbfd1635a FEXCore/docs: Adds programmer documentation about memory model emulation
I keep needing to look these up to remember the limitations. Add a doc
file so I can more easily point to the information.
2024-05-31 10:36:48 -07:00
Alyssa Rosenzweig 8b5ca303e3 JIT: add asserts for invalid TSO load/store
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-31 12:12:36 -04:00
Ryan Houdek f90d2aeb6d Merge pull request #3678 from marysaka/fix/test-harness-runner-ags-segfault
Fix segfault when starting TestHarnessRunner with missing arguments
2024-05-31 08:06:54 -07:00
Mary Guillemard cd4b520f72 Fix segfault when starting TestHarnessRunner with missing arguments
Signed-off-by: Mary Guillemard <mary@mary.zone>
2024-05-31 16:54:31 +02:00
Ryan Houdek 20d5a26a72 Merge pull request #3674 from alyssarosenzweig/opt/logical-flags
Optimize logical flags
2024-05-30 12:11:21 -07:00
Alyssa Rosenzweig 9346116485 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-30 14:42:29 -04:00
Alyssa Rosenzweig bb8336fcad OpcodeDispatcher: optimize logical flags
fuse the PF write in.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-30 14:42:22 -04:00
Ryan Houdek ee96d60983 Merge pull request #3673 from alyssarosenzweig/ra/tied
Track tied sources in the IR
2024-05-30 10:55:15 -07:00
Alyssa Rosenzweig 6052b335dc Merge pull request #3666 from Sonicadvance1/fix_initial_darwinia
FileManagement: Fix fstatat/statx with self and NOFOLLOW
2024-05-29 23:11:24 -04:00
Ryan Houdek ab0a6bbe9f Merge pull request #3669 from Sonicadvance1/fix_addshift_operation
ConstProp fixes for Darwinia
2024-05-29 19:43:13 -07:00
Ryan Houdek 9dd6d8ed94 Merge pull request #3639 from Sonicadvance1/cleanupFD
FEXLoader: Cleanup FD extraction from environment variables
2024-05-29 19:18:59 -07:00
Ryan Houdek 3b5d0e3e27 FEXLoader: Cleanup FD extraction from environment variables
In preparation for seccomp execve inheritance where we need to extract
another FD from a different environment variable.

- Small function to extract the FD and also unset the environment
  variable in the same place.
   - Keeping the fetch and unset together instead of spreading to
     another location in the source.
- Extract the FD upfront instead of passing the string_view around,
  since we are unsetting the environment variable at the same place.

Future seccomp inheritance will get the FD just after the FEXFD
   - `int FEXSeccompFD {GetFEXFDFromEnv("FEX_SECCOMPFD")};`
2024-05-29 18:47:28 -07:00
Ryan Houdek 37e13cf073 FileManagement: Fix fstatat with self and NOFOLLOW
When asked to not follow the symlink, FEX needs to return data about the
symlink itself rather than following to the target executable. In that
case we need to return symlink information otherwise games that sanity
check can break.

This is what happened with Darwinia in #3662.

We return the FEXInterpreter symlink information in this case since it
doesn't return any information that is relevent to leaking emulator
state. Once the application asks to follow through to the symlink target
is when we will replace.

Also adds a unit test to ensure we don't break it.
2024-05-29 18:41:24 -07:00
Alyssa Rosenzweig 9b1b9c26cc Merge pull request #3664 from Sonicadvance1/change_timestamp
FEXLogging: Changes representation of timestamp
2024-05-29 16:44:42 -04:00
Ryan Houdek f7f3024b92 unittests/ASM: Adds SIB transpose scale register test
With a bit of pointer math it will choose the incorrect address if the
base and offset registers were transposed.
2024-05-29 11:41:20 -07:00
Alyssa Rosenzweig 11ec71a4ce InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 665491adf8 OpcodeDispatcher: drop weird !flagm special case
now that bfi is coalesced, this is a win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 55391ccbc0 RegisterAllocationPass: try to coalesce tied sources
we'll do better in the future but this is already a win.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Alyssa Rosenzweig 7790d7a0b7 IR: track tied sources
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-29 12:32:07 -04:00
Ryan Houdek f3d8c2cbac unittests/ASM: Adds unittest for bug encountered in Darwinia 2024-05-29 07:29:49 -07:00
Ryan Houdek 1226069b4c InstCountCI: Update for fixes
Only prefetch hit currently since ConstProp is limited to optimizing the
ADD IROp atm.
2024-05-29 04:42:40 -07:00
Ryan Houdek 80687c8d2d ConstProp: Limits which addressing modes can be used for vector loadstores
This was causing us to generate invalid code in Darwinia, resulting in a
crash. With assertions enabled this would be picked up in the emitter.

Only implement AddShift optimizations for now because I don't want to do
the remaining optimizations in a bug fix PR.

Fixes Darwinia.
2024-05-29 04:42:11 -07:00
Ryan Houdek 920fe60492 ConstProp: Fix bug with transposed elements from AddShift op
Accidentally we were swapping which sources were the base and which was
the one getting shifted. This wasn't super common so it usually didn't
matter.

Fixes one crash in Darwinia.
2024-05-29 04:32:51 -07:00
Ryan Houdek 8c6ce2cb3b Passes/ConstProp: Have MemExtendedAddressing return a struct rather than a tuple
Makes it less confusing about which variable is the base versus the
offset.

NFC
2024-05-29 04:32:14 -07:00
Ryan Houdek 61f30d004c IR: Document AddShift behaviour
Just to clarify that Src2 is the shifted operation.
2024-05-29 04:29:09 -07:00
Ryan Houdek 95919a1ddf InstcountCI: Add addressing limit tests for base + offset<<shift
These need to be tested.
2024-05-29 04:28:24 -07:00
Ryan Houdek 35ec54f920 Merge pull request #3667 from alyssarosenzweig/opt/pcmp
Optimize PCMPESTRI flags a bit
2024-05-28 22:06:37 -07:00
Alyssa Rosenzweig 32e8a56093 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:32:14 -04:00
Alyssa Rosenzweig 136f1d0a0b OpcodeDispatcher: drop pcmpestri zext
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:32:14 -04:00
Alyssa Rosenzweig 0c042d1e85 VectorFallbacks: optimize PCMP*STRI flags
Return an NZCV.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-28 09:19:43 -04:00
Alyssa Rosenzweig ad13442be4 Merge pull request #3665 from Sonicadvance1/sse42_instcountci
InstCountCI: Adds SSE4.2 operations
2024-05-28 08:57:45 -04:00
Ryan Houdek d6b9252760 InstCountCI: Adds SSE4.2 operations
Doesn't handle all 127 combinations of the control immediate for all
four instructions. Although supplies the instruction control instruction
that A Hat in Time abuses heavily.

The SSE2 implementation of the function in vcruntime140 is likely faster
than our currently implementation but we should be able to get something
comparable. Not bad considering this is a required extension and this is
the first game we found that abuses the instruction heavily.
2024-05-28 00:55:32 -07:00
Ryan Houdek 22222ebaf5 FEXLogging: Changes representation of timestamp
This was a bit confusing to read and I had always expected to change
this at some point.

Previous:
```
[INFO][1579518391560577][1601857.1601857] clone: Unsupported flags w/o CLONE_THREAD (Shared Resources), 4100
```

Now:
```
[INFO][1590468.992593376][1629501.1629501] clone: Unsupported flags w/o CLONE_THREAD (Shared Resources), 4100
```
2024-05-27 23:36:58 -07:00
Alyssa Rosenzweig 734258e23b Merge pull request #3661 from Sonicadvance1/remove_warnings2
Removes warnings
2024-05-25 11:52:37 -04:00
Ryan Houdek 74916b3757 RAPass: Remove warnings 2024-05-24 18:41:30 -07:00
Ryan Houdek c5359264a3 VixlUtils: Remove warnings 2024-05-24 18:41:19 -07:00
Ryan Houdek 9d0ff7929e Merge pull request #3660 from alyssarosenzweig/opt/smash-less
Delete a big chunk of IR/Passes/*
2024-05-24 16:23:48 -07:00
Alyssa Rosenzweig d3eed27d17 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:45:32 -04:00
Alyssa Rosenzweig bc1669b163 DeadStoreElimination: eliminate map
use a vec. block indices will be dense in the new IR. This is memory intensive
but seems faster in practice.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 83e417b2c6 DeadStoreElimination: combine GPR/FPR handling
slight speed up per profile.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cb00d9171f IR: merge general DCE with flag DCE
Flag DCE needs to do general DCE anyway to converge in one pass. So we can move
the special syscall/atomic logic over to flag DCE and then drop the second DCE
pass altogether. Now local dead code of both is eliminated in a single pass.

Flag DCE is carefully written to converge in a single iteration which makes this
scheme work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig cf77f2ae5d RedundantFlagCalculationElimination: fix convergence issue
If both the destination and the flags are dead for an AddWithFlags, we need to
eliminate it in one pass. If we only replace without elimiating, we would need a
second DCE pass to eliminate. We want DCE to finish in one pass, so fix this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 273d086a7b ConstProp: merge const pooling passes
walk the IR less.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 94d9cf54bc ConstProp: don't push/pop cursor
pointless

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3089e0e6de ConstProp: merge masking opts with const folding
Single pass over the IR now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 3c088fb414 ConstProp: remove masking elimination opts
This has been deadcode since 2020. Drop it so we can focus on what *does* work
and what does matter.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Alyssa Rosenzweig 676c9a9be6 IR: do not return progress from passes
Generally, there are three reasons to track progress:

* Conditional optimizations. E.g. only run DCE if ConstProp succeeds.
* Fixed point optimizations. E.g. keep running the opt loop until convergence.
* Metadata shenianigans.

None of these apply to FEX. We explicitly do not want a nonlinear pass ordering,
instead we want just a few passes that each converge in a single iteration. We
expect them all to make progress when run. As such, tracking progress is a waste
of CPU cycles. Stop doing it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 15:44:49 -04:00
Ryan Houdek 314fea36b4 Merge pull request #3658 from Sonicadvance1/enable_afp 2024-05-24 12:23:29 -07:00
Ryan Houdek 3b7d30d26a Merge pull request #3637 from alyssarosenzweig/ra/mr 2024-05-24 12:14:17 -07:00
Alyssa Rosenzweig a8d32b9a2f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:26:02 -04:00
Alyssa Rosenzweig 24cb02f4ff FEXCore: remove IRCompaction
New RA does not need it for correctness, and the slight slow down to new RA from
not compacting first is much smaller than the cost of compaction. Overall speeds
up node.js start time by ~6% on top of new RA.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 725d0e187a RegisterAllocationPass: rewrite RA
I recommend viewing the new source file as the diff is quite messy.

---

The old RA commits every "how not to write an RA" sin in the book.

Chaitin spill-one loop? Check.

Potential spilling caused by alignment issues since there's no live range
splitting? Check.

Panic spilling? Check.

Generating an interference graph with linear live ranges, so you get the code
quality of linear scan with the cost of graph colouring? Check.

...

It is wholly unsuitable to any application, and specifically unsuitable for FEX.

---

The new RA exploits a key IR invariant unique to FEX: no values are live across
block boundaries. This is validated.

Because of this invariant, all RA is block local. This lets us use a dead simple
2 pass RA that generates ~optimal code in linear time.

The first pass walks the IR backwards, analyzing the IR. This is a souped up
analogue to liveness analysis.

The second pass walks the IR forward, blasting out registers. If necessary, it
will insert spill and/or shuffle code on the fly. Spilling uses the well-known
furthest-first heuristic, which has excellent results for straight line code.

That's it :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig d8603cb9bd RAValidation: weaken fill validation
Fails with the new RA when copies are inserted, since RAValidation loses
visibility. Nontrivial to fix, but this particular assert hopefully isn't buying
us a ton. Weaken it for now, we can revisit later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Alyssa Rosenzweig 04c2cb5feb IR: add GPR copy/swap instructions
To implement GPRPair "properly", we need to be able to shuffle scalars around
the register file. That means we need explicit copy/swap instructions that RA
can generate them. Add some.

Swap is split in a really sketchy way, because of the 1 instruction = 1 dest
requirement. Hopefully that requirement is lifted in the future and then this
goes away.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-24 09:25:44 -04:00
Ryan Houdek 32f2decb24 Merge pull request #3656 from alyssarosenzweig/opt/asr-masking
Optimize asr
2024-05-23 21:35:48 -07:00
Ryan Houdek 6954ebe3a0 Merge pull request #3651 from Sonicadvance1/change_default_tso
Config: Change default TSO options
2024-05-23 21:35:33 -07:00
Ryan Houdek 2e40da3d6b HostFeatures: Enable AFP and RPRES
This has been investigated. Theoretically should work.
2024-05-23 13:26:05 -07:00
Alyssa Rosenzweig 559772bb03 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig 2bbcf72e27 OpcodeDispatcher: optimize asr masking
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Alyssa Rosenzweig aa3a92aa60 OpcodeDispatcher: merge asr impls
so we can DRY the next patch

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-23 08:27:13 -04:00
Mai 5497240a25 Merge pull request #3655 from alyssarosenzweig/opt/wacky-imul
Optimize large sign-extended constants
2024-05-22 23:27:09 -04:00
Alyssa Rosenzweig 79c609d0f5 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:53:19 -04:00
Alyssa Rosenzweig b33788e765 Arm64Emitter: handle another class of constants
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:50:20 -04:00
Alyssa Rosenzweig 8340012466 InstCountCI: add wacky imul
found in bytemark, exposes an interesting constant case

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 22:50:19 -04:00
Ryan Houdek 0adcc779cf Merge pull request #3654 from alyssarosenzweig/opt/movsx
Optimize sign-extension
2024-05-22 19:29:51 -07:00
Alyssa Rosenzweig 9d0718fbc4 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Alyssa Rosenzweig 53567a6526 OpcodeDispatcher: optimize movsxd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Alyssa Rosenzweig 90fb5f038b OpcodeDispatcher: optimize movsx
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 21:00:25 -04:00
Mai 063b1eb936 Merge pull request #3652 from alyssarosenzweig/instcountci/more-bytemark
InstructionCountCI: add bytemark hot block
2024-05-22 14:20:13 -04:00
Alyssa Rosenzweig 9f243c8f7b InstructionCountCI: add bytemark hot block
this was basically practice for using perf with FEX, but a few things do stand
out in the assembly as suboptimal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-22 14:09:50 -04:00
Ryan Houdek 2dc600f283 Config: Change default TSO options
After two months of testing I finally have enough confidence that these
default options getting changed is safe enough that most games won't
notice the difference. But the performance differences can be wild.

Might want to come back and reenable this on platforms that support
hardware TSO and LRCPC3 (for the vector feature), but we can care once
those platforms come up to speed.
2024-05-21 18:08:58 -07:00
Ryan Houdek a01402d502 Merge pull request #3650 from alyssarosenzweig/unittests/xess
unittests: add XeSS test
2024-05-21 17:15:08 -07:00
Ryan Houdek c90036aeea Merge pull request #3649 from alyssarosenzweig/ra/validate-less-hard
Simplify/fix our validation passes
2024-05-21 17:12:49 -07:00
Alyssa Rosenzweig dfb751eea0 unittests: add XeSS test
Useful smoke test for the asymptoptic behaviour of our constant pooling code.
Not useful for correctness testing, so skip in CI.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:41:47 -04:00
Alyssa Rosenzweig 6e0f5eccb3 RAValidation: fix spillregister validation
It doesn't write to its node. Fixes spurious

  %7: Arg[0] expects reg0 to contain %4, but it actually contains %16

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 579fb42458 RAValidation: allow filling the same slot multiple times
This can be a reasonable thing to do!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig f4b487352c RAValidation: defeature control flow analysis
now that we've eliminated cross block liveness, we can do our validation locally
too for a massive simplification.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 4448f84f29 IRValidation: merge in ValueDominanceValidation
All we actually need to validate is that each source has been previously defined
within the block. That checks everything we care about now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:34:31 -04:00
Alyssa Rosenzweig 9e1e602e09 BitSet: fix memset/memclear logic
Missing a factored of 4, causing a buffer overflow.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 19:32:54 -04:00
Ryan Houdek ca70e387ec Merge pull request #3648 from alyssarosenzweig/ra/pair-extract
Slightly improve pair coalescing + memcpy fix from RA branch
2024-05-21 16:19:35 -07:00
Ryan Houdek 9a483107e3 Merge pull request #3647 from alyssarosenzweig/ir/pass-simpler
ConstProp, RCLSE: simplifications
2024-05-21 16:11:36 -07:00
Ryan Houdek 3bac767866 Merge pull request #3645 from neobrain/refactor_aotir
AOTIR: Refactor interfaces to clarify ownership flow
2024-05-21 16:03:06 -07:00
Ryan Houdek 7b4e48480b Merge pull request #3646 from alyssarosenzweig/opt/minor-disp
OpcodeDispatcher: eliminate some Bfe's
2024-05-21 15:57:52 -07:00
Alyssa Rosenzweig 101bba4808 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig a31c3c1c15 OpcodeDispatcher: use ExtractPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 1a18e392f8 OpcodeDispatcher: add ExtractPair helper
terser and will aid coalescing, as well as eventual transition to multidest
extracts which is what we'll actually want.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 4c7595c68a JIT: allow coalescing ExtractElementPair
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 17:14:57 -04:00
Alyssa Rosenzweig 1a467f0ebd IR: reduce memcpy worstcase reg pressure
This avoids a bunch of sharp edges for RA at a small cost when obscure
segment registers are used.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:54:06 -04:00
Alyssa Rosenzweig 06e7360f4c RCLSE: only run once, do not DCE
From a theoretical perspective, we should not need to run RCLSE more than once.
If there are convergence issues with the current implementation, they should be
fixed instead of bandaged around. Fortunately, this has no instcountci changes.

Brings RCLSE cost down from like 12% to 5%.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig ec3b72e17e ConstProp: drop select folding
no instcountci changes.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig 259e1b75a4 ConstProp: rm printfs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig f9642cba7a ConstProp: rm pointless opcode casts
Just use the headers directly

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:46:44 -04:00
Alyssa Rosenzweig c01c415030 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:40:59 -04:00
Alyssa Rosenzweig 50e56358c3 OpcodeDispatcher: eliminate Bfe's with cmpxchg
ConstProp was catching these but they're pointless.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:39:00 -04:00
Alyssa Rosenzweig 465dbc260f OpcodeDispatcher: eliminate Bfe's with lea
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-21 16:39:00 -04:00
Tony Wasserka 9f291f3adb AOTIR: Remove obsolete fields from serialized data 2024-05-21 17:54:28 +02:00
Tony Wasserka f2acc3da4c Core: Clean up ownership management for IRListView and RegisterAllocationData
IRListView is now purely a view type. Instead, ownership is managed on-demand
by a separate interface (IRStorageBase). Materialization of IRListViews to
owning types is moved to this interface as well.

This also avoids unneeded copies of the data.
2024-05-21 17:51:17 +02:00
Tony Wasserka f8f165d96d AOTIR: Remove redundant variable declarations
This code was already using structured bindings anyway and just reassigned
the values to different variables.
2024-05-21 17:38:41 +02:00
Tony Wasserka baef95992c AOTIR: Drop unneeded local variable 2024-05-21 17:38:41 +02:00
Tony Wasserka 97e18c8469 AOTIR: Drop effectively unused parameter from PreGenerateIRFetch 2024-05-21 17:38:41 +02:00
Tony Wasserka 6e04f7368b AOT: Use std::optional to replace a validity boolean in PreGenerateIRFetch 2024-05-21 17:38:41 +02:00
Tony Wasserka 55a835ebb8 AOTIR: Clarify serialization code
The comments weren't too helpful. Using the struct types directly conveys the
same information more clearly.
2024-05-21 17:38:41 +02:00
Alyssa Rosenzweig 85776c2537 Merge pull request #3643 from alyssarosenzweig/opt/shift-garbage
Allow garbage on more shifts
2024-05-21 11:05:47 -04:00
Ryan Houdek e3ec25d9db Merge pull request #3638 from alyssarosenzweig/jit/dedupe-vec
JIT/VectorOps: deduplicate common implementations
2024-05-20 14:44:52 -07:00
Alyssa Rosenzweig ebfcc1e835 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Alyssa Rosenzweig 769b2c2a46 OpcodeDispatcher: allow garbage on shift dests
doesn't matter for left shifts (we mask off the garbage), or 32-bit shifts, or
shifts where we explicitly sbfe after.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Alyssa Rosenzweig 3b2100307e OpcodeDispatcher: allow garbage on more shifts
we're masking anyway

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:33:15 -04:00
Ryan Houdek 28cc179214 LinuxSyscalls: Cleanup envp copying in execve
In preparation for seccomp execve inheritance.

We are going to need to add a new environment variable earlier in the
execve sequence to handle inheritance in the case of binfmt_misc.

No functional change in regards to envp handling.

Minor change around execveat with FD without binfmt_misc. In the case
that execveat returned an error and we did a `dup` of the FD then we
would have an FD leak. Make sure to close the duplicated FD in that
instance.
2024-05-20 07:27:37 -07:00
Alyssa Rosenzweig 3c0f243a2d JIT: dedupe MapCC
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:56 -04:00
Alyssa Rosenzweig bbf1563f80 JIT/ALUOps: extract DEF_BINOP_WITH_CONSTANT
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:56 -04:00
Alyssa Rosenzweig ed6b1011f9 JIT/VectorOps: deduplicate common implementations
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 10:25:55 -04:00
Ryan Houdek e3e7f0279c Merge pull request #3644 from alyssarosenzweig/clang-format/left
clang-format: left-align escaped newlines
2024-05-20 07:12:50 -07:00
Alyssa Rosenzweig a10f984b1c clang-format: left-align escaped newlines
alternative to #3638. this is theoretically better for side-by-side diffs. in
practice it may make other diffs worse since all the \'s change when part of the
macro change.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-20 09:47:21 -04:00
Ryan Houdek 663f3d8b5a Merge pull request #3641 from Sonicadvance1/instcountci_flake
InstCountCI: Hardcode the offset to load tests into
2024-05-20 06:45:47 -07:00
Ryan Houdek b1f7be2f6c InstCountCI: Update 2024-05-18 18:08:38 -07:00
Ryan Houdek b83cbcb33c ConstProp: Bandage fix for instcountci
Fixed offset x86 code doesn't quite solve the issue, so adjust this
heuristic just to get instcounci to stop flaking.

This code is going to heavily change soon anyway so +50 doesn't change
much.
2024-05-18 18:06:25 -07:00
Ryan Houdek ac1a096bae InstCountCI: Hardcode the offset to load tests into
Depending on where the assembly was getting loaded in to memory it was
causing slight code generation differences.

Map the entire file to the same fixed offset as our ASM tests to ensure
consistency and removing flakes in CI.
2024-05-18 17:00:28 -07:00
Ryan Houdek 048c8ded88 Merge pull request #3622 from Sonicadvance1/move_emitter 2024-05-17 10:41:51 -07:00
Alyssa Rosenzweig 948938bf4b Merge pull request #3636 from alyssarosenzweig/jit/factor-vec
JIT: factor out sub reg size conversion
2024-05-17 09:40:32 -04:00
Alyssa Rosenzweig bb064c7334 JIT/AtomicOps: factor out elementsize
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:28:58 -04:00
Alyssa Rosenzweig 2d3d49b900 JIT/ConversionOps: use ConvertSubRegSize*
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:18:59 -04:00
Alyssa Rosenzweig ea7096ed5b JIT/MemoryOps: use ConvertSubRegSize8
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:16:12 -04:00
Alyssa Rosenzweig 7a0f6c0a80 JIT: factor ConvertSize helper
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 17:13:00 -04:00
Alyssa Rosenzweig e4ee35a925 JIT: factor out sub reg size conversion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 15:56:05 -04:00
Ryan Houdek d3ab9bdef6 Remove Float16
We aren't using it. We won't be using it. We need unit tests in our
lives if we want this.
2024-05-16 12:06:54 -07:00
Ryan Houdek 926eefc86c Merge pull request #3635 from alyssarosenzweig/opt/flag-store
OpcodeDispatcher: reorder some moves
2024-05-16 10:58:40 -07:00
Ryan Houdek 3eb7a5b998 Merge pull request #3632 from pmatos/RemoveBlocks
Use erase-remove idiom to remove element
2024-05-16 10:50:50 -07:00
Ryan Houdek 58614ff131 Merge pull request #3634 from alyssarosenzweig/constprop/leftover
ConstProp: remove x86 jit leftover
2024-05-16 10:48:31 -07:00
Alyssa Rosenzweig bf3a09e5e3 ConstProp: remove x86 jit leftover
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 09:21:41 -04:00
Alyssa Rosenzweig 83c536c47f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 08:32:11 -04:00
Alyssa Rosenzweig 7b39e57e72 OpcodeDispatcher: defer overwritten store
this can save moves, as it's a bit easier to reason about the live ranges.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-16 08:31:58 -04:00
Alyssa Rosenzweig 5bedf32666 Merge pull request #3633 from pmatos/UndefShift
Fix left shift undefined behaviour
2024-05-16 07:45:38 -04:00
Paulo Matos 1eb7be9870 Check BitOffset instead of zeroext 2024-05-16 10:50:40 +02:00
Paulo Matos 93e4288c57 Fix left shift undefined behaviour
Here BitOffset can have values higher than 32.
2024-05-16 10:22:03 +02:00
Paulo Matos 9ca4868833 Use erase-remove idiom to remove element
Fixes #3631
2024-05-16 10:16:01 +02:00
Alyssa Rosenzweig c5f8ea58e9 Merge pull request #3629 from pmatos/Unsup-typo
NFC: Fix typo
2024-05-15 10:31:38 -04:00
Paulo Matos 5bee17bee1 NFC: Fix typo 2024-05-15 15:10:00 +02:00
Ryan Houdek efe7c54374 Merge pull request #3625 from Sonicadvance1/restricted_inst
FEXCore: Fixes the difference between CPL-0 and undefined instructions
2024-05-14 07:11:19 -07:00
Ryan Houdek 9e1840e974 FEXCore: Moves CodeEmitter to FHU
Now that the vixl dependency is gone, this gets moved to FHU since the
frontend is going to need it for a microjit.
2024-05-13 12:48:10 -07:00
Ryan Houdek 1f40590f9a Emitter: Inline IsImmLogical from vixl
The only core vixl usage we use in the emitter. Is a complete pain to
reimplement so keep it around.
2024-05-13 12:48:10 -07:00
Ryan Houdek 64a3bc235d unittests/Emitter: Ensures coverage of imm float encodings
To ensure we round everything correctly for the new float16 class
2024-05-13 12:48:10 -07:00
Ryan Houdek 7d9af246ea CodeEmitter: Removes vixl Float16 usage
Creating local Float16 helper which handles our needs
2024-05-13 12:48:09 -07:00
Ryan Houdek 6d3471bcaa Merge pull request #3627 from Sonicadvance1/timeout_merge_base
Github: Support a timeout on checkout
2024-05-13 11:53:11 -07:00
Ryan Houdek a8714dbd49 Github: Support a timeout on checkout
Sometimes github or the CI runner times out trying to checkout the
source and stays timing out forever.

Give it a three minute timeout otherwise the CI runner will stall
forever.
2024-05-13 11:26:54 -07:00
Ryan Houdek 512312fa06 FEXLinuxTests: Implements a test for the new instructions 2024-05-13 11:12:26 -07:00
Ryan Houdek 010028e381 FEXCore: Fixes the difference between CPL-0 and undefined instructions
undefined instructions are expected to return SIGILL, while implemented
instructions that aren't available in CPL-3 are expected to SIGSEGV.

Noticed this while testing out CPU-Z, it installs a kernel module and
does a bunch of `RDMSR` and `OUTS` instructions. Decided to walk through
the rest of the instructions in the `System Instruction Reference`
section.

Turns out there's a bunch of oddities in there that we don't support.
First step is to go through all the explicitl SIGILL and SIGSEGV and
implement a test for them.

Next step will be implementing the remaining operations that are
considered "System" operations but are still available in CPL-3.
This list includes:
- lar
- lgdt
- lsl
- sidt
- sldt
- stac
- clac
- verr
- verw
2024-05-13 11:12:26 -07:00
Ryan Houdek f27f1871e4 Merge pull request #3624 from Sonicadvance1/faulty_mc_fault_face
FEXCore: Get rid of DeferredSignalFaultAddress and use the InterruptFaultPage
2024-05-13 03:22:43 -07:00
Ryan Houdek 3a7aa83ab1 Merge pull request #3626 from pmatos/TestClangIgnore
Fix exec path where file needs to be ignored
2024-05-13 02:55:33 -07:00
Paulo Matos bcc136c3b9 Fix exec path where file needs to be ignored
Ignored files were not being checked. Both clang-format.py wrapper
and code-format-helper where not aligned.
2024-05-13 11:41:44 +02:00
Ryan Houdek 3da31830d1 InstcountCI: Update 2024-05-10 15:34:13 -07:00
Ryan Houdek d19b57a52e FEXCore: Get rid of DeferredSignalFaultAddress and use the InterruptFaultPage
Arm64ec introduced the InterruptFaultPage which is lower overhead since
instead of ldr+str it just turns in to a single str. We were already
allocating the space, FEXCore and the frontend signal delegator just
needed to be updated to understand the new location.

We can additionally use this in the future if we want to make deferred
async signals INSIDE the JIT only cost a single str as well.
2024-05-10 15:31:28 -07:00
Ryan Houdek ef6d640a8c Merge pull request #3612 from Sonicadvance1/threadmanager_move
FEXLoader: Changes frontend thread management to wrap FEXCore thread objects
2024-05-09 09:27:11 -07:00
Ryan Houdek 2cae2f2462 Merge pull request #3617 from bylaws/arm64ec-dispatcher
FEXCore: ARM64EC x64 entry/exit support
2024-05-08 12:25:26 -07:00
Ryan Houdek 1fde5d7fca Merge pull request #3621 from alyssarosenzweig/ra/drop-avx
RegisterAllocationPass: drop AVX flag
2024-05-08 11:42:53 -07:00
Ryan Houdek 10de2f83ac Merge pull request #3620 from alyssarosenzweig/ir/burn-parser
IR: drop IRParser
2024-05-08 11:29:37 -07:00
Alyssa Rosenzweig 9d86e11a47 RegisterAllocationPass: drop AVX flag
RA should not depend on whether we support AVX, that's a huge layering
violation! and fortunately, it does not.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:26:31 -04:00
Alyssa Rosenzweig 4d503d3155 RegisterAllocationPass: drop prewritable check
always true.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:24:41 -04:00
Alyssa Rosenzweig 7e663b91df IR: drop IRParser
Aside from its own self-test, the parser is unused and should remain that way,
since it's a maintenance burden with no real benefit. Burn it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:16:54 -04:00
Alyssa Rosenzweig 47242dc190 Merge pull request #3616 from alyssarosenzweig/sra/simplify-1
SRA controlled burn
2024-05-08 14:10:48 -04:00
Alyssa Rosenzweig e13c8e3295 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3c3ba62c10 MemoryOps: optimize 32-bit SRA case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 34fe56dfb2 DeadStoreElimination: CSE block info
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig ecf8cde5e0 DeadStoreElimination: group common logic
slightly less obnoxious copypaste.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 55284aad7e DeadStoreElimination: don't handle partial stores
SRA replaces the whole contents of the destination.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3afc35f7b4 DeadStoreElimination: simplify
use registers internally, not synthesized offsets

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1058428a51 IR: document invariant on SRA
This lets us simplify a lot!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig a2fc51fc7b IR: specify registers, not offsets for SRA
SRA is fundamentally about hardware registers, not stores into a
software-defined context. So, it should take a register instead of an offset.
This makes all the unaligned special cases unrepresentable (by design).

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 1848629ba5 RegisterAllocationPass: drop aliasable check
always true with the new ir invariants.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 3399577330 JIT: clean up fpr sra
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 18bfc8afd0 JIT: clean up gpr sra handling
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 76b023ed3e JIT: drop unaligned and partial SRA handling
This is all dead, assert as much so it stays that way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig b91b0e9d65 IR: infer SRA static class
no need to stick it in the IR.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Alyssa Rosenzweig 74489a4177 IR: remove dead SRA flags
I don't know what these were meant for, and I don't care (-:

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-05-08 14:01:42 -04:00
Ryan Houdek 55d1d6bcd4 Merge pull request #3615 from bylaws/wow64-fix
Fix WOW64 frontend with recent wine versions
2024-05-07 22:35:32 -07:00
Ryan Houdek cd249e2c3a Merge pull request #3614 from Sonicadvance1/remove_temporary_allocation
FEXServer: Removes temporary variable allocation
2024-05-07 22:35:23 -07:00
Billy Laws 61cd835754 Update InstCountCI 2024-05-06 17:37:43 +00:00
Billy Laws bd24364c1b FEXCore: Switch stacks before exiting the JIT on ARM64EC
This removes the need for the frontend to have any knowledge of FEX's
SRA layout.
2024-05-06 15:41:34 +00:00
Billy Laws ab516d7b79 Dispatcher: Implement ARM64EC SRA setup entrypoints
While the ARM64EC ABI mostly matches FEX's SRA, the stack still needs to
be switched to the emulator stack and target RIP stored into the FEX
context before jumping to the dispatcher loop.
2024-05-06 15:41:34 +00:00
Billy Laws d25ed4b0bf Dispatcher: Block system call callbacks when compiling code
These callbacks are used for code invalidation and setting the right
emulated CPU features, neither of which are necessary for syscalls made
from within FEX. Avoid calling them to prevent deadlocks caused by
nested locks during compilation.
2024-05-06 15:41:28 +00:00
Billy Laws c521d2b48d WOW64: Support unwinding past FEX from within syscall handlers
This is required by recent wine changes to use longjmp for user
callbacks. Switch to saving the context at every simulate call and
setting the unwind SP/PC to that context with a small SEH trampoline
for the syscall handler.
2024-05-06 15:26:36 +00:00
Billy Laws 9ed8165405 WOW64: Dynamically allocate unixcall/syscall entrypoints
Removes the requirement that FEX needs to be loaded as part of the lower
32-bit address space.
2024-05-06 14:55:59 +00:00
Ryan Houdek 5099b2b5dc FEXServer: Removes temporary variable allocation
Was causing unnecessary memory allocation churn when a FEXInterpreter
was asking for the rootfs folder path.
2024-05-05 14:11:26 -07:00
Ryan Houdek d372552593 FEXLoader: Changes frontend thread management to wrap FEXCore thread objects
A bit of refactoring necessary before we can move the remaining Linux
specific code to the frontend.

Most of this taken from #3535 but attempting to be NFC as much as
possible.
2024-05-05 07:43:09 -07:00
Ryan Houdek 729e32ccc2 Linux: Move ThreadManager to its own header 2024-05-05 06:32:59 -07:00
Mai 170204d6f1 Merge pull request #3609 from neobrain/fix_catch2_setting
CMake: Remove obsolete Catch2 setting
2024-05-04 00:04:52 -04:00
Mai f7bfecd3f1 Merge pull request #3610 from Sonicadvance1/support_oryon_named
CPUID: Adds Qualcomm Oryon product name
2024-05-04 00:04:29 -04:00
Ryan Houdek 5f0427c253 CPUID: Adds Qualcomm Oryon product name
From https://github.com/llvm/llvm-project/pull/91022

Easy enough
2024-05-03 20:16:46 -07:00
Tony Wasserka 472860a840 CMake: Remove obsolete Catch2 setting 2024-05-03 15:25:42 +02:00
Mai eddb7d12cc Merge pull request #3608 from Sonicadvance1/readme_remove_x86
Readme: Remove misleading text about x86 hosts being supported
2024-05-03 00:04:53 -04:00
Ryan Houdek 789a9f19c0 Readme: Remove misleading text about x86 hosts being supported
This is no longer the case as x86-64 hosts is purely a development
vehicle and it is not expected for users to try this.
2024-05-02 20:38:13 -07:00
Ryan Houdek f70aafb211 Merge pull request #3607 from teohhanhui/fix/open-mode
Pass compulsory `mode` argument to `open` when `O_CREAT` is used
2024-05-02 13:08:47 -07:00
Teoh Han Hui 7519af2819 Pass compulsory mode argument to open when O_CREAT is used
From `man 2 open`:

> The mode argument must be supplied if O_CREAT or O_TMPFILE is
> specified in flags; if it is not supplied, some arbitrary bytes
> from the stack will be applied as the file mode.
2024-05-03 03:16:29 +08:00
618 changed files with 271379 additions and 53809 deletions

No files matched your search

+1 -1
View File
@@ -7,7 +7,7 @@ AlignConsecutiveAssignments: None
AlignConsecutiveBitFields: Consecutive
AlignConsecutiveDeclarations: None
AlignConsecutiveMacros: None
AlignEscapedNewlines: DontAlign
AlignEscapedNewlines: Left
AlignOperands: Align
AlignTrailingComments: true
AllowAllParametersOfDeclarationOnNextLine: false
-2
View File
@@ -8,7 +8,5 @@ FEXCore/Source/Common/SoftFloat-3e/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
FEXCore/Source/Interface/Core/X86Tables/X87Tables.cpp
FEXCore/Source/Interface/Core/X86Tables/XOPTables.cpp
FEXCore/Source/Interface/Core/X86Tables/*
+1 -2
View File
@@ -13,7 +13,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
build_plus_test:
@@ -64,7 +63,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -2
View File
@@ -20,7 +20,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
glibc_fault_test:
@@ -71,7 +70,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -2
View File
@@ -13,7 +13,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
hostrunner_tests:
@@ -64,7 +63,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
+1 -2
View File
@@ -13,7 +13,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
instcountci_tests:
@@ -74,7 +73,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
+7 -3
View File
@@ -10,14 +10,13 @@ on:
env:
BUILD_TYPE: Debug
FEX_ENABLEAVX: 1
jobs:
mingw_build:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, mingw]]
arch: [[self-hosted, ARM64, mingw], [self-hosted, ARM64EC, mingw, ARM64]]
fail-fast: false
steps:
@@ -39,6 +38,11 @@ jobs:
run: |
echo "MINGW_TRIPLE=aarch64-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64EC
if: matrix.arch[1] == 'ARM64EC'
run: |
echo "MINGW_TRIPLE=arm64ec-w64-mingw32" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
@@ -74,7 +78,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+1
View File
@@ -20,6 +20,7 @@ jobs:
- name: Checkout through merge base
uses: rmacklin/fetch-through-merge-base@v0
timeout-minutes: 3
with:
base_ref: ${{ github.event.pull_request.base.ref }}
head_ref: ${{ github.event.pull_request.head.sha }}
+20 -8
View File
@@ -13,7 +13,6 @@ env:
BUILD_TYPE: Release
CC: clang
CXX: clang++
FEX_ENABLEAVX: 1
jobs:
vixl_simulator:
@@ -65,7 +64,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
- name: Build
working-directory: ${{runner.workspace}}/build
@@ -73,23 +72,22 @@ jobs:
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: ASM Tests
- name: ASM Tests - SVE256
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
- name: ASM Test SVE256 Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE256Bit.log || true
- name: ASM Tests 128-bit
- name: ASM Tests - SVE128
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disableavx"
FEX_FORCESVEWIDTH: "128"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
@@ -98,7 +96,21 @@ jobs:
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM128bit.log || true
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE128Bit.log || true
- name: ASM Tests - ASIMD
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disablesve"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test ASIMD Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_ASIMD.log || true
- name: Truncate test results
if: ${{ always() }}
+72 -10
View File
@@ -1,5 +1,5 @@
cmake_minimum_required(VERSION 3.14)
project(FEX)
project(FEX C CXX ASM)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
@@ -15,6 +15,7 @@ option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_COVERAGE "Enables Coverage" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
@@ -31,6 +32,7 @@ option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FAL
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -40,7 +42,16 @@ string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
message (STATUS "Mingw build")
set (MINGW_BUILD TRUE)
set (ENABLE_JEMALLOC FALSE)
set (ENABLE_JEMALLOC TRUE)
set (ENABLE_JEMALLOC_GLIBC_ALLOC FALSE)
endif()
if (NOT MINGW_BUILD)
message (STATUS "Clang version ${CMAKE_CXX_COMPILER_VERSION}")
set (CLANG_MINIMUM_VERSION 12.0)
if (CMAKE_CXX_COMPILER_VERSION VERSION_LESS ${CLANG_MINIMUM_VERSION})
message (FATAL_ERROR "Clang version too old for FEX. Need at least ${CLANG_MINIMUM_VERSION} but has ${CMAKE_CXX_COMPILER_VERSION}")
endif()
endif()
if (ENABLE_FEXCORE_PROFILER)
@@ -110,6 +121,12 @@ else()
endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
if (NOT ENABLE_X86_HOST_DEBUG)
message(FATAL_ERROR
" FEX-Emu doesn't support compiling for x86-64 hosts!"
" This is /only/ a supported configuration for FEX CI and nothing else!")
endif()
set(_M_X86_64 1)
add_definitions(-D_M_X86_64=1)
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
@@ -123,9 +140,39 @@ endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^arm64ec")
set(_M_ARM_64EC 1)
add_definitions(-D_M_ARM_64EC=1)
endif()
# Required as FEX is not allowed to lock the CRT heap lock during compilation or callbacks
set(ENABLE_JEMALLOC TRUE)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
int Testy(int a, int b, int c, int d, int e, int f) {
return a + b + c + d + e + f;
}
int main() {
return Testy(0, 1, 2, 3, 4, 5);
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
if (MINGW_BUILD)
message(STATUS "Ignoring broken clang::preserve_all support")
set(HAS_CLANG_PRESERVE_ALL FALSE)
else()
message(STATUS "Has clang::preserve_all")
endif()
endif ()
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
add_definitions("-DFEX_PRESERVE_ALL_ATTR=__attribute__((preserve_all))" "-DFEX_HAS_PRESERVE_ALL_ATTR=1")
else()
add_definitions("-DFEX_PRESERVE_ALL_ATTR=" "-DFEX_HAS_PRESERVE_ALL_ATTR=0")
endif()
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
add_definitions(-DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_CCACHE)
@@ -175,13 +222,18 @@ if (ENABLE_TSAN)
link_libraries(-fno-omit-frame-pointer -fsanitize=thread)
endif()
if (ENABLE_COVERAGE)
add_compile_options(-fprofile-instr-generate -fcoverage-mapping)
link_libraries(-fprofile-instr-generate -fcoverage-mapping)
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
# The glibc jemalloc subproject which hooks the glibc allocator.
# Required for thunks to work.
# All host native libraries will use this allocator, while *most* other FEX internal allocations will use the other jemalloc allocator.
add_definitions(-DENABLE_JEMALLOC_GLIBC=1)
add_subdirectory(External/jemalloc_glibc/)
else()
elseif (NOT MINGW_BUILD)
message (STATUS
" jemalloc glibc allocator disabled!\n"
" This is not a recommended configuration!\n"
@@ -194,7 +246,7 @@ if (ENABLE_JEMALLOC)
add_definitions(-DENABLE_JEMALLOC=1)
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
else()
elseif (NOT MINGW_BUILD)
message (STATUS
" jemalloc disabled!\n"
" This is not a recommended configuration!\n"
@@ -202,6 +254,11 @@ else()
" Use at your own risk!")
endif()
if (USE_PDB_DEBUGINFO)
add_compile_options(-g -gcodeview)
add_link_options(-g -Wl,--pdb=)
endif()
set (CMAKE_CXX_FLAGS_RELWITHDEBINFO "${CMAKE_CXX_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
@@ -235,8 +292,6 @@ add_definitions(-Wno-trigraphs)
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
if (BUILD_TESTS)
option(CATCH_BUILD_STATIC_LIBRARY "" ON)
set(CATCH_BUILD_STATIC_LIBRARY ON)
add_subdirectory(External/Catch2/)
# Pull in catch_discover_tests definition
@@ -257,8 +312,6 @@ include_directories(External/json-maker/)
add_subdirectory(External/tiny-json/)
include_directories(External/tiny-json/)
include_directories(External/xbyak/)
include_directories(Source/)
include_directories("${CMAKE_BINARY_DIR}/Source/")
@@ -356,9 +409,18 @@ if (BUILD_TESTS)
include(CTest)
enable_testing()
message(STATUS "Unit tests are enabled")
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
if (TEST_JOB_COUNT)
message(STATUS "Running tests with ${TEST_JOB_COUNT} jobs")
elseif(CMAKE_VERSION VERSION_LESS "3.29")
execute_process(COMMAND "nproc" OUTPUT_STRIP_TRAILING_WHITESPACE OUTPUT_VARIABLE TEST_JOB_COUNT)
endif()
set(TEST_JOB_FLAG "-j${TEST_JOB_COUNT}")
endif()
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
# Binfmt_misc files must be installed prior to Source/ installs
+2
View File
@@ -0,0 +1,2 @@
add_library(CodeEmitter INTERFACE)
target_include_directories(CodeEmitter INTERFACE .)
File diff suppressed because it is too large. Load diff
@@ -8,11 +8,11 @@ public:
public:
// Conditional branch immediate
///< Branch conditional
void b(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void b(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, Imm);
}
void b(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
void b(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
@@ -20,13 +20,13 @@ public:
}
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void b(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
void b(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 0, Cond, 0);
}
void b(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
void b(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
b(Cond, &Label->Backward);
}
@@ -36,11 +36,11 @@ public:
}
///< Branch consistent conditional
void bc(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void bc(ARMEmitter::Condition Cond, uint32_t Imm) {
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, Imm);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
void bc(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
constexpr uint32_t Op = 0b0101'010 << 25;
@@ -49,13 +49,13 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void bc(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
void bc(ARMEmitter::Condition Cond, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0101'010 << 25;
Branch_Conditional(Op, 0, 1, Cond, 0);
}
void bc(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
bc(Cond, &Label->Backward);
}
@@ -65,7 +65,7 @@ public:
}
// Unconditional branch register
void br(FEXCore::ARMEmitter::Register rn) {
void br(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'000 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -74,7 +74,7 @@ public:
UnconditionalBranch(Op, rn);
}
void blr(FEXCore::ARMEmitter::Register rn) {
void blr(ARMEmitter::Register rn) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'001 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -83,7 +83,7 @@ public:
UnconditionalBranch(Op, rn);
}
void ret(FEXCore::ARMEmitter::Register rn = FEXCore::ARMEmitter::Reg::r30) {
void ret(ARMEmitter::Register rn = ARMEmitter::Reg::r30) {
constexpr uint32_t Op = 0b1101011 << 25 |
0b0'010 << 21 | // opc
0b1'1111 << 16 | // op2
@@ -156,13 +156,13 @@ public:
}
// Compare and branch
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0100 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -173,7 +173,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0100 << 24;
@@ -181,7 +181,7 @@ public:
CompareAndBranch(Op, s, rt, 0);
}
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbz(s, rt, &Label->Backward);
}
@@ -190,13 +190,13 @@ public:
}
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0101 << 24;
CompareAndBranch(Op, s, rt, Imm);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -207,7 +207,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
constexpr uint32_t Op = 0b0011'0101 << 24;
@@ -215,7 +215,7 @@ public:
CompareAndBranch(Op, s, rt, 0);
}
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
cbnz(s, rt, &Label->Backward);
}
@@ -225,12 +225,12 @@ public:
}
// Test and branch immediate
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0110 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -241,7 +241,7 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0110 << 24;
@@ -249,7 +249,7 @@ public:
TestAndBranch(Op, rt, Bit, 0);
}
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbz(rt, Bit, &Label->Backward);
}
@@ -258,12 +258,12 @@ public:
}
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, Imm);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -274,14 +274,14 @@ public:
template<typename LabelType>
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
constexpr uint32_t Op = 0b0011'0111 << 24;
TestAndBranch(Op, rt, Bit, 0);
}
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
if (Label->Backward.Location) {
tbnz(rt, Bit, &Label->Backward);
}
@@ -292,7 +292,7 @@ public:
private:
// Conditional branch immediate
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, ARMEmitter::Condition Cond, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= Op1 << 24;
@@ -304,7 +304,7 @@ private:
}
// Unconditional branch register
void UnconditionalBranch(uint32_t Op, FEXCore::ARMEmitter::Register rn) {
void UnconditionalBranch(uint32_t Op, ARMEmitter::Register rn) {
uint32_t Instr = Op;
Instr |= Encode_rn(rn);
dc32(Instr);
@@ -318,8 +318,8 @@ private:
}
// Compare and branch
void CompareAndBranch(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
void CompareAndBranch(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
uint32_t Instr = Op;
@@ -330,7 +330,7 @@ private:
}
// Test and branch - immediate
void TestAndBranch(uint32_t Op, FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
void TestAndBranch(uint32_t Op, ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
uint32_t Instr = Op;
Instr |= (Bit >> 5) << 31;
@@ -4,7 +4,7 @@
#include <cstdint>
#include <cstring>
namespace FEXCore::ARMEmitter {
namespace ARMEmitter {
class Buffer {
public:
Buffer() {
@@ -103,4 +103,4 @@ protected:
uint8_t* CurrentOffset;
uint64_t Size;
};
} // namespace FEXCore::ARMEmitter
} // namespace ARMEmitter
@@ -1,9 +1,6 @@
// SPDX-License-Identifier: MIT
#pragma once
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/EnumUtils.h>
#include <FEXCore/Utils/LogManager.h>
@@ -11,8 +8,8 @@
#include <FEXCore/fextl/vector.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <aarch64/assembler-aarch64.h>
#include <CodeEmitter/Buffer.h>
#include <CodeEmitter/Registers.h>
#include <array>
#include <cstdint>
@@ -56,7 +53,7 @@
* it easier to select the correct load-store instruction. Mostly because these are a nightmare selecting
* the right instruction.
*/
namespace FEXCore::ARMEmitter {
namespace ARMEmitter {
/*
* This `Size` enum is used for most ALU operations.
* These follow the AArch64 encoding style in most cases.
@@ -605,13 +602,25 @@ constexpr bool AreVectorsSequential(T first, const Args&... args) {
return (fn(first, args) && ...);
}
// Returns if the immediate can fit in to add/sub immediate instruction encodings.
constexpr bool IsImmAddSub(uint64_t imm) {
constexpr uint64_t U12Mask = 0xFFF;
auto FitsWithin12Bits = [](uint64_t imm) {
return (imm & ~U12Mask) == 0;
};
// Can fit in to the instruction encoding:
// - if only bits [11:0] are set.
// - if only bits [23:12] are set.
return FitsWithin12Bits(imm) || (FitsWithin12Bits(imm >> 12) && (imm & U12Mask) == 0);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
// Choices:
// - Size of ops passed as an argument rather than template to let the compiler use csel instead of branching.
// - Registers are unsized so they can be passed in a GPR and not need conversion operations
class Emitter : public FEXCore::ARMEmitter::Buffer {
class Emitter : public ARMEmitter::Buffer {
public:
Emitter() = default;
@@ -759,15 +768,17 @@ public:
Bind<false>(&Label->Forward);
}
#include <CodeEmitter/VixlUtils.inl>
public:
// TODO: Implement SME when it matters.
#include "Interface/Core/ArchHelpers/CodeEmitter/ALUOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/BranchOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/LoadstoreOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SystemOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ScalarOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/ASIMDOps.inl"
#include "Interface/Core/ArchHelpers/CodeEmitter/SVEOps.inl"
#include <CodeEmitter/ALUOps.inl>
#include <CodeEmitter/BranchOps.inl>
#include <CodeEmitter/LoadstoreOps.inl>
#include <CodeEmitter/SystemOps.inl>
#include <CodeEmitter/ScalarOps.inl>
#include <CodeEmitter/ASIMDOps.inl>
#include <CodeEmitter/SVEOps.inl>
private:
template<typename T>
@@ -829,4 +840,4 @@ private:
return FEXCore::ToUnderlying(Reg);
}
};
} // namespace FEXCore::ARMEmitter
} // namespace ARMEmitter
@@ -6,7 +6,7 @@
#include <compare>
#include <cstdint>
namespace FEXCore::ARMEmitter {
namespace ARMEmitter {
class WRegister;
class XRegister;
@@ -1024,4 +1024,4 @@ enum class OpType : uint32_t {
Destructive = 0,
Constructive,
};
} // namespace FEXCore::ARMEmitter
} // namespace ARMEmitter
@@ -1506,13 +1506,14 @@ public:
}
// SVE broadcast floating-point immediate (unpredicated)
void fdup(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_AA_FMT(size == FEXCore::ARMEmitter::SubRegSize::i16Bit ||
size == FEXCore::ARMEmitter::SubRegSize::i32Bit ||
size == FEXCore::ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
LOGMAN_THROW_AA_FMT(size == ARMEmitter::SubRegSize::i16Bit ||
size == ARMEmitter::SubRegSize::i32Bit ||
size == ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
uint32_t Imm{};
if (size == SubRegSize::i16Bit) {
Imm = FP16ToImm8(vixl::Float16(Value));
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
} else if (size == SubRegSize::i32Bit) {
Imm = FP32ToImm8(Value);
} else if (size == SubRegSize::i64Bit) {
@@ -1521,7 +1522,7 @@ public:
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
}
void fmov(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
fdup(size, zd, Value);
}
@@ -2568,7 +2569,19 @@ public:
}
// SVE contiguous non-temporal load (scalar plus immediate)
// XXX:
void ldnt1b(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b00, zt, pg, rn, Imm);
}
void ldnt1h(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b01, zt, pg, rn, Imm);
}
void ldnt1w(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b10, zt, pg, rn, Imm);
}
void ldnt1d(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b11, zt, pg, rn, Imm);
}
// SVE contiguous non-temporal load (scalar plus scalar)
// XXX:
// SVE load multiple structures (scalar plus immediate)
@@ -3320,7 +3333,18 @@ public:
// SVE Memory - Contiguous Store with Immediate Offset
// SVE contiguous non-temporal store (scalar plus immediate)
// XXX:
void stnt1b(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalStore(0b00, zt, pg, rn, Imm);
}
void stnt1h(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalStore(0b01, zt, pg, rn, Imm);
}
void stnt1w(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalStore(0b10, zt, pg, rn, Imm);
}
void stnt1d(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalStore(0b11, zt, pg, rn, Imm);
}
// SVE store multiple structures (scalar plus immediate)
void st2b(ZRegister zt1, ZRegister zt2, PRegister pg, Register rn, int32_t Imm = 0) {
@@ -3513,7 +3537,8 @@ private:
size == SubRegSize::i64Bit, "Unsupported fcpy/fmov size");
uint32_t imm{};
if (size == SubRegSize::i16Bit) {
imm = FP16ToImm8(vixl::Float16(value));
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
} else if (size == SubRegSize::i32Bit) {
imm = FP32ToImm8(value);
} else if (size == SubRegSize::i64Bit) {
@@ -3717,7 +3742,7 @@ private:
// SVE bitwise logical operations (predicated)
void SVEBitwiseLogicalPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zdn, ZRegister zm, ZRegister zd) {
LOGMAN_THROW_AA_FMT(size != FEXCore::ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
LOGMAN_THROW_AA_FMT(size != ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
LOGMAN_THROW_A_FMT(zd == zdn, "zd needs to equal zdn");
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
@@ -4479,6 +4504,38 @@ private:
dc32(Instr);
}
// SVE contiguous non-temporal load (scalar plus immediate)
void SVEContiguousNontemporalLoad(uint32_t msz, ZRegister zt, PRegister pg, Register rn, int32_t imm) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
"Invalid loadstore offset ({}). Must be between [-8, 7]", imm);
const auto imm4 = static_cast<uint32_t>(imm) & 0xF;
uint32_t Instr = 0b1010'0100'0000'0000'1110'0000'0000'0000;
Instr |= msz << 23;
Instr |= imm4 << 16;
Instr |= pg.Idx() << 10;
Instr |= Encode_rn(rn);
Instr |= zt.Idx();
dc32(Instr);
}
// SVE contiguous non-temporal store (scalar plus immediate)
void SVEContiguousNontemporalStore(uint32_t msz, ZRegister zt, PRegister pg, Register rn, int32_t imm) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
"Invalid loadstore offset ({}). Must be between [-8, 7]", imm);
const auto imm4 = static_cast<uint32_t>(imm) & 0xF;
uint32_t Instr = 0b1110'0100'0001'0000'1110'0000'0000'0000;
Instr |= msz << 23;
Instr |= imm4 << 16;
Instr |= pg.Idx() << 10;
Instr |= Encode_rn(rn);
Instr |= zt.Idx();
dc32(Instr);
}
void SVEContiguousLoadImm(bool is_store, uint32_t dtype, int32_t imm, PRegister pg, Register rn, ZRegister zt) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
@@ -4743,7 +4800,7 @@ private:
dc32(Instr);
}
void SVEPermuteVector(uint32_t op0, FEXCore::ARMEmitter::ZRegister zd, FEXCore::ARMEmitter::ZRegister zm, uint32_t Imm) {
void SVEPermuteVector(uint32_t op0, ARMEmitter::ZRegister zd, ARMEmitter::ZRegister zm, uint32_t Imm) {
constexpr uint32_t Op = 0b0000'0101'0010'0000'000 << 13;
uint32_t Instr = Op;
@@ -5228,15 +5285,14 @@ private:
// Alias that returns the equivalently sized unsigned type for a floating-point type T.
template <typename T>
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, vixl::Float16>, uint16_t,
std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>>;
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>;
// Determines if a floating-point value is capable of being converted
// into an 8-bit immediate. See pseudocode definition of VFPExpandImm
// in ARM A-profile reference manual for a general overview of how this was derived.
template <typename T>
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
[[nodiscard, maybe_unused]] static bool IsValidFPValueForImm8(T value) {
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
@@ -5277,18 +5333,6 @@ private:
return true;
}
static uint32_t FP16ToImm8(vixl::Float16 value) {
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
"Value cannot be encoded into an 8-bit immediate");
const uint32_t bits = vixl::Float16ToRawbits(value);
const uint32_t sign = (bits & 0x8000) >> 8;
const uint32_t expb2 = (bits & 0x2000) >> 7;
const uint32_t b5_to_0 = (bits >> 6) & 0x3F;
return sign | expb2 | b5_to_0;
}
static uint32_t FP32ToImm8(float value) {
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
"Value ({}) cannot be encoded into an 8-bit immediate", value);
@@ -33,7 +33,7 @@ public:
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
}
void mov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn, uint32_t Index) {
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
dup(size, rd, rn, Index);
}
@@ -1052,21 +1052,21 @@ public:
}
// Floating-point immediate
void fmov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, float Value) {
void fmov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, float Value) {
uint32_t M = 0;
uint32_t S = 0;
uint32_t ptype;
uint32_t imm8;
uint32_t imm5 = 0b0'0000;
if (size == FEXCore::ARMEmitter::ScalarRegSize::i16Bit) {
ptype = 0b11;
imm8 = FP16ToImm8(vixl::Float16(Value));
if (size == ARMEmitter::ScalarRegSize::i16Bit) {
LOGMAN_MSG_A_FMT("Unsupported");
FEX_UNREACHABLE;
}
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i32Bit) {
else if (size == ARMEmitter::ScalarRegSize::i32Bit) {
ptype = 0b00;
imm8 = FP32ToImm8(Value);
}
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i64Bit) {
else if (size == ARMEmitter::ScalarRegSize::i64Bit) {
ptype = 0b01;
imm8 = FP64ToImm8(Value);
}
@@ -1077,7 +1077,7 @@ public:
FloatScalarImmediate(M, S, ptype, imm8, imm5, rd);
}
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, FEXCore::ARMEmitter::VRegister rd) {
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, ARMEmitter::VRegister rd) {
constexpr uint32_t Op = 0b0001'1110'0010'0000'0001'00 << 10;
uint32_t Instr = Op;
@@ -1286,7 +1286,7 @@ public:
private:
// Advanced SIMD scalar copy
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn) {
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
uint32_t Instr = Op;
Instr |= Q << 30;
@@ -11,7 +11,7 @@ public:
// TODO: AT
// TODO: CFP
// TODO: CPP
void dc(FEXCore::ARMEmitter::DataCacheOperation DCOp, FEXCore::ARMEmitter::Register rt) {
void dc(ARMEmitter::DataCacheOperation DCOp, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
}
@@ -48,67 +48,67 @@ public:
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
}
// System instructions with register argument
void wfet(FEXCore::ARMEmitter::Register rt) {
void wfet(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b000, rt);
}
void wfit(FEXCore::ARMEmitter::Register rt) {
void wfit(ARMEmitter::Register rt) {
SystemInstructionWithReg(0b0000, 0b001, rt);
}
// Hints
void nop() {
Hint(FEXCore::ARMEmitter::HintRegister::NOP);
Hint(ARMEmitter::HintRegister::NOP);
}
void yield() {
Hint(FEXCore::ARMEmitter::HintRegister::YIELD);
Hint(ARMEmitter::HintRegister::YIELD);
}
void wfe() {
Hint(FEXCore::ARMEmitter::HintRegister::WFE);
Hint(ARMEmitter::HintRegister::WFE);
}
void wfi() {
Hint(FEXCore::ARMEmitter::HintRegister::WFI);
Hint(ARMEmitter::HintRegister::WFI);
}
void sev() {
Hint(FEXCore::ARMEmitter::HintRegister::SEV);
Hint(ARMEmitter::HintRegister::SEV);
}
void sevl() {
Hint(FEXCore::ARMEmitter::HintRegister::SEVL);
Hint(ARMEmitter::HintRegister::SEVL);
}
void dgh() {
Hint(FEXCore::ARMEmitter::HintRegister::DGH);
Hint(ARMEmitter::HintRegister::DGH);
}
void csdb() {
Hint(FEXCore::ARMEmitter::HintRegister::CSDB);
Hint(ARMEmitter::HintRegister::CSDB);
}
// Barriers
void clrex(uint32_t imm = 15) {
LOGMAN_THROW_AA_FMT(imm < 16, "Immediate out of range");
Barrier(FEXCore::ARMEmitter::BarrierRegister::CLREX, imm);
Barrier(ARMEmitter::BarrierRegister::CLREX, imm);
}
void dsb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
void dsb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
}
void dmb(FEXCore::ARMEmitter::BarrierScope Scope) {
Barrier(FEXCore::ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
void dmb(ARMEmitter::BarrierScope Scope) {
Barrier(ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
}
void isb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(FEXCore::ARMEmitter::BarrierScope::SY));
Barrier(ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(ARMEmitter::BarrierScope::SY));
}
void sb() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::SB, 0);
Barrier(ARMEmitter::BarrierRegister::SB, 0);
}
void tcommit() {
Barrier(FEXCore::ARMEmitter::BarrierRegister::TCOMMIT, 0);
Barrier(ARMEmitter::BarrierRegister::TCOMMIT, 0);
}
// System register move
void msr(FEXCore::ARMEmitter::SystemRegister reg, FEXCore::ARMEmitter::Register rt) {
void msr(ARMEmitter::SystemRegister reg, ARMEmitter::Register rt) {
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
SystemRegisterMove(Op, rt, reg);
}
void mrs(FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::SystemRegister reg) {
void mrs(ARMEmitter::Register rd, ARMEmitter::SystemRegister reg) {
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
SystemRegisterMove(Op, rd, reg);
}
@@ -130,7 +130,7 @@ private:
}
// System instructions with register argument
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, FEXCore::ARMEmitter::Register rt) {
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, ARMEmitter::Register rt) {
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
Instr |= CRm << 8;
@@ -140,13 +140,13 @@ private:
}
// Hints
void Hint(FEXCore::ARMEmitter::HintRegister Reg) {
void Hint(ARMEmitter::HintRegister Reg) {
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
Instr |= FEXCore::ToUnderlying(Reg);
dc32(Instr);
}
// Barriers
void Barrier(FEXCore::ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
void Barrier(ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
Instr |= CRm << 8;
Instr |= FEXCore::ToUnderlying(Reg);
@@ -154,7 +154,7 @@ private:
}
// System Instruction
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, FEXCore::ARMEmitter::Register rt) {
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, ARMEmitter::Register rt) {
uint32_t Instr = Op;
Instr |= L << 21;
@@ -165,7 +165,7 @@ private:
}
// System register move
void SystemRegisterMove(uint32_t Op, FEXCore::ARMEmitter::Register rt, FEXCore::ARMEmitter::SystemRegister reg) {
void SystemRegisterMove(uint32_t Op, ARMEmitter::Register rt, ARMEmitter::SystemRegister reg) {
uint32_t Instr = Op;
Instr |= FEXCore::ToUnderlying(reg);
+351
View File
@@ -0,0 +1,351 @@
// Collection of utilities from vixl.
// Following is the vixl license.
// Copyright 2015, VIXL authors
// All rights reserved.
//
// Redistribution and use in source and binary forms, with or without
// modification, are permitted provided that the following conditions are met:
//
// * Redistributions of source code must retain the above copyright notice,
// this list of conditions and the following disclaimer.
// * Redistributions in binary form must reproduce the above copyright notice,
// this list of conditions and the following disclaimer in the documentation
// and/or other materials provided with the distribution.
// * Neither the name of ARM Limited nor the names of its contributors may be
// used to endorse or promote products derived from this software without
// specific prior written permission.
//
// THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS CONTRIBUTORS "AS IS" AND
// ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
// WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
// DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
// FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
// DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
// SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
// CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
// OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
// OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
// Test if a given value can be encoded in the immediate field of a logical
// instruction.
// If it can be encoded, the function returns true, and values pointed to by n,
// imm_s and imm_r are updated with immediates encoded in the format required
// by the corresponding fields in the logical instruction.
// If it can not be encoded, the function returns false, and the values pointed
// to by n, imm_s and imm_r are undefined.
static bool IsImmLogical(uint64_t value,
unsigned width,
unsigned* n = nullptr,
unsigned* imm_s = nullptr,
unsigned* imm_r = nullptr) {
[[maybe_unused]] constexpr auto kBRegSize = 8;
[[maybe_unused]] constexpr auto kHRegSize = 16;
[[maybe_unused]] constexpr auto kSRegSize = 32;
[[maybe_unused]] constexpr auto kDRegSize = 64;
constexpr auto kWRegSize = 32;
constexpr auto kXRegSize = 64;
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) ||
(width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
bool negate = false;
// Logical immediates are encoded using parameters n, imm_s and imm_r using
// the following table:
//
// N imms immr size S R
// 1 ssssss rrrrrr 64 UInt(ssssss) UInt(rrrrrr)
// 0 0sssss xrrrrr 32 UInt(sssss) UInt(rrrrr)
// 0 10ssss xxrrrr 16 UInt(ssss) UInt(rrrr)
// 0 110sss xxxrrr 8 UInt(sss) UInt(rrr)
// 0 1110ss xxxxrr 4 UInt(ss) UInt(rr)
// 0 11110s xxxxxr 2 UInt(s) UInt(r)
// (s bits must not be all set)
//
// A pattern is constructed of size bits, where the least significant S+1 bits
// are set. The pattern is rotated right by R, and repeated across a 32 or
// 64-bit value, depending on destination register width.
//
// Put another way: the basic format of a logical immediate is a single
// contiguous stretch of 1 bits, repeated across the whole word at intervals
// given by a power of 2. To identify them quickly, we first locate the
// lowest stretch of 1 bits, then the next 1 bit above that; that combination
// is different for every logical immediate, so it gives us all the
// information we need to identify the only logical immediate that our input
// could be, and then we simply check if that's the value we actually have.
//
// (The rotation parameter does give the possibility of the stretch of 1 bits
// going 'round the end' of the word. To deal with that, we observe that in
// any situation where that happens the bitwise NOT of the value is also a
// valid logical immediate. So we simply invert the input whenever its low bit
// is set, and then we know that the rotated case can't arise.)
if (value & 1) {
// If the low bit is 1, negate the value, and set a flag to remember that we
// did (so that we can adjust the return values appropriately).
negate = true;
value = ~value;
}
if (width <= kWRegSize) {
// To handle 8/16/32-bit logical immediates, the very easiest thing is to repeat
// the input value to fill a 64-bit word. The correct encoding of that as a
// logical immediate will also be the correct encoding of the value.
// Avoid making the assumption that the most-significant 56/48/32 bits are zero by
// shifting the value left and duplicating it.
for (unsigned bits = width; bits <= kWRegSize; bits *= 2) {
value <<= bits;
uint64_t mask = (UINT64_C(1) << bits) - 1;
value |= ((value >> bits) & mask);
}
}
// The basic analysis idea: imagine our input word looks like this.
//
// 0011111000111110001111100011111000111110001111100011111000111110
// c b a
// |<--d-->|
//
// We find the lowest set bit (as an actual power-of-2 value, not its index)
// and call it a. Then we add a to our original number, which wipes out the
// bottommost stretch of set bits and replaces it with a 1 carried into the
// next zero bit. Then we look for the new lowest set bit, which is in
// position b, and subtract it, so now our number is just like the original
// but with the lowest stretch of set bits completely gone. Now we find the
// lowest set bit again, which is position c in the diagram above. Then we'll
// measure the distance d between bit positions a and c (using CLZ), and that
// tells us that the only valid logical immediate that could possibly be equal
// to this number is the one in which a stretch of bits running from a to just
// below b is replicated every d bits.
uint64_t a = LowestSetBit(value);
uint64_t value_plus_a = value + a;
uint64_t b = LowestSetBit(value_plus_a);
uint64_t value_plus_a_minus_b = value_plus_a - b;
uint64_t c = LowestSetBit(value_plus_a_minus_b);
int d, clz_a, out_n;
uint64_t mask;
if (c != 0) {
// The general case, in which there is more than one stretch of set bits.
// Compute the repeat distance d, and set up a bitmask covering the basic
// unit of repetition (i.e. a word with the bottom d bits set). Also, in all
// of these cases the N bit of the output will be zero.
clz_a = CountLeadingZeros(a, kXRegSize);
int clz_c = CountLeadingZeros(c, kXRegSize);
d = clz_a - clz_c;
mask = ((UINT64_C(1) << d) - 1);
out_n = 0;
} else {
// Handle degenerate cases.
//
// If any of those 'find lowest set bit' operations didn't find a set bit at
// all, then the word will have been zero thereafter, so in particular the
// last lowest_set_bit operation will have returned zero. So we can test for
// all the special case conditions in one go by seeing if c is zero.
if (a == 0) {
// The input was zero (or all 1 bits, which will come to here too after we
// inverted it at the start of the function), for which we just return
// false.
return false;
} else {
// Otherwise, if c was zero but a was not, then there's just one stretch
// of set bits in our word, meaning that we have the trivial case of
// d == 64 and only one 'repetition'. Set up all the same variables as in
// the general case above, and set the N bit in the output.
clz_a = CountLeadingZeros(a, kXRegSize);
d = 64;
mask = ~UINT64_C(0);
out_n = 1;
}
}
// If the repeat period d is not a power of two, it can't be encoded.
if (!IsPowerOf2(d)) {
return false;
}
if (((b - a) & ~mask) != 0) {
// If the bit stretch (b - a) does not fit within the mask derived from the
// repeat period, then fail.
return false;
}
// The only possible option is b - a repeated every d bits. Now we're going to
// actually construct the valid logical immediate derived from that
// specification, and see if it equals our original input.
//
// To repeat a value every d bits, we multiply it by a number of the form
// (1 + 2^d + 2^(2d) + ...), i.e. 0x0001000100010001 or similar. These can
// be derived using a table lookup on CLZ(d).
static const uint64_t multipliers[] = {
0x0000000000000001UL,
0x0000000100000001UL,
0x0001000100010001UL,
0x0101010101010101UL,
0x1111111111111111UL,
0x5555555555555555UL,
};
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
uint64_t candidate = (b - a) * multiplier;
if (value != candidate) {
// The candidate pattern doesn't match our input value, so fail.
return false;
}
// We have a match! This is a valid logical immediate, so now we have to
// construct the bits and pieces of the instruction encoding that generates
// it.
// Count the set bits in our basic stretch. The special case of clz(0) == -1
// makes the answer come out right for stretches that reach the very top of
// the word (e.g. numbers like 0xffffc00000000000).
int clz_b = (b == 0) ? -1 : CountLeadingZeros(b, kXRegSize);
int s = clz_a - clz_b;
// Decide how many bits to rotate right by, to put the low bit of that basic
// stretch in position a.
int r;
if (negate) {
// If we inverted the input right at the start of this function, here's
// where we compensate: the number of set bits becomes the number of clear
// bits, and the rotation count is based on position b rather than position
// a (since b is the location of the 'lowest' 1 bit after inversion).
s = d - s;
r = (clz_b + 1) & (d - 1);
} else {
r = (clz_a + 1) & (d - 1);
}
// Now we're done, except for having to encode the S output in such a way that
// it gives both the number of set bits and the length of the repeated
// segment. The s field is encoded like this:
//
// imms size S
// ssssss 64 UInt(ssssss)
// 0sssss 32 UInt(sssss)
// 10ssss 16 UInt(ssss)
// 110sss 8 UInt(sss)
// 1110ss 4 UInt(ss)
// 11110s 2 UInt(s)
//
// So we 'or' (2 * -d) with our computed s to form imms.
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
*n = out_n;
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
*imm_r = r;
}
return true;
}
static inline bool IsIntN(unsigned n, int64_t x) {
if (n == 64) return true;
int64_t limit = INT64_C(1) << (n - 1);
return (-limit <= x) && (x < limit);
}
static inline bool IsUintN(unsigned n, int64_t x) {
// Convert to an unsigned integer to avoid implementation-defined behavior.
return !(static_cast<uint64_t>(x) >> n);
}
// clang-format off
#define INT_1_TO_32_LIST(V) \
V(1) V(2) V(3) V(4) V(5) V(6) V(7) V(8) \
V(9) V(10) V(11) V(12) V(13) V(14) V(15) V(16) \
V(17) V(18) V(19) V(20) V(21) V(22) V(23) V(24) \
V(25) V(26) V(27) V(28) V(29) V(30) V(31) V(32)
#define INT_33_TO_63_LIST(V) \
V(33) V(34) V(35) V(36) V(37) V(38) V(39) V(40) \
V(41) V(42) V(43) V(44) V(45) V(46) V(47) V(48) \
V(49) V(50) V(51) V(52) V(53) V(54) V(55) V(56) \
V(57) V(58) V(59) V(60) V(61) V(62) V(63)
#define INT_1_TO_63_LIST(V) INT_1_TO_32_LIST(V) INT_33_TO_63_LIST(V)
// clang-format on
#define DECLARE_IS_INT_N(N) \
static inline bool IsInt##N(int64_t x) { return IsIntN(N, x); }
#define DECLARE_IS_UINT_N(N) \
static inline bool IsUint##N(int64_t x) { return IsUintN(N, x); }
INT_1_TO_63_LIST(DECLARE_IS_INT_N)
INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
#undef DECLARE_IS_INT_N
#undef DECLARE_IS_UINT_N
private:
template <typename V>
static inline bool IsPowerOf2(V value) {
return (value != 0) && ((value & (value - 1)) == 0);
}
// Some compilers dislike negating unsigned integers,
// so we provide an equivalent.
template <typename T>
static inline T UnsignedNegate(T value) {
static_assert(std::is_unsigned<T>::value);
return ~value + 1;
}
static inline uint64_t LowestSetBit(uint64_t value) {
return value & UnsignedNegate(value);
}
template <typename V>
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
#if COMPILER_HAS_BUILTIN_CLZ
if (width == 32) {
return (value == 0) ? 32 : __builtin_clz(static_cast<unsigned>(value));
} else if (width == 64) {
return (value == 0) ? 64 : __builtin_clzll(value);
}
#endif
return CountLeadingZerosFallBack(value, width);
}
static inline int CountLeadingZerosFallBack(uint64_t value, int width) {
LOGMAN_THROW_A_FMT(IsPowerOf2(width) && (width <= 64), "Invalid width");
if (value == 0) {
return width;
}
int count = 0;
value = value << (64 - width);
if ((value & UINT64_C(0xffffffff00000000)) == 0) {
count += 32;
value = value << 32;
}
if ((value & UINT64_C(0xffff000000000000)) == 0) {
count += 16;
value = value << 16;
}
if ((value & UINT64_C(0xff00000000000000)) == 0) {
count += 8;
value = value << 8;
}
if ((value & UINT64_C(0xf000000000000000)) == 0) {
count += 4;
value = value << 4;
}
if ((value & UINT64_C(0xc000000000000000)) == 0) {
count += 2;
value = value << 2;
}
if ((value & UINT64_C(0x8000000000000000)) == 0) {
count += 1;
}
count += (value == 0);
return count;
}
public:
+15 -73
View File
@@ -20,8 +20,7 @@ the coding style of LLVM. It can also be installed as a pre-commit git hook to
check the coding style before submitting it. The canonical source of this script
is in the LLVM source tree under llvm/utils/git.
For C/C++ code it uses clang-format and for Python code it uses darker (which
in turn invokes black).
For C/C++ code it uses clang-format.
You can learn more about the LLVM coding style on llvm.org:
https://llvm.org/docs/CodingStandards.html
@@ -31,8 +30,8 @@ directory:
ln -s $(pwd)/llvm/utils/git/code-format-helper.py .git/hooks/pre-commit
You can control the exact path to clang-format or darker with the following
environment variables: $CLANG_FORMAT_PATH and $DARKER_FORMAT_PATH.
You can control the exact path to clang-format with the following
environment variable: $CLANG_FORMAT_PATH.
"""
@@ -173,6 +172,12 @@ class ClangFormatHelper(FormatHelper):
name = "clang-format"
friendly_name = "C/C++ code formatter"
@property
def cformat_wrapper_path(self) -> str:
relpath = "../../Scripts/clang-format.py"
curpath = os.path.dirname(os.path.abspath(__file__))
return os.path.abspath(os.path.normpath(os.path.join(curpath, relpath)))
@property
def instructions(self) -> str:
return " ".join(self.cf_cmd)
@@ -210,7 +215,11 @@ class ClangFormatHelper(FormatHelper):
if not cpp_files:
return None
cf_cmd = [self.clang_fmt_path, "--diff"]
cf_cmd = [
self.clang_fmt_path,
f"--binary={self.cformat_wrapper_path}",
"--diff",
]
if args.start_rev and args.end_rev:
cf_cmd.append(args.start_rev)
@@ -235,74 +244,7 @@ class ClangFormatHelper(FormatHelper):
else:
return None
class DarkerFormatHelper(FormatHelper):
name = "darker"
friendly_name = "Python code formatter"
@property
def instructions(self) -> str:
return " ".join(self.darker_cmd)
def filter_changed_files(self, changed_files: List[str]) -> List[str]:
filtered_files = []
for path in changed_files:
name, ext = os.path.splitext(path)
if ext == ".py":
filtered_files.append(path)
return filtered_files
@property
def darker_fmt_path(self) -> str:
if "DARKER_FORMAT_PATH" in os.environ:
return os.environ["DARKER_FORMAT_PATH"]
return "darker"
def has_tool(self) -> bool:
cmd = [self.darker_fmt_path, "--version"]
proc = None
try:
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
except:
return False
return proc.returncode == 0
def format_run(self, changed_files: List[str], args: FormatArgs) -> Optional[str]:
py_files = self.filter_changed_files(changed_files)
if not py_files:
return None
darker_cmd = [
self.darker_fmt_path,
"--check",
"--diff",
]
if args.start_rev and args.end_rev:
darker_cmd += ["-r", f"{args.start_rev}...{args.end_rev}"]
darker_cmd += py_files
if args.verbose:
print(f"Running: {' '.join(darker_cmd)}")
self.darker_cmd = darker_cmd
proc = subprocess.run(
darker_cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE
)
if args.verbose:
sys.stdout.write(proc.stderr.decode("utf-8"))
if proc.returncode != 0:
# formatting needed, or the command otherwise failed
if args.verbose:
print(f"error: {self.name} exited with code {proc.returncode}")
# Print the diff in the log so that it is viewable there
print(proc.stdout.decode("utf-8"))
return proc.stdout.decode("utf-8")
else:
sys.stdout.write(proc.stdout.decode("utf-8"))
return None
ALL_FORMATTERS = (DarkerFormatHelper(), ClangFormatHelper())
ALL_FORMATTERS = [ClangFormatHelper()]
def hook_main():
# fill out args
+1 -1
-21
View File
@@ -24,27 +24,6 @@ include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
int Testy(int a, int b, int c, int d, int e, int f) {
return a + b + c + d + e + f;
}
int main() {
return Testy(0, 1, 2, 3, 4, 5);
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
if (MINGW_BUILD)
message(STATUS "Ignoring broken clang::preserve_all support")
set(HAS_CLANG_PRESERVE_ALL FALSE)
else()
message(STATUS "Has clang::preserve_all")
endif()
endif ()
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
+4 -6
View File
@@ -148,9 +148,8 @@ def print_man_options(options):
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
output_man.write(", ".join(f"{enum_op_val}" for [_, enum_op_val] in Enums.items()))
output_man.write("\n.sp\n")
output_man.write(".El\n")
@@ -179,9 +178,8 @@ def print_man_environment(options):
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
output_man.write(", ".join(f"{enum_op_val}" for [_, enum_op_val] in Enums.items()))
output_man.write("\n.sp\n")
print_man_environment_tail()
output_man.write(".El\n")
+80 -94
View File
@@ -2,6 +2,7 @@
import json
import sys
from dataclasses import dataclass, field
import textwrap
def ExitError(msg):
print(msg)
@@ -53,8 +54,10 @@ class OpDefinition:
SSAArgNum: int
NonSSAArgNum: int
DynamicDispatch: bool
LoweredX87: bool
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
Arguments: list
EmitValidation: list
Desc: list
@@ -75,8 +78,10 @@ class OpDefinition:
self.SSAArgNum = 0
self.NonSSAArgNum = 0
self.DynamicDispatch = False
self.LoweredX87 = False
self.JITDispatch = True
self.JITDispatchOverride = None
self.TiedSource = -1
self.Arguments = []
self.EmitValidation = []
self.Desc = []
@@ -202,7 +207,7 @@ def parse_ops(ops):
(OpArg.Type == "GPR" or
OpArg.Type == "GPRPair" or
OpArg.Type == "FPR")):
OpDef.EmitValidation.append("GetOpRegClass({}) == InvalidClass || WalkFindRegClass({}) == {}Class".format(NameWithPrefix, NameWithPrefix, OpArg.Type))
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == InvalidClass || WalkFindRegClass({ArgName}) == {OpArg.Type}Class")
OpArg.Name = ArgName
OpArg.NameWithPrefix = NameWithPrefix
@@ -248,17 +253,23 @@ def parse_ops(ops):
if "JITDispatchOverride" in op_val:
OpDef.JITDispatchOverride = op_val["JITDispatchOverride"]
if "X87" in op_val:
OpDef.LoweredX87 = op_val["X87"]
# X87 implies !JITDispatch
assert("JITDispatch" not in op_val)
OpDef.JITDispatch = False
if "TiedSource" in op_val:
OpDef.TiedSource = op_val["TiedSource"]
# Do some fixups of the data here
if len(OpDef.EmitValidation) != 0:
for i in range(len(OpDef.EmitValidation)):
# Patch up all the argument names
for Arg in OpDef.Arguments:
if Arg.Temporary:
# Temporary ops just replace all instances no prefix variant
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, Arg.Name)
else:
# All other ops replace $ with _ variant for argument passed in
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, "_{}".format(Arg.Name))
# Temporary ops just replace all instances no prefix variant
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, Arg.Name)
#OpDef.print()
@@ -363,25 +374,28 @@ def print_ir_sizes():
if op.Name == "Last":
output_file.write("\t-1ULL,\n")
else:
output_file.write("\tsizeof(IROp_{}),\n".format(op.Name))
output_file.write(f"\tsizeof(IROp_{op.Name}),\n")
output_file.write("};\n\n")
output_file.write(textwrap.dedent("""
};
output_file.write("// Make sure our array maps directly to the IROps enum\n")
output_file.write("static_assert(IRSizes[IROps::OP_LAST] == -1ULL);\n\n")
// Make sure our array maps directly to the IROps enum
static_assert(IRSizes[IROps::OP_LAST] == -1ULL);
output_file.write("[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] std::string_view const& GetName(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetRAArgs(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool ImplicitFlagClobber(IROps Op);\n")
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool GetHasDest(IROps Op);\n")
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
#undef IROP_SIZES
#endif
"""))
def print_ir_reg_classes():
output_file.write("#ifdef IROP_REG_CLASSES_IMPL\n")
@@ -471,16 +485,27 @@ def print_ir_getraargs():
def print_ir_hassideeffects():
output_file.write("#ifdef IROP_HASSIDEEFFECTS_IMPL\n")
for array, prop in [("SideEffects", "HasSideEffects"),
("ImplicitFlagClobbers", "ImplicitFlagClobber")]:
output_file.write(f"constexpr std::array<uint8_t, OP_LAST + 1> {array} = {{\n")
for prop, T in [
("HasSideEffects", "bool"),
("ImplicitFlagClobber", "bool"),
("LoweredX87", "bool"),
("TiedSource", "int8_t"),
]:
output_file.write(
f"constexpr std::array<{'uint8_t' if T == 'bool' else T}, OP_LAST + 1> {prop}_ = {{\n"
)
for op in IROps:
output_file.write("\t{},\n".format(("true" if getattr(op, prop) else "false")))
if T == "bool":
output_file.write(
"\t{},\n".format(("true" if getattr(op, prop) else "false"))
)
else:
output_file.write(f"\t{getattr(op, prop)},\n")
output_file.write("};\n\n")
output_file.write(f"bool {prop}(IROps Op) {{\n")
output_file.write(f" return {array}[Op];\n")
output_file.write(f"{T} {prop}(IROps Op) {{\n")
output_file.write(f" return {prop}_[Op];\n")
output_file.write("}\n")
output_file.write("#undef IROP_HASSIDEEFFECTS_IMPL\n")
@@ -520,14 +545,20 @@ def print_ir_arg_printer():
output_file.write("\t*out << \" \";\n")
SSAArgNum = 0
FirstArg = True
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
# No point printing temporaries that we can't recover
if arg.Temporary:
# Temporary that we can't recover
output_file.write("\t*out << \"{}:Tmp:{}\";\n".format(arg.Type, arg.Name))
elif arg.IsSSA:
continue
if FirstArg:
FirstArg = False
else:
output_file.write('\t*out << ", ";\n')
if arg.IsSSA:
# SSA value
output_file.write("\tPrintArg(out, IR, Op->Header.Args[{}], RAData);\n".format(SSAArgNum))
SSAArgNum = SSAArgNum + 1
@@ -535,9 +566,6 @@ def print_ir_arg_printer():
# User defined op that is stored
output_file.write("\tPrintArg(out, IR, Op->{});\n".format(arg.Name))
if not LastArg:
output_file.write("\t*out << \", \";\n")
output_file.write("break;\n")
output_file.write("}\n")
@@ -636,11 +664,11 @@ def print_ir_allocator_helpers():
output_file.write("{} {}".format(CType, arg.Name));
elif arg.IsSSA:
# SSA value
output_file.write("OrderedNode *_{}".format(arg.Name))
output_file.write("OrderedNode *{}".format(arg.Name))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} _{}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name));
if arg.DefaultInitializer != None:
output_file.write(" = {}".format(arg.DefaultInitializer))
@@ -654,23 +682,28 @@ def print_ir_allocator_helpers():
if op.ImplicitFlagClobber:
output_file.write("\t\tSaveNZCV(IROps::OP_{});".format(op.Name.upper()))
output_file.write("\t\tauto Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
# We gather the "has x87?" flag as we go. This saves the user from
# having to keep track of whether they emitted any x87.
if op.LoweredX87:
output_file.write("\t\tRecordX87Use();\n")
output_file.write("\t\tauto _Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
if op.SSAArgNum != 0:
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tOp.first->{} = _{}->Wrapped(ListDataBegin);\n".format(arg.Name, arg.Name))
output_file.write("\t\t_Op.first->{} = {}->Wrapped(ListDataBegin);\n".format(arg.Name, arg.Name))
if op.SSAArgNum != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t_{}->AddUse();\n".format(arg.Name))
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
if len(op.Arguments) != 0:
for arg in op.Arguments:
if not arg.Temporary and not arg.IsSSA:
output_file.write("\t\tOp.first->{} = _{};\n".format(arg.Name, arg.Name))
output_file.write("\t\t_Op.first->{} = {};\n".format(arg.Name, arg.Name))
if (op.HasDest):
# We can only infer a size if we have arguments
@@ -680,22 +713,22 @@ def print_ir_allocator_helpers():
if len(op.Arguments) != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tuint8_t Size{} = GetOpSize(_{});\n".format(arg.Name, arg.Name))
output_file.write("\t\tuint8_t Size{} = GetOpSize({});\n".format(arg.Name, arg.Name))
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tInferSize = std::max(InferSize, Size{});\n".format(arg.Name))
output_file.write("\t\tOp.first->Header.Size = InferSize;\n")
output_file.write("\t\t_Op.first->Header.Size = InferSize;\n")
# Some ops without a destination still need an operating size
# Effectively reusing the destination size value for operation size
if op.DestSize != None:
output_file.write("\t\tOp.first->Header.Size = {};\n".format(op.DestSize))
output_file.write("\t\t_Op.first->Header.Size = {};\n".format(op.DestSize))
if op.NumElements == None:
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / ({});\n".format(1))
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(1))
else:
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / ({});\n".format(op.NumElements))
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(op.NumElements))
# Insert validation here
if op.EmitValidation != None:
@@ -706,58 +739,12 @@ def print_ir_allocator_helpers():
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
output_file.write("\t\treturn Op;\n")
output_file.write("\t\treturn _Op;\n")
output_file.write("\t}\n\n")
output_file.write("#undef IROP_ALLOCATE_HELPERS\n")
output_file.write("#endif\n")
def print_ir_parser_switch_helper():
output_file.write("#ifdef IROP_PARSER_SWITCH_HELPERS\n")
for op in IROps:
if op.Name != "Last" and op.SwitchGen:
output_file.write("\tcase FEXCore::IR::IROps::OP_%s: {\n" % (op.Name.upper()))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("\t\tauto arg{} = DecodeValue<{}>(Def.Args[{}]);\n".format(i, CType, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
elif arg.IsSSA:
# SSA value
output_file.write("\t\tauto arg{} = DecodeValue<OrderedNode*>(Def.Args[{}]);\n".format(i, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("\t\tauto arg{} = DecodeValue<{}>(Def.Args[{}]);\n".format(i, CType, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
output_file.write("\t\tDef.Node = _{}(\n".format(op.Name))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
output_file.write("\t\t\targ{}.second".format(i))
if not LastArg:
output_file.write(",\n")
else:
output_file.write("\n")
output_file.write("\t\t);\n")
output_file.write("\t\tSSANameMapper[Def.Definition] = Def.Node;\n")
output_file.write("\t\tbreak;\n")
output_file.write("\t}\n")
output_file.write("#undef IROP_PARSER_SWITCH_HELPERS\n")
output_file.write("#endif\n")
def print_ir_dispatcher_defs():
output_dispatch_file.write("#ifdef IROP_DISPATCH_DEFS\n")
for op in IROps:
@@ -816,7 +803,6 @@ print_ir_hassideeffects()
print_ir_gethasdest()
print_ir_arg_printer()
print_ir_allocator_helpers()
print_ir_parser_switch_helper()
output_file.close()
+4 -27
View File
@@ -67,7 +67,6 @@ set (SRCS
Common/SoftFloat-3e/s_approxRecipSqrt32_1.c
Common/SoftFloat-3e/s_approxRecipSqrt_1Ks.c
Common/SoftFloat-3e/softfloat_raiseFlags.c
Common/SoftFloat-3e/softfloat_state.c
Common/SoftFloat-3e/f64_to_extF80.c
Common/SoftFloat-3e/s_commonNaNToExtF80UI.c
Common/SoftFloat-3e/s_normSubnormalF64Sig.c
@@ -91,10 +90,10 @@ set (SRCS
Interface/Core/CPUBackend.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/HostFeatures.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
Interface/Core/OpcodeDispatcher/AVX_128.cpp
Interface/Core/OpcodeDispatcher/Crypto.cpp
Interface/Core/OpcodeDispatcher/Flags.cpp
Interface/Core/OpcodeDispatcher/Vector.cpp
@@ -112,7 +111,6 @@ set (SRCS
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
@@ -134,22 +132,16 @@ set (SRCS
Interface/GDBJIT/GDBJIT.cpp
Interface/IR/AOTIR.cpp
Interface/IR/IRDumper.cpp
Interface/IR/IRParser.cpp
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/DeadCodeElimination.cpp
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRCompaction.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/LongDivideRemovalPass.cpp
Interface/IR/Passes/ValueDominanceValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/InlineCallOptimization.cpp
Interface/IR/Passes/x87StackOptimizationPass.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
@@ -166,7 +158,7 @@ if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
Utils/AllocatorOverride.cpp)
endif()
set(DEFINES -DTHREAD_LOCAL=_Thread_local -DJIT_ARM64)
set(DEFINES -DJIT_ARM64)
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
@@ -176,11 +168,6 @@ if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
endif()
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
list(APPEND DEFINES -DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
@@ -194,7 +181,7 @@ endif()
# Some defines for the softfloat library
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils)
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils CodeEmitter)
if (NOT MINGW_BUILD)
list (APPEND LIBS dl)
@@ -373,16 +360,6 @@ function(AddLibrary Name Type)
target_link_libraries(${Name} FEXCore_Base)
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
if (MINGW_BUILD)
# Mingw build isn't building a linux shared library, so it can't have a SONAME.
set_target_properties(${Name} PROPERTIES NO_SONAME ON)
# Change the suffixes otherwise cmake continues using .a and .so
if (${Type} STREQUAL SHARED)
set_target_properties(${Name} PROPERTIES SUFFIX ".dll")
elseif(${Type} STREQUAL STATIC)
set_target_properties(${Name} PROPERTIES SUFFIX ".lib")
endif()
endif()
AddDefaultOptionsToTarget(${Name})
endfunction()
+7 -4
View File
@@ -20,12 +20,12 @@ struct BitSet final {
ElementType* Memory;
void Allocate(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
}
void Realloc(size_t Elements) {
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
size_t AllocateSize = ToBytes(Elements);
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
}
@@ -43,10 +43,13 @@ struct BitSet final {
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
}
void MemClear(size_t Elements) {
memset(Memory, 0, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0, ToBytes(Elements));
}
void MemSet(size_t Elements) {
memset(Memory, 0xFF, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
memset(Memory, 0xFF, ToBytes(Elements));
}
uint32_t ToBytes(size_t Elements) {
return AlignUp(Elements, MinimumSizeBits) / MinimumSize;
}
// This very explicitly doesn't let you take an address
@@ -41,7 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_add( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -53,7 +53,7 @@ extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
bool signB;
extFloat80_t
(*magsFuncPtr)(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
uA.f = a;
uiA64 = uA.s.signExp;
@@ -65,6 +65,6 @@ extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
signB = signExtF80UI64( uiB64 );
magsFuncPtr =
(signA == signB) ? softfloat_addMagsExtF80 : softfloat_subMagsExtF80;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
return (*magsFuncPtr)( state, uiA64, uiA0, uiB64, uiB0, signA );
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_div( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -107,7 +107,7 @@ extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
if ( ! (sigB & UINT64_C( 0x8000000000000000 )) ) {
if ( ! sigB ) {
if ( ! sigA ) goto invalid;
softfloat_raiseFlags( softfloat_flag_infinite );
softfloat_raiseFlags( state, softfloat_flag_infinite );
goto infinity;
}
normExpSig = softfloat_normSubnormalExtF80Sig( sigB );
@@ -169,18 +169,18 @@ extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
sigZExtra = (uint64_t) ((uint_fast64_t) q<<41);
return
softfloat_roundPackToExtF80(
signZ, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, signZ, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t a, extFloat80_t b )
bool extF80_eq( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -62,7 +62,7 @@ bool extF80_eq( extFloat80_t a, extFloat80_t b )
softfloat_isSigNaNExtF80UI( uiA64, uiA0 )
|| softfloat_isSigNaNExtF80UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t a, extFloat80_t b )
bool extF80_lt( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -59,7 +59,7 @@ bool extF80_lt( extFloat80_t a, extFloat80_t b )
uiB64 = uB.s.signExp;
uiB0 = uB.s.signif;
if ( isNaNExtF80UI( uiA64, uiA0 ) || isNaNExtF80UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signExtF80UI64( uiA64 );
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_mul( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -125,11 +125,11 @@ extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
}
return
softfloat_roundPackToExtF80(
signZ, expZ, sig128Z.v64, sig128Z.v0, extF80_roundingPrecision );
state, signZ, expZ, sig128Z.v64, sig128Z.v0, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
@@ -137,7 +137,7 @@ extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
*------------------------------------------------------------------------*/
infArg:
if ( ! magBits ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
} else {
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_rem( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -193,18 +193,18 @@ extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
}
return
softfloat_normRoundPackToExtF80(
signRem, rem.v64 | rem.v0 ? expB + 32 : 0, rem.v64, rem.v0, 80 );
state, signRem, rem.v64 | rem.v0 ? expB + 32 : 0, rem.v64, rem.v0, 80 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
extF80_roundToInt( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_roundToInt( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64, signUI64;
@@ -80,7 +80,7 @@ extFloat80_t
if ( 0x403E <= exp ) {
if ( exp == 0x7FFF ) {
if ( sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
uiZ = softfloat_propagateNaNExtF80UI( uiA64, sigA, 0, 0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, sigA, 0, 0 );
uiZ64 = uiZ.v64;
sigZ = uiZ.v0;
goto uiZ;
@@ -93,7 +93,7 @@ extFloat80_t
goto uiZ;
}
if ( exp <= 0x3FFE ) {
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
switch ( roundingMode ) {
case softfloat_round_near_even:
if ( !(sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF )) ) break;
@@ -145,7 +145,7 @@ extFloat80_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) sigZ |= lastBitMask;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
uiZ:
uZ.s.signExp = uiZ64;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t a )
extFloat80_t extF80_sqrt( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -74,7 +74,7 @@ extFloat80_t extF80_sqrt( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, 0, 0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, 0, 0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
@@ -155,11 +155,11 @@ extFloat80_t extF80_sqrt( extFloat80_t a )
}
return
softfloat_roundPackToExtF80(
0, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, 0, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -41,7 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_sub( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -54,7 +54,7 @@ extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
extFloat80_t
(*magsFuncPtr)(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
#endif
uA.f = a;
@@ -67,14 +67,14 @@ extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
signB = signExtF80UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_subMagsExtF80( uiA64, uiA0, uiB64, uiB0, signA );
return softfloat_subMagsExtF80( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_addMagsExtF80( uiA64, uiA0, uiB64, uiB0, signA );
return softfloat_addMagsExtF80( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_subMagsExtF80 : softfloat_addMagsExtF80;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
return (*magsFuncPtr)( state, uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t a )
float128_t extF80_to_f128( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -61,7 +61,7 @@ float128_t extF80_to_f128( extFloat80_t a )
exp = expExtF80UI64( uiA64 );
frac = uiA0 & UINT64_C( 0x7FFFFFFFFFFFFFFF );
if ( (exp == 0x7FFF) && frac ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF128UI( &commonNaN );
} else {
sign = signExtF80UI64( uiA64 );
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t a )
float32_t extF80_to_f32( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -66,7 +66,7 @@ float32_t extF80_to_f32( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF32UI( &commonNaN );
} else {
uiZ = packToF32UI( sign, 0xFF, 0 );
@@ -86,7 +86,7 @@ float32_t extF80_to_f32( extFloat80_t a )
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF32( sign, exp, sig32 );
return softfloat_roundPackToF32( state, sign, exp, sig32 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t a )
float64_t extF80_to_f64( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -72,7 +72,7 @@ float64_t extF80_to_f64( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF64UI( &commonNaN );
} else {
uiZ = packToF64UI( sign, 0x7FF, 0 );
@@ -86,7 +86,7 @@ float64_t extF80_to_f64( extFloat80_t a )
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF64( sign, exp, sig );
return softfloat_roundPackToF64( state, sign, exp, sig );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
extF80_to_i32( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_i32( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -68,7 +68,7 @@ int_fast32_t
#elif (i32_fromNaN == i32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return i32_fromNaN;
#endif
}
@@ -78,7 +78,7 @@ int_fast32_t
shiftDist = 0x4032 - exp;
if ( shiftDist <= 0 ) shiftDist = 1;
sig = softfloat_shiftRightJam64( sig, shiftDist );
return softfloat_roundToI32( sign, sig, roundingMode, exact );
return softfloat_roundToI32( state, sign, sig, roundingMode, exact );
}
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
extF80_to_i64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_i64( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -68,7 +68,7 @@ int_fast64_t
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
? i64_fromNaN
@@ -84,7 +84,7 @@ int_fast64_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
}
return softfloat_roundToI64( sign, sig, sigExtra, roundingMode, exact );
return softfloat_roundToI64( state, sign, sig, sigExtra, roundingMode, exact );
}
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t
extF80_to_ui64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_ui64( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -65,7 +65,7 @@ uint_fast64_t
*------------------------------------------------------------------------*/
shiftDist = 0x403E - exp;
if ( shiftDist < 0 ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
? ui64_fromNaN
@@ -79,7 +79,7 @@ uint_fast64_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
}
return softfloat_roundToUI64( sign, sig, sigExtra, roundingMode, exact );
return softfloat_roundToUI64( state, sign, sig, sigExtra, roundingMode, exact );
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t a )
extFloat80_t f128_to_extF80( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
@@ -70,7 +70,7 @@ extFloat80_t f128_to_extF80( float128_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 | frac0 ) {
softfloat_f128UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -98,7 +98,7 @@ extFloat80_t f128_to_extF80( float128_t a )
sig128 =
softfloat_shortShiftLeft128(
frac64 | UINT64_C( 0x0001000000000000 ), frac0, 15 );
return softfloat_roundPackToExtF80( sign, exp, sig128.v64, sig128.v0, 80 );
return softfloat_roundPackToExtF80( state, sign, exp, sig128.v64, sig128.v0, 80 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t a )
extFloat80_t f32_to_extF80( struct softfloat_state *state, float32_t a )
{
union ui32_f32 uA;
uint_fast32_t uiA;
@@ -67,7 +67,7 @@ extFloat80_t f32_to_extF80( float32_t a )
*------------------------------------------------------------------------*/
if ( exp == 0xFF ) {
if ( frac ) {
softfloat_f32UIToCommonNaN( uiA, &commonNaN );
softfloat_f32UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t a )
extFloat80_t f64_to_extF80( struct softfloat_state *state, float64_t a )
{
union ui64_f64 uA;
uint_fast64_t uiA;
@@ -67,7 +67,7 @@ extFloat80_t f64_to_extF80( float64_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FF ) {
if ( frac ) {
softfloat_f64UIToCommonNaN( uiA, &commonNaN );
softfloat_f64UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -63,19 +63,19 @@ uint_fast32_t softfloat_roundToUI32( bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
uint_fast64_t
softfloat_roundToUI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
struct softfloat_state *, bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
#else
uint_fast64_t softfloat_roundMToUI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t softfloat_roundToI32( bool, uint_fast64_t, uint_fast8_t, bool );
int_fast32_t softfloat_roundToI32( struct softfloat_state *, bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
struct softfloat_state *, bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
#else
int_fast64_t softfloat_roundMToI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
@@ -115,7 +115,7 @@ FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t softfloat_roundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_roundPackToF32( struct softfloat_state *, bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_normRoundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_addMagsF32( uint_fast32_t, uint_fast32_t );
@@ -138,7 +138,7 @@ FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t softfloat_roundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_roundPackToF64( struct softfloat_state *, bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_normRoundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_addMagsF64( uint_fast64_t, uint_fast64_t, bool );
@@ -167,18 +167,18 @@ struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
struct softfloat_state *, bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
struct softfloat_state *, bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
extFloat80_t
softfloat_addMagsExtF80(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
extFloat80_t
softfloat_subMagsExtF80(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
/*----------------------------------------------------------------------------
*----------------------------------------------------------------------------*/
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extFloat80_t
softfloat_addMagsExtF80(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -140,11 +141,11 @@ extFloat80_t
roundAndPack:
return
softfloat_roundPackToExtF80(
signZ, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, signZ, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
uiZ:
@@ -49,11 +49,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
struct softfloat_state *state, uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNExtF80UI( uiA64, uiA0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA64>>15;
zPtr->v64 = uiA0<<1;
@@ -50,12 +50,12 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
struct softfloat_state *state, uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
{
struct uint128 NaNSig;
if ( softfloat_isSigNaNF128UI( uiA64, uiA0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
NaNSig = softfloat_shortShiftLeft128( uiA64, uiA0, 16 );
zPtr->sign = uiA64>>63;
@@ -46,11 +46,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr )
void softfloat_f32UIToCommonNaN( struct softfloat_state *state, uint_fast32_t uiA, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNF32UI( uiA ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA>>31;
zPtr->v64 = (uint_fast64_t) uiA<<41;
@@ -46,11 +46,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr )
void softfloat_f64UIToCommonNaN( struct softfloat_state *state, uint_fast64_t uiA, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNF64UI( uiA ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA>>63;
zPtr->v64 = uiA<<12;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
struct softfloat_state *state,
bool sign,
int_fast32_t exp,
uint_fast64_t sig,
@@ -66,7 +67,7 @@ extFloat80_t
}
return
softfloat_roundPackToExtF80(
sign, exp, sig, sigExtra, roundingPrecision );
state, sign, exp, sig, sigExtra, roundingPrecision );
}
@@ -53,6 +53,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -76,7 +77,7 @@ struct uint128
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( isSigNaNA | isSigNaNB ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
if ( isSigNaNA ) {
if ( isSigNaNB ) goto returnLargerMag;
if ( isNaNExtF80UI( uiB64, uiB0 ) ) goto returnB;
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
struct softfloat_state *state,
bool sign,
int_fast32_t exp,
uint_fast64_t sig,
@@ -59,7 +60,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
if ( roundingPrecision == 80 ) goto precision80;
if ( roundingPrecision == 64 ) {
@@ -87,15 +88,15 @@ extFloat80_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess
(state->detectTininess
== softfloat_tininess_beforeRounding)
|| (exp < 0)
|| (sig <= (uint64_t) (sig + roundIncrement));
sig = softfloat_shiftRightJam64( sig, 1 - exp );
roundBits = sig & roundMask;
if ( roundBits ) {
if ( isTiny ) softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( isTiny ) softfloat_raiseFlags( state, softfloat_flag_underflow );
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= roundMask + 1;
@@ -121,7 +122,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig = (sig & ~roundMask) | (roundMask + 1);
@@ -157,7 +158,7 @@ extFloat80_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess
(state->detectTininess
== softfloat_tininess_beforeRounding)
|| (exp < 0)
|| ! doIncrement
@@ -168,8 +169,8 @@ extFloat80_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
if ( sigExtra ) {
if ( isTiny ) softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( isTiny ) softfloat_raiseFlags( state, softfloat_flag_underflow );
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -207,7 +208,7 @@ extFloat80_t
roundMask = 0;
overflow:
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
if (
roundNearEven
|| (roundingMode == softfloat_round_near_maxMag)
@@ -226,7 +227,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( sigExtra ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
float32_t
softfloat_roundPackToF32( bool sign, int_fast16_t exp, uint_fast32_t sig )
softfloat_roundPackToF32( struct softfloat_state *state, bool sign, int_fast16_t exp, uint_fast32_t sig )
{
uint_fast8_t roundingMode;
bool roundNearEven;
@@ -53,7 +53,7 @@ float32_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
roundIncrement = 0x40;
if ( ! roundNearEven && (roundingMode != softfloat_round_near_maxMag) ) {
@@ -71,19 +71,19 @@ float32_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess == softfloat_tininess_beforeRounding)
(state->detectTininess == softfloat_tininess_beforeRounding)
|| (exp < -1) || (sig + roundIncrement < 0x80000000);
sig = softfloat_shiftRightJam32( sig, -exp );
exp = 0;
roundBits = sig & 0x7F;
if ( isTiny && roundBits ) {
softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_raiseFlags( state, softfloat_flag_underflow );
}
} else if ( (0xFD < exp) || (0x80000000 <= sig + roundIncrement) ) {
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
uiZ = packToF32UI( sign, 0xFF, 0 ) - ! roundIncrement;
goto uiZ;
}
@@ -92,7 +92,7 @@ float32_t
*------------------------------------------------------------------------*/
sig = (sig + roundIncrement)>>7;
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
float64_t
softfloat_roundPackToF64( bool sign, int_fast16_t exp, uint_fast64_t sig )
softfloat_roundPackToF64( struct softfloat_state *state, bool sign, int_fast16_t exp, uint_fast64_t sig )
{
uint_fast8_t roundingMode;
bool roundNearEven;
@@ -53,7 +53,7 @@ float64_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
roundIncrement = 0x200;
if ( ! roundNearEven && (roundingMode != softfloat_round_near_maxMag) ) {
@@ -71,14 +71,14 @@ float64_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess == softfloat_tininess_beforeRounding)
(state->detectTininess == softfloat_tininess_beforeRounding)
|| (exp < -1)
|| (sig + roundIncrement < UINT64_C( 0x8000000000000000 ));
sig = softfloat_shiftRightJam64( sig, -exp );
exp = 0;
roundBits = sig & 0x3FF;
if ( isTiny && roundBits ) {
softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_raiseFlags( state, softfloat_flag_underflow );
}
} else if (
(0x7FD < exp)
@@ -87,7 +87,7 @@ float64_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
uiZ = packToF64UI( sign, 0x7FF, 0 ) - ! roundIncrement;
goto uiZ;
}
@@ -96,7 +96,7 @@ float64_t
*------------------------------------------------------------------------*/
sig = (sig + roundIncrement)>>10;
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -44,7 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
softfloat_roundToI32(
bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
struct softfloat_state *state, bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
{
uint_fast16_t roundIncrement, roundBits;
uint_fast32_t sig32;
@@ -86,13 +86,13 @@ int_fast32_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) z |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? i32_fromNegOverflow : i32_fromPosOverflow;
}
@@ -44,6 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
struct softfloat_state *state,
bool sign,
uint_fast64_t sig,
uint_fast64_t sigExtra,
@@ -89,13 +90,13 @@ int_fast64_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) z |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? i64_fromNegOverflow : i64_fromPosOverflow;
}
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
uint_fast64_t
softfloat_roundToUI64(
struct softfloat_state *state,
bool sign,
uint_fast64_t sig,
uint_fast64_t sigExtra,
@@ -84,13 +85,13 @@ uint_fast64_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) sig |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return sig;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? ui64_fromNegOverflow : ui64_fromPosOverflow;
}
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extFloat80_t
softfloat_subMagsExtF80(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -77,7 +78,7 @@ extFloat80_t
if ( (sigA | sigB) & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
goto propagateNaN;
}
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -90,7 +91,7 @@ extFloat80_t
if ( sigB < sigA ) goto aBigger;
if ( sigA < sigB ) goto bBigger;
uiZ64 =
packToExtF80UI64( (softfloat_roundingMode == softfloat_round_min), 0 );
packToExtF80UI64( (state->roundingMode == softfloat_round_min), 0 );
uiZ0 = 0;
goto uiZ;
/*------------------------------------------------------------------------
@@ -142,11 +143,11 @@ extFloat80_t
normRoundPack:
return
softfloat_normRoundPackToExtF80(
signZ, expZ, sig128.v64, sig128.v0, extF80_roundingPrecision );
state, signZ, expZ, sig128.v64, sig128.v0, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
uiZ:
+19 -64
View File
@@ -50,50 +50,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include <stdint.h>
#include "softfloat_types.h"
#ifndef THREAD_LOCAL
#define THREAD_LOCAL
#endif
/*----------------------------------------------------------------------------
| Software floating-point underflow tininess-detection mode.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_detectTininess;
enum {
softfloat_tininess_beforeRounding = 0,
softfloat_tininess_afterRounding = 1
};
/*----------------------------------------------------------------------------
| Software floating-point rounding mode. (Mode "odd" is supported only if
| SoftFloat is compiled with macro 'SOFTFLOAT_ROUND_ODD' defined.)
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_roundingMode;
enum {
softfloat_round_near_even = 0,
softfloat_round_minMag = 1,
softfloat_round_min = 2,
softfloat_round_max = 3,
softfloat_round_near_maxMag = 4,
softfloat_round_odd = 6
};
/*----------------------------------------------------------------------------
| Software floating-point exception flags.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_exceptionFlags;
enum {
softfloat_flag_inexact = 1,
softfloat_flag_underflow = 2,
softfloat_flag_overflow = 4,
softfloat_flag_infinite = 8,
softfloat_flag_invalid = 16
};
/*----------------------------------------------------------------------------
| Routine to raise any or all of the software floating-point exception flags.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t );
void softfloat_raiseFlags( struct softfloat_state *, uint_fast8_t );
/*----------------------------------------------------------------------------
| Integer-to-floating-point conversion routines.
@@ -187,7 +148,7 @@ float16_t f32_to_f16( float32_t );
float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t );
extFloat80_t f32_to_extF80( struct softfloat_state *, float32_t );
float128_t f32_to_f128( float32_t );
#endif
void f32_to_extF80M( float32_t, extFloat80_t * );
@@ -223,7 +184,7 @@ float16_t f64_to_f16( float64_t );
float32_t f64_to_f32( float64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t );
extFloat80_t f64_to_extF80( struct softfloat_state *, float64_t );
float128_t f64_to_f128( float64_t );
#endif
void f64_to_extF80M( float64_t, extFloat80_t * );
@@ -244,53 +205,47 @@ bool f64_le_quiet( float64_t, float64_t );
bool f64_lt_quiet( float64_t, float64_t );
bool f64_isSignalingNaN( float64_t );
/*----------------------------------------------------------------------------
| Rounding precision for 80-bit extended double-precision floating-point.
| Valid values are 32, 64, and 80.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t extF80_roundingPrecision;
/*----------------------------------------------------------------------------
| 80-bit extended double-precision floating-point operations.
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t extF80_to_ui32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t extF80_to_ui64( extFloat80_t, uint_fast8_t, bool );
uint_fast64_t extF80_to_ui64( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t extF80_to_i32( extFloat80_t, uint_fast8_t, bool );
int_fast32_t extF80_to_i32( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t extF80_to_i64( extFloat80_t, uint_fast8_t, bool );
int_fast64_t extF80_to_i64( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
uint_fast32_t extF80_to_ui32_r_minMag( extFloat80_t, bool );
uint_fast64_t extF80_to_ui64_r_minMag( extFloat80_t, bool );
int_fast32_t extF80_to_i32_r_minMag( extFloat80_t, bool );
int_fast64_t extF80_to_i64_r_minMag( extFloat80_t, bool );
float16_t extF80_to_f16( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t );
float32_t extF80_to_f32( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t );
float64_t extF80_to_f64( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t );
float128_t extF80_to_f128( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_roundToInt( extFloat80_t, uint_fast8_t, bool );
extFloat80_t extF80_roundToInt( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t, extFloat80_t );
extFloat80_t extF80_add( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t, extFloat80_t );
extFloat80_t extF80_sub( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t, extFloat80_t );
extFloat80_t extF80_mul( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t, extFloat80_t );
extFloat80_t extF80_div( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t, extFloat80_t );
extFloat80_t extF80_rem( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t );
extFloat80_t extF80_sqrt( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t, extFloat80_t );
bool extF80_eq( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t, extFloat80_t );
bool extF80_lt( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
bool extF80_lt_quiet( extFloat80_t, extFloat80_t );
@@ -341,7 +296,7 @@ float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t );
extFloat80_t f128_to_extF80( struct softfloat_state *, float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
float128_t f128_sub( float128_t, float128_t );
@@ -44,10 +44,10 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| should be simply `softfloat_exceptionFlags |= flags;'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t flags )
void softfloat_raiseFlags( struct softfloat_state *state, uint_fast8_t flags )
{
softfloat_exceptionFlags |= flags;
state->exceptionFlags |= flags;
}
@@ -1,52 +0,0 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All Rights Reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
#ifndef THREAD_LOCAL
#define THREAD_LOCAL
#endif
THREAD_LOCAL uint_fast8_t softfloat_roundingMode = softfloat_round_near_even;
THREAD_LOCAL uint_fast8_t softfloat_detectTininess = init_detectTininess;
THREAD_LOCAL uint_fast8_t softfloat_exceptionFlags = 0;
THREAD_LOCAL uint_fast8_t extF80_roundingPrecision = 80;
@@ -77,5 +77,50 @@ struct extFloat80M { uint16_t signExp; uint64_t signif; };
*----------------------------------------------------------------------------*/
typedef struct extFloat80M extFloat80_t;
enum {
softfloat_tininess_beforeRounding = 0,
softfloat_tininess_afterRounding = 1
};
enum {
softfloat_round_near_even = 0,
softfloat_round_minMag = 1,
softfloat_round_min = 2,
softfloat_round_max = 3,
softfloat_round_near_maxMag = 4,
softfloat_round_odd = 6
};
enum {
softfloat_flag_inexact = 1,
softfloat_flag_underflow = 2,
softfloat_flag_overflow = 4,
softfloat_flag_infinite = 8,
softfloat_flag_invalid = 16
};
struct softfloat_state {
/*----------------------------------------------------------------------------
| Software floating-point underflow tininess-detection mode.
*----------------------------------------------------------------------------*/
uint8_t detectTininess; /* = init_detectTininess */
/*----------------------------------------------------------------------------
| Software floating-point rounding mode. (Mode "odd" is supported only if
| SoftFloat is compiled with macro 'SOFTFLOAT_ROUND_ODD' defined.)
*----------------------------------------------------------------------------*/
uint8_t roundingMode; /* = softfloat_round_near_even */
/*----------------------------------------------------------------------------
| Software floating-point exception flags.
*----------------------------------------------------------------------------*/
uint8_t exceptionFlags; /* = 0 */
/*----------------------------------------------------------------------------
| Rounding precision for 80-bit extended double-precision floating-point.
| Valid values are 32, 64, and 80.
*----------------------------------------------------------------------------*/
uint8_t roundingPrecision; /* = 80 */
};
#endif
@@ -136,7 +136,7 @@ uint_fast16_t
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr );
void softfloat_f32UIToCommonNaN( struct softfloat_state *, uint_fast32_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 32-bit floating-point
@@ -173,7 +173,7 @@ uint_fast32_t
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr );
void softfloat_f64UIToCommonNaN( struct softfloat_state *, uint_fast64_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 64-bit floating-point
@@ -222,7 +222,7 @@ uint_fast64_t
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
struct softfloat_state *, uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into an 80-bit extended
@@ -244,6 +244,7 @@ struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
struct softfloat_state *,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -274,7 +275,7 @@ struct uint128
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
struct softfloat_state *, uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 128-bit floating-point
+73 -79
View File
@@ -63,7 +63,7 @@ struct FEX_PACKED X80SoftFloat {
}
// Ops
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FADD(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FADD(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -79,11 +79,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_add(lhs, rhs);
return extF80_add(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSUB(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSUB(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -99,11 +99,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_sub(lhs, rhs);
return extF80_sub(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FMUL(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FMUL(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -119,11 +119,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_mul(lhs, rhs);
return extF80_mul(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FDIV(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FDIV(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -139,11 +139,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_div(lhs, rhs);
return extF80_div(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm(R"(
@@ -160,11 +160,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_rem(lhs, rhs);
return extF80_rem(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM1(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM1(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm(R"(
@@ -181,16 +181,16 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_rem(lhs, rhs);
return extF80_rem(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(softfloat_state* state, const X80SoftFloat& lhs) {
return extF80_roundToInt(state, lhs, state->roundingMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(softfloat_state* state, const X80SoftFloat& lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(state, lhs, RoundMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FXTRACT_SIG(const X80SoftFloat& lhs) {
@@ -237,13 +237,14 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static void FCMP(const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(lhs, rhs);
*lt = extF80_lt(lhs, rhs);
FEXCORE_PRESERVE_ALL_ATTR static void
FCMP(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(state, lhs, rhs);
*lt = extF80_lt(state, lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -261,16 +262,16 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
LIBRARY_PRECISION Src2_d = Int;
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
LIBRARY_PRECISION Src2_d = Int.ToFMax(state);
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
X80SoftFloat Src2_X80(state, Src2_d);
X80SoftFloat Result = extF80_mul(state, lhs, Src2_X80);
return Result;
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat F2XM1(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat F2XM1(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -286,14 +287,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
return X80SoftFloat(state, Result);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FYL2X(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FYL2X(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -310,14 +311,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Src2_d = rhs.ToFMax(state);
LIBRARY_PRECISION Tmp = Src2_d * log2l(Src1_d);
return Tmp;
return X80SoftFloat(state, Tmp);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FATAN(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FATAN(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -334,14 +335,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Src2_d = rhs.ToFMax(state);
LIBRARY_PRECISION Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
return X80SoftFloat(state, Tmp);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FTAN(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FTAN(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -358,13 +359,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = tanl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSIN(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSIN(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -380,13 +381,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = sinl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FCOS(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FCOS(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -402,13 +403,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = cosl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSQRT(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSQRT(softfloat_state* state, const X80SoftFloat& lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -423,62 +424,55 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_sqrt(lhs);
return extF80_sqrt(state, lhs);
#endif
}
operator float() const {
const float32_t Result = extF80_to_f32(*this);
float ToF32(softfloat_state* state) const {
const float32_t Result = extF80_to_f32(state, *this);
return FEXCore::BitCast<float>(Result);
}
operator double() const {
const float64_t Result = extF80_to_f64(*this);
double ToF64(softfloat_state* state) const {
const float64_t Result = extF80_to_f64(state, *this);
return FEXCore::BitCast<double>(Result);
}
#ifndef _WIN32
operator BIGFLOAT() const {
LIBRARY_PRECISION ToFMax(softfloat_state* state) const {
#ifdef _WIN32
return ToF64(state);
#else
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(*this);
const float128_t Result = extF80_to_f128(state, *this);
return FEXCore::BitCast<BIGFLOAT>(Result);
#else
BIGFLOAT result {};
memcpy(&result, this, sizeof(result));
return result;
#endif
}
#endif
}
operator int16_t() const {
auto rv = extF80_to_i32(*this, softfloat_roundingMode, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
int16_t ToI16(softfloat_state* state) const {
auto rv = extF80_to_i32(state, *this, state->roundingMode, false);
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
return INT16_MIN;
} else {
return rv;
}
}
operator int32_t() const {
return extF80_to_i32(*this, softfloat_roundingMode, false);
int32_t ToI32(softfloat_state* state) const {
return extF80_to_i32(state, *this, state->roundingMode, false);
}
operator int64_t() const {
return extF80_to_i64(*this, softfloat_roundingMode, false);
int64_t ToI64(softfloat_state* state) const {
return extF80_to_i64(state, *this, state->roundingMode, false);
}
operator uint64_t() const {
return extF80_to_ui64(*this, softfloat_roundingMode, false);
}
void operator=(const float rhs) {
*this = f32_to_extF80(FEXCore::BitCast<float32_t>(rhs));
}
void operator=(const double rhs) {
*this = f64_to_extF80(FEXCore::BitCast<float64_t>(rhs));
uint64_t ToUI64(softfloat_state* state) const {
return extF80_to_ui64(state, *this, state->roundingMode, false);
}
void operator=(const int16_t rhs) {
@@ -509,18 +503,18 @@ struct FEX_PACKED X80SoftFloat {
Sign = rhs.signExp >> 15;
}
X80SoftFloat(const float rhs) {
*this = f32_to_extF80(FEXCore::BitCast<float32_t>(rhs));
X80SoftFloat(softfloat_state* state, const float rhs) {
*this = f32_to_extF80(state, FEXCore::BitCast<float32_t>(rhs));
}
X80SoftFloat(const double rhs) {
*this = f64_to_extF80(FEXCore::BitCast<float64_t>(rhs));
X80SoftFloat(softfloat_state* state, const double rhs) {
*this = f64_to_extF80(state, FEXCore::BitCast<float64_t>(rhs));
}
#ifndef _WIN32
X80SoftFloat(BIGFLOAT rhs) {
X80SoftFloat(softfloat_state* state, BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(FEXCore::BitCast<float128_t>(rhs));
*this = f128_to_extF80(state, FEXCore::BitCast<float128_t>(rhs));
#else
*this = FEXCore::BitCast<long double>(rhs);
#endif
@@ -50,8 +50,6 @@
"DISABLESVE": "disablesve",
"ENABLEAVX": "enableavx",
"DISABLEAVX": "disableavx",
"ENABLEAVX2": "enableavx2",
"DISABLEAVX2": "disableavx2",
"ENABLEAFP": "enableafp",
"DISABLEAFP": "disableafp",
"ENABLELRCPC": "enablelrcpc",
@@ -78,6 +76,8 @@
"DISABLECRYPTO": "disablecrypto",
"ENABLERPRES": "enablerpres",
"DISABLERPRES": "disablerpres",
"ENABLESVEBITPERM": "enablesvebitperm",
"DISABLESVEBITPERM": "disablesvebitperm",
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi"
},
@@ -86,7 +86,6 @@
"\toff: Default CPU features queried from CPU features",
"\t{enable,disable}sve: Will force enable or disable sve even if the host doesn't support it",
"\t{enable,disable}avx: Will force enable or disable avx even if the host doesn't support it",
"\t{enable,disable}avx2: Will force enable or disable avx2 even if the host doesn't support it",
"\t{enable,disable}afp: Will force enable or disable afp even if the host doesn't support it",
"\t{enable,disable}lrcpc: Will force enable or disable lrcpc even if the host doesn't support it",
"\t{enable,disable}lrcpc2: Will force enable or disable lrcpc2 even if the host doesn't support it",
@@ -100,6 +99,7 @@
"\t{enable,disable}flagm2: Will force enable or disable flagm2 even if the host doesn't support it",
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it"
]
},
@@ -401,14 +401,14 @@
},
"VectorTSOEnabled": {
"Type": "bool",
"Default": "true",
"Default": "false",
"Desc": [
"When TSO emulation is enabled, controls if vector loadstores should also be atomic."
]
},
"MemcpySetTSOEnabled": {
"Type": "bool",
"Default": "true",
"Default": "false",
"Desc": [
"When TSO emulation is enabled, controls if memcpy and memset should also be atomic.",
"Only affects REP MOVS and REP STOS instructions"
+3 -6
View File
@@ -6,6 +6,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "FEXCore/Debug/InternalThreadState.h"
@@ -18,8 +19,8 @@ void InitializeStaticTables(OperatingMode Mode) {
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext() {
return fextl::make_unique<FEXCore::Context::ContextImpl>();
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext(const FEXCore::HostFeatures& Features) {
return fextl::make_unique<FEXCore::Context::ContextImpl>(Features);
}
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
@@ -42,10 +43,6 @@ void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryT
CustomCPUFactory = std::move(Factory);
}
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator* _SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
+8 -16
View File
@@ -57,6 +57,7 @@ namespace HLE {
namespace FEXCore::IR {
class RegisterAllocationData;
struct IRListCopy;
class IRListView;
namespace Validation {
class IRValidation;
@@ -95,14 +96,15 @@ public:
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
HostFeatures GetHostFeatures() const override;
void HandleCallback(FEXCore::Core::InternalThreadState* Thread, uint64_t RIP) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, uint64_t* HostGPRs, uint64_t PSTATE) override;
void SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) override;
void ReconstructXMMRegisters(const FEXCore::Core::InternalThreadState* Thread, __uint128_t* XMM_Low, __uint128_t* YMM_High) override;
void SetXMMRegistersFromState(FEXCore::Core::InternalThreadState* Thread, const __uint128_t* XMM_Low, const __uint128_t* YMM_High) override;
/**
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
*
@@ -191,7 +193,7 @@ public:
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void* Creator = nullptr, void* Data = nullptr);
void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) override;
void AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) override;
public:
friend class FEXCore::HLE::SyscallHandler;
@@ -205,9 +207,6 @@ public:
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
uint64_t VirtualMemSize {1ULL << 36};
// this is for internal use
bool ValidateIRarser {false};
// Used if the JIT needs to have its interrupt fault code emitted.
bool NeedsPendingInterruptFaultCheck {false};
@@ -263,7 +262,7 @@ public:
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl();
ContextImpl(const FEXCore::HostFeatures& Features);
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
@@ -282,9 +281,6 @@ public:
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LOGMAN_THROW_A_FMT(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}",
Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
ThreadRemoveCodeEntry(Thread, GuestRIP);
@@ -293,8 +289,7 @@ public:
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
struct GenerateIRResult {
FEXCore::IR::IRListView* IRList;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
uint64_t TotalInstructions;
uint64_t TotalInstructionsLength;
uint64_t StartAddr;
@@ -305,9 +300,8 @@ public:
struct CompileCodeResult {
void* CompiledCode;
FEXCore::IR::IRListView* IRData;
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData;
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
bool GeneratedIR;
uint64_t StartAddr;
uint64_t Length;
@@ -379,8 +373,6 @@ public:
return ExitOnHLT;
}
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -2,8 +2,6 @@
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "FEXCore/Core/X86Enums.h"
#include "FEXCore/Utils/AllocatorHooks.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Context/Context.h"
#include "Interface/HLE/Thunks/Thunks.h"
@@ -13,128 +11,118 @@
#include <FEXCore/Utils/MathUtils.h>
#include <FEXHeaderUtils/BitUtils.h>
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#ifdef VIXL_DISASSEMBLER
#include <aarch64/cpu-aarch64.h>
#include <aarch64/instructions-aarch64.h>
#include <cpu-features.h>
#include <utils-vixl.h>
#endif
#include <array>
#include <tuple>
#include <utility>
namespace FEXCore::CPU {
// Register x18 is unused in the current configuration.
// This is due to it being a platform register on wine platforms.
// TODO: Allow x18 register allocation on Linux in the future to gain one more register.
namespace x64 {
#ifndef _M_ARM_64EC
// All but x19 and x29 are caller saved
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
FEXCore::ARMEmitter::Reg::r4,
FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6,
FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
FEXCore::ARMEmitter::Reg::r9,
FEXCore::ARMEmitter::Reg::r10,
FEXCore::ARMEmitter::Reg::r11,
FEXCore::ARMEmitter::Reg::r12,
FEXCore::ARMEmitter::Reg::r13,
FEXCore::ARMEmitter::Reg::r14,
FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16,
FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r19,
FEXCore::ARMEmitter::Reg::r29,
constexpr std::array<ARMEmitter::Register, 18> SRA = {
ARMEmitter::Reg::r4,
ARMEmitter::Reg::r5,
ARMEmitter::Reg::r6,
ARMEmitter::Reg::r7,
ARMEmitter::Reg::r8,
ARMEmitter::Reg::r9,
ARMEmitter::Reg::r10,
ARMEmitter::Reg::r11,
ARMEmitter::Reg::r12,
ARMEmitter::Reg::r13,
ARMEmitter::Reg::r14,
ARMEmitter::Reg::r15,
ARMEmitter::Reg::r16,
ARMEmitter::Reg::r17,
ARMEmitter::Reg::r19,
ARMEmitter::Reg::r29,
// PF/AF must be last.
REG_PF,
REG_AF,
};
constexpr std::array<FEXCore::ARMEmitter::Register, 7> RA = {
constexpr std::array<ARMEmitter::Register, 8> RA = {
// All these callee saved
FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21, FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23,
FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25, FEXCore::ARMEmitter::Reg::r30,
ARMEmitter::Reg::r20, ARMEmitter::Reg::r21, ARMEmitter::Reg::r22, ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24, ARMEmitter::Reg::r25, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
};
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 3> RAPair = {{
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
}};
constexpr unsigned RAPairs = 6;
// All are caller saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
ARMEmitter::VReg::v16, ARMEmitter::VReg::v17, ARMEmitter::VReg::v18, ARMEmitter::VReg::v19,
ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22, ARMEmitter::VReg::v23,
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27,
ARMEmitter::VReg::v28, ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
constexpr std::array<ARMEmitter::VRegister, 14> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// ARMEmitter::VReg::v0, ARMEmitter::VReg::v1,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
};
#else
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
FEXCore::ARMEmitter::Reg::r8,
FEXCore::ARMEmitter::Reg::r0,
FEXCore::ARMEmitter::Reg::r1,
FEXCore::ARMEmitter::Reg::r27,
constexpr std::array<ARMEmitter::Register, 18> SRA = {
ARMEmitter::Reg::r8,
ARMEmitter::Reg::r0,
ARMEmitter::Reg::r1,
ARMEmitter::Reg::r27,
// SP's register location isn't specified by the ARM64EC ABI, we choose to use r23
FEXCore::ARMEmitter::Reg::r23,
FEXCore::ARMEmitter::Reg::r29,
FEXCore::ARMEmitter::Reg::r25,
FEXCore::ARMEmitter::Reg::r26,
FEXCore::ARMEmitter::Reg::r2,
FEXCore::ARMEmitter::Reg::r3,
FEXCore::ARMEmitter::Reg::r4,
FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r19,
FEXCore::ARMEmitter::Reg::r20,
FEXCore::ARMEmitter::Reg::r21,
FEXCore::ARMEmitter::Reg::r22,
ARMEmitter::Reg::r23,
ARMEmitter::Reg::r29,
ARMEmitter::Reg::r25,
ARMEmitter::Reg::r26,
ARMEmitter::Reg::r2,
ARMEmitter::Reg::r3,
ARMEmitter::Reg::r4,
ARMEmitter::Reg::r5,
ARMEmitter::Reg::r19,
ARMEmitter::Reg::r20,
ARMEmitter::Reg::r21,
ARMEmitter::Reg::r22,
REG_PF,
REG_AF,
};
constexpr std::array<FEXCore::ARMEmitter::Register, 7> RA = {
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r30,
constexpr std::array<ARMEmitter::Register, 7> RA = {
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15,
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
};
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 3> RAPair = {{
{FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7},
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17},
}};
constexpr unsigned RAPairs = 6;
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1, FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9, FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13, FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
ARMEmitter::VReg::v0, ARMEmitter::VReg::v1, ARMEmitter::VReg::v2, ARMEmitter::VReg::v3,
ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
};
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19, FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21,
FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23, FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27, FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
constexpr std::array<ARMEmitter::VRegister, 14> RAFPR = {
ARMEmitter::VReg::v18, ARMEmitter::VReg::v19, ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22,
ARMEmitter::VReg::v23, ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27,
ARMEmitter::VReg::v28, ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
#endif
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 7> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
constexpr std::array<ARMEmitter::Register, 7> PreserveAll_SRA = {
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r6, ARMEmitter::Reg::r7,
ARMEmitter::Reg::r8, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
};
constexpr uint32_t PreserveAll_SRAMask = {[]() -> uint32_t {
@@ -160,12 +148,12 @@ namespace x64 {
}()};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 1> PreserveAll_Dynamic = {
constexpr std::array<ARMEmitter::Register, 1> PreserveAll_Dynamic = {
// Only LR needs to get saved.
FEXCore::ARMEmitter::Reg::r30};
ARMEmitter::Reg::r30};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
constexpr std::array<ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
@@ -179,15 +167,14 @@ namespace x64 {
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
constexpr std::array<ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4,
FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr std::array<ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {[]() -> uint32_t {
uint32_t Mask {};
@@ -198,89 +185,77 @@ namespace x64 {
}()};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
constexpr std::array<ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
};
} // namespace x64
namespace x32 {
// All but x19 and x29 are caller saved
constexpr std::array<FEXCore::ARMEmitter::Register, 10> SRA = {
FEXCore::ARMEmitter::Reg::r4,
FEXCore::ARMEmitter::Reg::r5,
FEXCore::ARMEmitter::Reg::r6,
FEXCore::ARMEmitter::Reg::r7,
FEXCore::ARMEmitter::Reg::r8,
FEXCore::ARMEmitter::Reg::r9,
FEXCore::ARMEmitter::Reg::r10,
FEXCore::ARMEmitter::Reg::r11,
constexpr std::array<ARMEmitter::Register, 10> SRA = {
ARMEmitter::Reg::r4,
ARMEmitter::Reg::r5,
ARMEmitter::Reg::r6,
ARMEmitter::Reg::r7,
ARMEmitter::Reg::r8,
ARMEmitter::Reg::r9,
ARMEmitter::Reg::r10,
ARMEmitter::Reg::r11,
// PF/AF must be last.
REG_PF,
REG_AF,
};
constexpr std::array<FEXCore::ARMEmitter::Register, 15> RA = {
constexpr std::array<ARMEmitter::Register, 15> RA = {
// All these callee saved
FEXCore::ARMEmitter::Reg::r20,
FEXCore::ARMEmitter::Reg::r21,
FEXCore::ARMEmitter::Reg::r22,
FEXCore::ARMEmitter::Reg::r23,
FEXCore::ARMEmitter::Reg::r24,
FEXCore::ARMEmitter::Reg::r25,
ARMEmitter::Reg::r20,
ARMEmitter::Reg::r21,
ARMEmitter::Reg::r22,
ARMEmitter::Reg::r23,
ARMEmitter::Reg::r24,
ARMEmitter::Reg::r25,
// Registers only available on 32-bit
// All these are caller saved (except for r19).
FEXCore::ARMEmitter::Reg::r12,
FEXCore::ARMEmitter::Reg::r13,
FEXCore::ARMEmitter::Reg::r14,
FEXCore::ARMEmitter::Reg::r15,
FEXCore::ARMEmitter::Reg::r16,
FEXCore::ARMEmitter::Reg::r17,
FEXCore::ARMEmitter::Reg::r29,
FEXCore::ARMEmitter::Reg::r30,
ARMEmitter::Reg::r12,
ARMEmitter::Reg::r13,
ARMEmitter::Reg::r14,
ARMEmitter::Reg::r15,
ARMEmitter::Reg::r16,
ARMEmitter::Reg::r17,
ARMEmitter::Reg::r29,
ARMEmitter::Reg::r30,
FEXCore::ARMEmitter::Reg::r19,
ARMEmitter::Reg::r19,
};
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 7> RAPair = {{
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
{FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r13},
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17},
{FEXCore::ARMEmitter::Reg::r29, FEXCore::ARMEmitter::Reg::r30},
}};
constexpr unsigned RAPairs = 12;
// All are caller saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> SRAFPR = {
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
constexpr std::array<ARMEmitter::VRegister, 8> SRAFPR = {
ARMEmitter::VReg::v16, ARMEmitter::VReg::v17, ARMEmitter::VReg::v18, ARMEmitter::VReg::v19,
ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22, ARMEmitter::VReg::v23,
};
// v8..v15 = (lower 64bits) Callee saved
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> RAFPR = {
constexpr std::array<ARMEmitter::VRegister, 22> RAFPR = {
// v0 ~ v1 are used as temps.
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
// ARMEmitter::VReg::v0, ARMEmitter::VReg::v1,
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27, ARMEmitter::VReg::v28,
ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 5> PreserveAll_SRA = {
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6,
FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r8,
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_SRA = {
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r8,
};
constexpr uint32_t PreserveAll_SRAMask = {[]() -> uint32_t {
@@ -306,11 +281,10 @@ namespace x32 {
}()};
// Dynamic GPRs
constexpr std::array<FEXCore::ARMEmitter::Register, 3> PreserveAll_Dynamic = {
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r30};
constexpr std::array<ARMEmitter::Register, 3> PreserveAll_Dynamic = {ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30};
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
constexpr std::array<ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
// None.
};
@@ -324,15 +298,14 @@ namespace x32 {
// Dynamic FPRs
// - v0-v7
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
constexpr std::array<ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
// v0 ~ v1 are temps
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4,
FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
};
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
// This is /all/ of the SRA registers
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr std::array<ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {[]() -> uint32_t {
uint32_t Mask {};
@@ -343,15 +316,14 @@ namespace x32 {
}()};
// Dynamic FPRs when the host supports SVE-256bit.
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
constexpr std::array<ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
// v0 ~ v1 are used as temps.
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27, ARMEmitter::VReg::v28,
ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
} // namespace x32
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
@@ -379,24 +351,22 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr
}
#endif
CPU.SetUp();
// Number of register available is dependent on what operating mode the proccess is in.
if (EmitterCTX->Config.Is64BitMode()) {
StaticRegisters = x64::SRA;
GeneralRegisters = x64::RA;
GeneralPairRegisters = x64::RAPair;
StaticFPRegisters = x64::SRAFPR;
GeneralFPRegisters = x64::RAFPR;
PairRegisters = x64::RAPairs;
#ifdef _M_ARM_64EC
ConfiguredDynamicRegisterBase = std::span(x64::RA.begin(), 7);
#endif
} else {
ConfiguredDynamicRegisterBase = std::span(x32::RA.begin() + 6, 8);
PairRegisters = x32::RAPairs;
StaticRegisters = x32::SRA;
GeneralRegisters = x32::RA;
GeneralPairRegisters = x32::RAPair;
StaticFPRegisters = x32::SRAFPR;
GeneralFPRegisters = x32::RAFPR;
@@ -451,7 +421,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
if (RequiredMoveSegments > 1) {
// Only try to use this path if the number of segments is > 1.
// `movz` is better than `orr` since hardware will rename or merge if possible when `movz` is used.
const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Constant, RegSizeInBits(s));
const auto IsImm = ARMEmitter::Emitter::IsImmLogical(Constant, RegSizeInBits(s));
if (IsImm) {
orr(s, Reg, ARMEmitter::Reg::zr, Constant);
if (NOPPad) {
@@ -463,6 +433,17 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
}
}
// If we can't handle negatives with the orr, try with movn+movk
if (Is64Bit && ((~Constant) >> 32) == 0) {
movn(s, Reg, (~Constant) & 0xFFFF);
movk(s, Reg, (Constant >> 16) & 0xFFFF, 16);
if (NOPPad) {
nop();
nop();
}
return;
}
// ADRP+ADD is specifically optimized in hardware
// Check if we can use this
auto PC = GetCursorAddress<uint64_t>();
@@ -477,7 +458,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
// If the aligned offset is within the 4GB window then we can use ADRP+ADD
// and the number of move segments more than 1
if (RequiredMoveSegments > 1 && vixl::IsInt32(AlignedOffset)) {
if (RequiredMoveSegments > 1 && ARMEmitter::Emitter::IsInt32(AlignedOffset)) {
// If this is 4k page aligned then we only need ADRP
if ((AlignedOffset & 0xFFF) == 0) {
adrp(Reg, AlignedOffset >> 12);
@@ -485,7 +466,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
if (vixl::IsInt21(SmallOffset)) {
if (ARMEmitter::Emitter::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
} else {
// Need to use ADRP + ADD
@@ -589,12 +570,53 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
}
void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
// Enable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
orr(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
if (SetFIZ) {
// Insert MXCSR.DAZ in to FIZ
ldr(TmpReg2.W(), STATE.R(), offsetof(FEXCore::Core::CPUState, mxcsr));
bfxil(ARMEmitter::Size::i64Bit, TmpReg, TmpReg2, 6, 1);
}
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
if (SetPredRegs) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
if (EmitterCTX->HostFeatures.SupportsSVE256) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
}
if (EmitterCTX->HostFeatures.SupportsSVE128) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
//
// Disable FPCR.NEP and FPCR.AH
// Disable FPCR.NEP and FPCR.AH and FPCR.FIZ
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
// Also interacts with RPRES to change reciprocal/rsqrt precision from 8-bit mantissa to 12-bit.
@@ -604,7 +626,8 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
bic(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
(1U << 1) | // AH
(1U << 0)); // FIZ
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
@@ -643,7 +666,7 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
}
if (FPRs) {
if (EmitterCTX->HostFeatures.SupportsAVX) {
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
@@ -682,37 +705,37 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
FEXCore::ARMEmitter::Register TmpReg = FEXCore::ARMEmitter::Reg::r0;
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
[[maybe_unused]] bool FoundRegister {};
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & GPRFillMask)) {
TmpReg = Reg;
FoundRegister = true;
break;
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
*GPRFillMask &= ~(1U << Reg.Idx());
return std::make_optional(Reg);
}
}
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(FoundRegister, "Didn't have an SRA register to use as a temporary while spilling!");
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
// Enable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
orr(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
#ifdef _M_ARM_64EC
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
ldr(TmpReg.X(), ARMEmitter::Reg::r18, TEB_CPU_AREA_OFFSET);
ldr(STATE, TmpReg, CPU_AREA_EMULATOR_DATA_OFFSET);
#endif
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
@@ -723,18 +746,10 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
if (FPRs) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
if (EmitterCTX->HostFeatures.SupportsSVE) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
if (EmitterCTX->HostFeatures.SupportsAVX) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
@@ -798,8 +813,8 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
}
}
void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
void Arm64Emitter::PushVectorRegisters(ARMEmitter::Register TmpReg, bool SVE256Regs, std::span<const ARMEmitter::VRegister> VRegs) {
if (SVE256Regs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
@@ -835,7 +850,7 @@ void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, boo
}
}
void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs) {
void Arm64Emitter::PushGeneralRegisters(ARMEmitter::Register TmpReg, std::span<const ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
@@ -849,8 +864,8 @@ void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, st
}
}
void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
if (SVERegs) {
void Arm64Emitter::PopVectorRegisters(bool SVE256Regs, std::span<const ARMEmitter::VRegister> VRegs) {
if (SVE256Regs) {
size_t i = 0;
for (; i < (VRegs.size() % 4); i += 2) {
const auto Reg1 = VRegs[i];
@@ -885,7 +900,7 @@ void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARM
}
}
void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs) {
void Arm64Emitter::PopGeneralRegisters(std::span<const ARMEmitter::Register> Regs) {
size_t i = 0;
for (; i < (Regs.size() % 2); ++i) {
const auto Reg1 = Regs[i];
@@ -898,10 +913,10 @@ void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Regi
}
}
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
void Arm64Emitter::PushDynamicRegsAndLR(ARMEmitter::Register TmpReg) {
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
const auto GPRSize = (ConfiguredDynamicRegisterBase.size() + 1) * Core::CPUState::GPR_REG_SIZE;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
const auto FPRRegSize = CanUseSVE256 ? 32 : 16;
const auto FPRSize = GeneralFPRegisters.size() * FPRRegSize;
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
@@ -913,7 +928,7 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 2 == 0, "Needs to have multiple of 2 FPRs for RA");
// Push the vector registers
PushVectorRegisters(TmpReg, CanUseSVE, GeneralFPRegisters);
PushVectorRegisters(TmpReg, CanUseSVE256, GeneralFPRegisters);
// Push the general registers.
PushGeneralRegisters(TmpReg, ConfiguredDynamicRegisterBase);
@@ -924,10 +939,10 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
}
void Arm64Emitter::PopDynamicRegsAndLR() {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
// Pop vectors first
PopVectorRegisters(CanUseSVE, GeneralFPRegisters);
PopVectorRegisters(CanUseSVE256, GeneralFPRegisters);
// Pop GPRs second
PopGeneralRegisters(ConfiguredDynamicRegisterBase);
@@ -937,12 +952,12 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
#endif
}
void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
void Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, bool FPRs) {
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
const auto FPRRegSize = CanUseSVE256 ? 32 : 16;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs {};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs {};
std::span<const ARMEmitter::Register> DynamicGPRs {};
std::span<const ARMEmitter::VRegister> DynamicFPRs {};
uint32_t PreserveSRAMask {};
uint32_t PreserveSRAFPRMask {};
if (EmitterCTX->Config.Is64BitMode()) {
@@ -951,7 +966,7 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
if (CanUseSVE256) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
@@ -961,7 +976,7 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
if (CanUseSVE256) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
@@ -980,17 +995,17 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
// Push the vector registers.
PushVectorRegisters(TmpReg, CanUseSVE, DynamicFPRs);
PushVectorRegisters(TmpReg, CanUseSVE256, DynamicFPRs);
// Push the general registers.
PushGeneralRegisters(TmpReg, DynamicGPRs);
}
void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs {};
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs {};
std::span<const ARMEmitter::Register> DynamicGPRs {};
std::span<const ARMEmitter::VRegister> DynamicFPRs {};
uint32_t PreserveSRAMask {};
uint32_t PreserveSRAFPRMask {};
@@ -1000,7 +1015,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
PreserveSRAMask = x64::PreserveAll_SRAMask;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
if (CanUseSVE256) {
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
}
@@ -1010,7 +1025,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
PreserveSRAMask = x32::PreserveAll_SRAMask;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
if (CanUseSVE) {
if (CanUseSVE256) {
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
}
@@ -1020,7 +1035,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
FillStaticRegs(true, PreserveSRAMask, PreserveSRAFPRMask);
// Pop the vector registers.
PopVectorRegisters(CanUseSVE, DynamicFPRs);
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
// Pop the general registers.
PopGeneralRegisters(DynamicGPRs);
@@ -2,9 +2,6 @@
#pragma once
#include "FEXCore/Utils/EnumUtils.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
#include "Interface/Core/ObjectCache/Relocations.h"
#include <aarch64/assembler-aarch64.h>
@@ -22,6 +19,8 @@
#include <FEXCore/Config/Config.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#include <array>
#include <cstddef>
@@ -35,86 +34,85 @@ class ContextImpl;
namespace FEXCore::CPU {
// Contains the address to the currently available CPU state
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
constexpr auto STATE = ARMEmitter::XReg::x28;
#ifndef _M_ARM_64EC
// GPR temporaries. Only x3 can be used across spill boundaries
// so if these ever need to change, be very careful about that.
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x0;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x1;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x2;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
constexpr auto TMP1 = ARMEmitter::XReg::x0;
constexpr auto TMP2 = ARMEmitter::XReg::x1;
constexpr auto TMP3 = ARMEmitter::XReg::x2;
constexpr auto TMP4 = ARMEmitter::XReg::x3;
constexpr bool TMP_ABIARGS = true;
// We pin r26/r27 as PF/AF respectively, this is internal FEX ABI.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r26;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r27;
constexpr auto REG_PF = ARMEmitter::Reg::r26;
constexpr auto REG_AF = ARMEmitter::Reg::r27;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
#else
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x10;
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x11;
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x12;
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x13;
constexpr auto TMP1 = ARMEmitter::XReg::x10;
constexpr auto TMP2 = ARMEmitter::XReg::x11;
constexpr auto TMP3 = ARMEmitter::XReg::x12;
constexpr auto TMP4 = ARMEmitter::XReg::x13;
constexpr bool TMP_ABIARGS = false;
// We pin r11/r12 as PF/AF respectively for arm64ec, as r26/r27 are used for SRA.
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r9;
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r24;
constexpr auto REG_PF = ARMEmitter::Reg::r9;
constexpr auto REG_AF = ARMEmitter::Reg::r24;
// Vector temporaries
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v16;
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v17;
constexpr auto VTMP1 = ARMEmitter::VReg::v16;
constexpr auto VTMP2 = ARMEmitter::VReg::v17;
// Entry/Exit ABI
constexpr auto EC_CALL_CHECKER_PC_REG = ARMEmitter::XReg::x9;
constexpr auto EC_ENTRY_CPUAREA_REG = ARMEmitter::XReg::x17;
// These structures are not included in the standard Windows headers, define the offsets of members we care about for EC here.
constexpr size_t TEB_CPU_AREA_OFFSET = 0x1788;
constexpr size_t TEB_PEB_OFFSET = 0x60;
constexpr size_t PEB_EC_CODE_BITMAP_OFFSET = 0x368;
constexpr size_t CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET = 0x1;
constexpr size_t CPU_AREA_EMULATOR_STACK_BASE_OFFSET = 0x8;
constexpr size_t CPU_AREA_EMULATOR_DATA_OFFSET = 0x30;
#endif
// Predicate register temporaries (used when AVX support is enabled)
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_16B = FEXCore::ARMEmitter::PReg::p6;
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PReg::p7;
constexpr ARMEmitter::PRegister PRED_TMP_16B = ARMEmitter::PReg::p6;
constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
// This class contains common emitter utility functions that can
// be used by both Arm64 JIT and ARM64 Dispatcher
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
class Arm64Emitter : public ARMEmitter::Emitter {
protected:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
FEXCore::Context::ContextImpl* EmitterCTX;
vixl::aarch64::CPU CPU;
std::span<const FEXCore::ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
std::span<const FEXCore::ARMEmitter::Register> StaticRegisters {};
std::span<const FEXCore::ARMEmitter::Register> GeneralRegisters {};
std::span<const std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>> GeneralPairRegisters {};
std::span<const FEXCore::ARMEmitter::VRegister> StaticFPRegisters {};
std::span<const FEXCore::ARMEmitter::VRegister> GeneralFPRegisters {};
std::span<const ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
std::span<const ARMEmitter::Register> StaticRegisters {};
std::span<const ARMEmitter::Register> GeneralRegisters {};
std::span<const ARMEmitter::VRegister> StaticFPRegisters {};
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
uint32_t PairRegisters = 0;
/**
* @name Register Allocation
* @{ */
constexpr static uint32_t RegisterClasses = 6;
constexpr static uint64_t GPRBase = (0ULL << 32);
constexpr static uint64_t FPRBase = (1ULL << 32);
constexpr static uint64_t GPRPairBase = (2ULL << 32);
/** @} */
constexpr static uint8_t RA_32 = 0;
constexpr static uint8_t RA_64 = 1;
constexpr static uint8_t RA_FPR = 2;
void LoadConstant(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt);
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -124,13 +122,13 @@ protected:
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
// Generic push and pop vector registers.
void PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs);
void PushVectorRegisters(ARMEmitter::Register TmpReg, bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
void PushGeneralRegisters(ARMEmitter::Register TmpReg, std::span<const ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs);
void PopVectorRegisters(bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
void PopGeneralRegisters(std::span<const ARMEmitter::Register> Regs);
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
void PushDynamicRegsAndLR(ARMEmitter::Register TmpReg);
void PopDynamicRegsAndLR();
void PushCalleeSavedRegisters();
@@ -146,10 +144,10 @@ protected:
// Callee Saved:
// - X9-X15, X19-X31
// - Low 128-bits of v8-v31
void SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true);
void SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, bool FPRs = true);
void FillForPreserveAllABICall(bool FPRs = true);
void SpillForABICall(bool SupportsPreserveAllABI, FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true) {
void SpillForABICall(bool SupportsPreserveAllABI, ARMEmitter::Register TmpReg, bool FPRs = true) {
if (SupportsPreserveAllABI) {
SpillForPreserveAllABICall(TmpReg, FPRs);
} else {
File diff suppressed because it is too large. Load diff
@@ -19,6 +19,10 @@ namespace CPU {
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT_UPPER
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT_UPPER
{0x8000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPS_INVERT
{0x8000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPS_INVERT_UPPER
{0x0000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPD_INVERT
{0x0000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPD_INVERT_UPPER
{0x0000'0001'0000'0000ULL, 0x0000'0003'0000'0002ULL}, // NAMED_VECTOR_MOVMSKPS_SHIFT
{0x040B'0E01'0B0E'0104ULL, 0x0C03'0609'0306'090CULL}, // NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE
{0x0706'0504'FFFF'FFFFULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0110B
+1 -7
View File
@@ -34,12 +34,6 @@ namespace CodeSerialize {
}
namespace CPU {
struct CPUBackendFeatures {
bool SupportsFlags = false;
bool SupportsSaturatingRoundingShifts = false;
bool SupportsVTBL2 = false;
};
class CPUBackend {
public:
struct CodeBuffer {
@@ -140,7 +134,7 @@ namespace CPU {
*/
[[nodiscard]]
virtual CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
FEXCore::IR::RegisterAllocationData* RAData) = 0;
const FEXCore::IR::RegisterAllocationData* RAData) = 0;
/**
* @brief Relocates a block of code from the JIT code object cache
+62 -48
View File
@@ -34,6 +34,8 @@ namespace ProductNames {
static const char ARM_A76AE[] = "Cortex-A76AE";
static const char ARM_V1[] = "Neoverse V1";
static const char ARM_V2[] = "Neoverse V2";
static const char ARM_V3[] = "Neoverse V3";
static const char ARM_V3AE[] = "Neoverse V3AE";
static const char ARM_A77[] = "Cortex-A77";
static const char ARM_A78[] = "Cortex-A78";
static const char ARM_A78AE[] = "Cortex-A78AE";
@@ -41,13 +43,16 @@ namespace ProductNames {
static const char ARM_A710[] = "Cortex-A710";
static const char ARM_A715[] = "Cortex-A715";
static const char ARM_A720[] = "Cortex-A720";
static const char ARM_A725[] = "Cortex-A725";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X1C[] = "Cortex-X1C";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_X3[] = "Cortex-X3";
static const char ARM_X4[] = "Cortex-X4";
static const char ARM_X925[] = "Cortex-X925";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
static const char ARM_N3[] = "Neoverse N3";
static const char ARM_E1[] = "Neoverse E1";
static const char ARM_A35[] = "Cortex-A35";
static const char ARM_A53[] = "Cortex-A53";
@@ -69,6 +74,8 @@ namespace ProductNames {
static const char ARM_Firestorm[] = "Apple Firestorm";
static const char ARM_Icestorm[] = "Apple Icestorm";
static const char ARM_ORYON_1[] = "Oryon-1";
#else
#endif
} // namespace ProductNames
@@ -140,10 +147,17 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 42> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 48> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
{0x41, 0xd85, 1, ProductNames::ARM_X925}, // X925
{0x41, 0xd87, 1, ProductNames::ARM_A725}, // A725
{0x41, 0xd84, 1, ProductNames::ARM_V3}, // V3
{0x41, 0xd83, 1, ProductNames::ARM_V3AE}, // V3AE
{0x41, 0xd8e, 1, ProductNames::ARM_N3}, // N3
{0x41, 0xd82, 1, ProductNames::ARM_X4}, // X4
{0x41, 0xd81, 1, ProductNames::ARM_A720}, // A720
{0x41, 0xd4e, 1, ProductNames::ARM_X3}, // X3
@@ -343,8 +357,7 @@ void CPUIDEmu::SetupHostHybridFlag() {}
void CPUIDEmu::SetupFeatures() {
// TODO: Enable once AVX is supported.
if (false && CTX->HostFeatures.SupportsAVX) {
if (CTX->HostFeatures.SupportsAVX) {
XCR0 |= XCR0_AVX;
}
@@ -355,14 +368,14 @@ void CPUIDEmu::SetupFeatures() {
return;
}
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::DISABLE##enum_name) != 0; \
const bool Enable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::ENABLE##enum_name) != 0; \
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::DISABLE##enum_name) != 0; \
const bool Enable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features.FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features.FeatureName = Result; \
const bool AlreadyEnabled = Features.FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features.FeatureName = Result; \
} while (0)
ENABLE_DISABLE_OPTION(SHA, SHA, SHA);
@@ -413,7 +426,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(1 << 9) | // SSSE3
(0 << 10) | // L1 context ID
(0 << 11) | // Silicon debug
(0 << 12) | // FMA3
(SupportsAVX() << 12) | // FMA3
(1 << 13) | // CMPXCHG16B
(0 << 14) | // xTPR update control
(0 << 15) | // Perfmon and debug capability
@@ -430,7 +443,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
(SupportsAVX() << 26) | // XSAVE
(SupportsAVX() << 27) | // OSXSAVE
(SupportsAVX() << 28) | // AVX
(0 << 29) | // F16C
(SupportsAVX() << 29) | // F16C
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
(Hypervisor << 31);
@@ -597,6 +610,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
// This is due to LRCPC performance on Cortex being abysmal.
// Only enable EnhancedREPMOVS if SoftwareTSO isn't required OR if MemcpySetTSO is not enabled.
const uint32_t SupportsEnhancedREPMOVS = CTX->SoftwareTSORequired() == false || MemcpySetTSOEnabled() == false;
const uint32_t SupportsVPCLMULQDQ = CTX->HostFeatures.SupportsPMULL_128Bit && SupportsAVX();
// Number of subfunctions
Res.eax = 0x0;
@@ -605,7 +619,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 2) | // SGX
(SupportsAVX() << 3) | // BMI1
(0 << 4) | // Intel Hardware Lock Elison
(0 << 5) | // AVX2 support
(SupportsAVX() << 5) | // AVX2 support
(1 << 6) | // FPU data pointer updated only on exception
(1 << 7) | // SMEP support
(SupportsAVX() << 8) | // BMI2
@@ -624,7 +638,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(1 << 23) | // CLFLUSHOPT instruction
(CTX->HostFeatures.SupportsCLWB << 24) | // CLWB instruction
(1 << 24) | // CLWB instruction
(0 << 25) | // Intel processor trace
(0 << 26) | // Reserved
(0 << 27) | // Reserved
@@ -633,38 +647,38 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 30) | // Reserved
(0 << 31); // Reserved
Res.ecx = (1 << 0) | // PREFETCHWT1
(0 << 1) | // AVX512VBMI
(0 << 2) | // Usermode instruction prevention
(0 << 3) | // Protection keys for user mode pages
(0 << 4) | // OS protection keys
(0 << 5) | // waitpkg
(0 << 6) | // AVX512_VBMI2
(0 << 7) | // CET shadow stack
(0 << 8) | // GFNI
(0 << 9) | // VAES
(0 << 10) | // VPCLMULQDQ
(0 << 11) | // AVX512_VNNI
(0 << 12) | // AVX512_BITALG
(0 << 13) | // Intel Total Memory Encryption
(0 << 14) | // AVX512_VPOPCNTDQ
(0 << 15) | // Reserved
(0 << 16) | // 5 Level page tables
(0 << 17) | // MPX MAWAU
(0 << 18) | // MPX MAWAU
(0 << 19) | // MPX MAWAU
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(0 << 23) | // Reserved
(0 << 24) | // Reserved
(0 << 25) | // CLDEMOTE
(0 << 26) | // Reserved
(0 << 27) | // MOVDIRI
(0 << 28) | // MOVDIR64B
(0 << 29) | // Reserved
(0 << 30) | // SGX Launch configuration
(0 << 31); // Reserved
Res.ecx = (1 << 0) | // PREFETCHWT1
(0 << 1) | // AVX512VBMI
(0 << 2) | // Usermode instruction prevention
(0 << 3) | // Protection keys for user mode pages
(0 << 4) | // OS protection keys
(0 << 5) | // waitpkg
(0 << 6) | // AVX512_VBMI2
(0 << 7) | // CET shadow stack
(0 << 8) | // GFNI
(CTX->HostFeatures.SupportsAES256 << 9) | // VAES
(SupportsVPCLMULQDQ << 10) | // VPCLMULQDQ
(0 << 11) | // AVX512_VNNI
(0 << 12) | // AVX512_BITALG
(0 << 13) | // Intel Total Memory Encryption
(0 << 14) | // AVX512_VPOPCNTDQ
(0 << 15) | // Reserved
(0 << 16) | // 5 Level page tables
(0 << 17) | // MPX MAWAU
(0 << 18) | // MPX MAWAU
(0 << 19) | // MPX MAWAU
(0 << 20) | // MPX MAWAU
(0 << 21) | // MPX MAWAU
(1 << 22) | // RDPID Read Processor ID
(0 << 23) | // Reserved
(0 << 24) | // Reserved
(0 << 25) | // CLDEMOTE
(0 << 26) | // Reserved
(0 << 27) | // MOVDIRI
(0 << 28) | // MOVDIR64B
(0 << 29) | // Reserved
(0 << 30) | // SGX Launch configuration
(0 << 31); // Reserved
Res.edx = (0 << 0) | // Reserved
(0 << 1) | // Reserved
@@ -883,7 +897,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
(0 << 18) | // Reserved
(0 << 19) | // Reserved
(0 << 20) | // Reserved
(0 << 21) | // Reserved
(0 << 21) | // XOP-TBM
(0 << 22) | // Topology extensions support
(0 << 23) | // Core performance counter extensions
(0 << 24) | // NB performance counter extensions
@@ -891,7 +905,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
(0 << 26) | // Data breakpoints extensions
(0 << 27) | // Performance TSC
(0 << 28) | // L2 perf counter extensions
(0 << 29) | // Reserved
(0 << 29) | // MONITORX
(0 << 30) | // Reserved
(0 << 31); // Reserved
+143 -132
View File
@@ -38,7 +38,6 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/File.h>
@@ -75,8 +74,9 @@ $end_info$
#include <xxhash.h>
namespace FEXCore::Context {
ContextImpl::ContextImpl()
: CPUID {this}
ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
: HostFeatures {Features}
, CPUID {this}
, IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
@@ -216,6 +216,55 @@ uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadSt
return EFLAGS;
}
void ContextImpl::ReconstructXMMRegisters(const FEXCore::Core::InternalThreadState* Thread, __uint128_t* XMM_Low, __uint128_t* YMM_High) {
const size_t MaximumRegisters = Config.Is64BitMode ? FEXCore::Core::CPUState::NUM_XMMS : 8;
if (YMM_High != nullptr && HostFeatures.SupportsAVX) {
const bool SupportsConvergedRegisters = HostFeatures.SupportsSVE256;
if (SupportsConvergedRegisters) {
///< Output wants to de-interleave
for (size_t i = 0; i < MaximumRegisters; ++i) {
memcpy(&XMM_Low[i], &Thread->CurrentFrame->State.xmm.avx.data[i][0], sizeof(__uint128_t));
memcpy(&YMM_High[i], &Thread->CurrentFrame->State.xmm.avx.data[i][2], sizeof(__uint128_t));
}
} else {
///< Matches what FEX wants with non-converged registers
for (size_t i = 0; i < MaximumRegisters; ++i) {
memcpy(&XMM_Low[i], &Thread->CurrentFrame->State.xmm.sse.data[i][0], sizeof(__uint128_t));
memcpy(&YMM_High[i], &Thread->CurrentFrame->State.avx_high[i][0], sizeof(__uint128_t));
}
}
} else {
// Only support SSE, no AVX here, even if requested.
memcpy(XMM_Low, Thread->CurrentFrame->State.xmm.sse.data, MaximumRegisters * sizeof(__uint128_t));
}
}
void ContextImpl::SetXMMRegistersFromState(FEXCore::Core::InternalThreadState* Thread, const __uint128_t* XMM_Low, const __uint128_t* YMM_High) {
const size_t MaximumRegisters = Config.Is64BitMode ? FEXCore::Core::CPUState::NUM_XMMS : 8;
if (YMM_High != nullptr && HostFeatures.SupportsAVX) {
const bool SupportsConvergedRegisters = HostFeatures.SupportsSVE256;
if (SupportsConvergedRegisters) {
///< Output wants to de-interleave
for (size_t i = 0; i < MaximumRegisters; ++i) {
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][0], &XMM_Low[i], sizeof(__uint128_t));
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][2], &YMM_High[i], sizeof(__uint128_t));
}
} else {
///< Matches what FEX wants with non-converged registers
for (size_t i = 0; i < MaximumRegisters; ++i) {
memcpy(&Thread->CurrentFrame->State.xmm.sse.data[i][0], &XMM_Low[i], sizeof(__uint128_t));
memcpy(&Thread->CurrentFrame->State.avx_high[i][0], &YMM_High[i], sizeof(__uint128_t));
}
}
} else {
// Only support SSE, no AVX here, even if requested.
memcpy(Thread->CurrentFrame->State.xmm.sse.data, XMM_Low, MaximumRegisters * sizeof(__uint128_t));
}
}
void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) {
const auto Frame = Thread->CurrentFrame;
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_EFLAG_BITS; ++i) {
@@ -260,14 +309,6 @@ void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState
bool ContextImpl::InitCore() {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
case FEXCore::Config::CONFIG_IRJIT: BackendFeatures = FEXCore::CPU::GetArm64JITBackendFeatures(); break;
case FEXCore::Config::CONFIG_CUSTOM:
// Do nothing
break;
default: LogMan::Msg::EFmt("Unknown core configuration"); return false;
}
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
// Set up the SignalDelegator config since core is initialized.
@@ -324,14 +365,15 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
FEXCore::Context::ExitReason ContextImpl::RunUntilExit(FEXCore::Core::InternalThreadState* Thread) {
ExecutionThread(Thread);
while (true) {
auto reason = Thread->ExitReason;
// Don't return if a custom exit handling the exit
if (!CustomExitHandler || reason == ExitReason::EXIT_SHUTDOWN) {
return reason;
}
CoreShuttingDown.store(true);
if (CustomExitHandler) {
CustomExitHandler(Thread, FEXCore::Context::ExitReason::EXIT_SHUTDOWN);
return Thread->ExitReason;
}
return FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
}
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
@@ -341,9 +383,6 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState* Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
if (ThunkHandler) {
ThunkHandler->RegisterTLSState(Thread);
}
@@ -366,7 +405,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
Thread->CTX = this;
Thread->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT);
Thread->PassManager->AddDefaultPasses(this);
Thread->PassManager->AddDefaultValidationPasses();
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
@@ -374,7 +413,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
// Create CPU backend
switch (Config.Core) {
case FEXCore::Config::CONFIG_IRJIT:
Thread->PassManager->InsertRegisterAllocationPass(HostFeatures.SupportsAVX);
Thread->PassManager->InsertRegisterAllocationPass();
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
break;
case FEXCore::Config::CONFIG_CUSTOM: Thread->CPUBackend = CustomCPUFactory(this, Thread); break;
@@ -397,14 +436,11 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, FEXCore::C
}
// Set up the thread manager state
Thread->ThreadManager.parent_tid = ParentTID;
Thread->CurrentFrame->Thread = Thread;
InitializeCompiler(Thread);
Thread->CurrentFrame->State.DeferredSignalRefCount.Store(0);
Thread->CurrentFrame->State.DeferredSignalFaultAddress =
reinterpret_cast<Core::NonAtomicRefCounter<uint64_t>*>(FEXCore::Allocator::VirtualAlloc(4096));
if (Config.BlockJITNaming() || Config.GlobalJITNaming() || Config.LibraryJITNaming()) {
// Allocate a JIT symbol buffer only if enabled.
@@ -421,7 +457,8 @@ void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool
#endif
}
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096);
FEXCore::Allocator::VirtualProtect(&Thread->InterruptFaultPage, sizeof(Thread->InterruptFaultPage),
Allocator::ProtectOptions::Read | Allocator::ProtectOptions::Write);
delete Thread;
}
@@ -469,29 +506,39 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
};
static void ValidateIR(ContextImpl* ctx, IR::IREmitter* IREmitter) {
// Convert to text, Parse, Convert to text again and make sure the texts match
fextl::stringstream out;
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
compaction->Run(IREmitter);
auto NewIR = IREmitter->ViewIR();
Dump(&out, &NewIR, nullptr);
out.seekg(0);
FEXCore::Utils::PooledAllocatorMalloc Allocator;
auto reparsed = IR::Parse(Allocator, out);
if (reparsed == nullptr) {
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
} else {
fextl::stringstream out2;
auto NewIR2 = reparsed->ViewIR();
Dump(&out2, &NewIR2, nullptr);
if (out.str() != out2.str()) {
LogMan::Msg::IFmt("one:\n {}", out.str());
LogMan::Msg::IFmt("two:\n {}", out2.str());
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
}
// IRStorageBase with fully owned memory
struct IRListCopy : public IR::IRStorageBase {
std::span<std::byte> IRData;
std::span<std::byte> ListData;
// TODO: Consider defaulting to empty RAData instead?
IR::RegisterAllocationData::UniquePtr RADataInternal;
IRListCopy(const IR::IRListView& view, IR::RegisterAllocationData::UniquePtr RAData)
: RADataInternal(std::move(RAData)) {
std::byte* Storage = reinterpret_cast<std::byte*>(FEXCore::Allocator::malloc(view.GetDataSize() + view.GetListSize()));
IRData = {Storage, Storage + view.GetDataSize()};
ListData = {Storage + view.GetDataSize(), Storage + view.GetDataSize() + view.GetListSize()};
memcpy(IRData.data(), (char*)view.GetData(), IRData.size());
memcpy(ListData.data(), (char*)view.GetListData(), ListData.size());
}
}
IRListCopy(const IRListCopy& other) = delete;
IRListCopy(IRListCopy&& other) = delete;
~IRListCopy() {
FEXCore::Allocator::free(IRData.data());
}
const IR::RegisterAllocationData* RAData() override {
return RADataInternal.get();
}
IR::IRListView GetIRView() override {
return IR::IRListView {IRData.data(), ListData.data(), IRData.size(), ListData.size()};
}
};
ContextImpl::GenerateIRResult
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
@@ -556,6 +603,20 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
DecodedInfo = &Block.DecodedInstructions[i];
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
// Do a partial register cache flush before every instruction. This
// prevents cross-instruction static register caching, while allowing
// context load/stores to be optimized within a block. Theoretically,
// this flush is not required for correctness, all mandatory flushes are
// included in instruction-specific handlers. Instead, this is a blunt
// heuristic to make the register cache less aggressive, as the current
// RA generates bad code in common cases with tied registers otherwise.
//
// However, it makes our exception handling behaviour more predictable.
// It is potentially correctness bearing in that sense, but that is a
// side effect here and (if that behaviour is required) we should handle
// that more explicitly later.
Thread->OpDispatcher->FlushRegisterCache(true);
if (ExtendedDebugInfo || Thread->OpDispatcher->CanHaveSideEffects(TableInfo, DecodedInfo)) {
Thread->OpDispatcher->_GuestOpcode(Block.Entry + BlockInstructionsLength - GuestRIP);
}
@@ -574,7 +635,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
Thread->OpDispatcher->_ExitFunction(
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
@@ -605,7 +666,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
// Invalid instruction
Thread->OpDispatcher->InvalidOp(DecodedInfo);
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
}
const bool NeedsBlockEnd =
@@ -615,14 +676,14 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
if (HadDispatchError && TotalInstructions == 0) {
// Couldn't handle any instruction in op dispatcher
Thread->OpDispatcher->ResetWorkingList();
return {nullptr, nullptr, 0, 0, 0, 0};
return {nullptr, 0, 0, 0, 0};
}
if (NeedsBlockEnd) {
const uint8_t GPRSize = GetGPRSize();
// We had some instructions. Early exit
Thread->OpDispatcher->_ExitFunction(
Thread->OpDispatcher->ExitFunction(
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
break;
}
@@ -643,35 +704,26 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
auto ShouldDump = Thread->OpDispatcher->ShouldDumpIR();
// Debug
{
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
}
if (static_cast<ContextImpl*>(Thread->CTX)->Config.ValidateIRarser) {
ValidateIR(this, IREmitter);
}
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
}
// Run the passmanager over the IR from the dispatcher
Thread->PassManager->Run(IREmitter);
// Debug
{
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP,
Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
if (ShouldDump) {
IRDumper(Thread, IREmitter, GuestRIP,
Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
}
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
auto IRList = IREmitter->CreateIRCopy();
auto IRList = fextl::make_unique<IRListCopy>(IREmitter->ViewIR(), std::move(RAData));
IREmitter->DelayedDisownBuffer();
return {
.IRList = IRList,
.RAData = std::move(RAData),
.IR = std::move(IRList),
.TotalInstructions = TotalInstructions,
.TotalInstructionsLength = TotalInstructionsLength,
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
@@ -680,13 +732,6 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
}
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
FEXCore::IR::IRListView* IRList {};
FEXCore::Core::DebugData* DebugData {};
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
bool GeneratedIR {};
uint64_t StartAddr {};
uint64_t Length {};
// JIT Code object cache lookup
if (CodeObjectCacheService) {
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
@@ -695,9 +740,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
if (CompiledCode) {
return {
.CompiledCode = CompiledCode,
.IRData = nullptr, // No IR data generated
.IR = nullptr, // No IR/RA data generated
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
.RAData = nullptr, // No RA data generated
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
.StartAddr = 0, // Unused
.Length = 0, // Unused
@@ -713,49 +757,47 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
}
}
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
FEXCore::Core::DebugData* DebugData {};
uint64_t StartAddr {};
uint64_t Length {};
// AOT IR bookkeeping and cache
{
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP, IRList);
if (_GeneratedIR) {
auto IRFromAOT = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP);
if (IRFromAOT) {
// Setup pointers to internal structures
IRList = IRCopy;
RAData = std::move(RACopy);
DebugData = DebugDataCopy;
StartAddr = _StartAddr;
Length = _Length;
GeneratedIR = _GeneratedIR;
IR = std::move(IRFromAOT->IR);
DebugData = IRFromAOT->DebugData;
StartAddr = IRFromAOT->StartAddr;
Length = IRFromAOT->Length;
}
}
if (IRList == nullptr) {
if (!IR) {
// Generate IR + Meta Info
auto [IRCopy, RACopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] =
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
auto [IRCopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] = GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
// Setup pointers to internal structures
IRList = IRCopy;
RAData = std::move(RACopy);
IR = std::move(IRCopy);
DebugData = new FEXCore::Core::DebugData();
StartAddr = _StartAddr;
Length = _Length;
// These blocks aren't already in the cache
GeneratedIR = true;
}
if (IRList == nullptr) {
if (!IR) {
return {};
}
// Attempt to get the CPU backend to compile this code
auto IRView = IR->GetIRView();
return {
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
// In the future with code caching getting wired up, we will pass the rest of the data forward.
// TODO: Pass the data forward when code caching is wired up to this.
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get()).BlockEntry,
.IRData = IRList,
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, &IRView, DebugData, IR->RAData()).BlockEntry,
.IR = std::move(IR),
.DebugData = DebugData,
.RAData = std::move(RAData),
.GeneratedIR = GeneratedIR,
.GeneratedIR = true,
.StartAddr = StartAddr,
.Length = Length,
};
@@ -765,14 +807,6 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
#ifdef _M_ARM_64EC
// If the target PC is EC code, mark it in the L2 and return straight to the dispatcher
// so it can handle the call/return.
if (Thread->LookupCache->CheckPageEC(GuestRIP)) {
return GuestRIP;
}
#endif
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
@@ -782,21 +816,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
return HostCode;
}
void* CodePtr {};
FEXCore::IR::IRListView* IRList {};
FEXCore::Core::DebugData* DebugData {};
bool GeneratedIR {};
uint64_t StartAddr {}, Length {};
auto [Code, IR, Data, RAData, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP, MaxInst);
CodePtr = Code;
IRList = IR;
DebugData = Data;
GeneratedIR = Generated;
StartAddr = _StartAddr;
Length = _Length;
auto [CodePtr, IR, DebugData, GeneratedIR, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
if (CodePtr == nullptr) {
return 0;
}
@@ -847,7 +867,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
// Clear any relocations that might have been generated
Thread->CPUBackend->ClearRelocations();
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, std::move(RAData), IRList, DebugData, GeneratedIR)) {
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, std::move(IR), DebugData, GeneratedIR)) {
// Early exit
return (uintptr_t)CodePtr;
}
@@ -893,15 +913,6 @@ void ContextImpl::ExecutionThread(FEXCore::Core::InternalThreadState* Thread) {
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
if (Thread->ThreadManager.parent_tid == 0) {
CoreShuttingDown.store(true);
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
if (CustomExitHandler) {
CustomExitHandler(Thread->ThreadManager.TID, Thread->ExitReason);
}
}
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
@@ -996,7 +1007,7 @@ void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry* Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
}
void ContextImpl::AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) {
void ContextImpl::AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) {
if (ThunkHandler) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: MIT
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/X86HelperGen.h"
@@ -17,6 +16,8 @@
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/MathUtils.h>
#include <CodeEmitter/Emitter.h>
#include <atomic>
#include <condition_variable>
#include <csignal>
@@ -61,9 +62,6 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::SingleUseForwardLabel l_Sleep;
#ifdef _M_ARM_64EC
ARMEmitter::SingleUseForwardLabel ExitEC;
#endif
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
// Push all the register we need to save
@@ -82,20 +80,43 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
FillStaticRegs();
ARMEmitter::BiDirectionalLabel LoopTop {};
#ifdef _M_ARM_64EC
b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
FillStaticRegs();
// Enter JIT
b(&LoopTop);
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
str(EC_CALL_CHECKER_PC_REG, STATE_PTR(CpuStateFrame, State.rip));
// Swap stacks to the emulator stack
ldr(TMP1, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_STACK_BASE_OFFSET);
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
FillSpecialRegs(TMP1, TMP2, false, true);
// Enter JIT
#endif
// We want to ensure that we are 16 byte aligned at the top of this loop
Align16B();
ARMEmitter::BiDirectionalLabel FullLookup {};
ARMEmitter::BiDirectionalLabel CallBlock {};
ARMEmitter::BackwardLabel LoopTop {};
Bind(&LoopTop);
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
// Load in our RIP
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
// IMPORTANT: Pointers.Common.ExitFunctionEC callsites/implementations need to be
// adjusted accordingly if this changes.
auto RipReg = TMP3;
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
@@ -137,10 +158,6 @@ void Dispatcher::EmitDispatcher() {
// If page pointer is zero then we have no block
cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
#ifdef _M_ARM_64EC
// The LSB of an L2 page entry indicates if this page contains EC code
tbnz(TMP1, 0, &ExitEC);
#endif
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
@@ -167,22 +184,13 @@ void Dispatcher::EmitDispatcher() {
and_(ARMEmitter::Size::i64Bit, TMP2, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP3, TMP1);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, RipReg, TMP1);
// Jump to the block
br(TMP4);
}
}
#ifdef _M_ARM_64EC
{
Bind(&ExitEC);
// Target PC is already loaded into TMP3 at the start of the dispatcher
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
}
#endif
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
@@ -200,9 +208,17 @@ void Dispatcher::EmitDispatcher() {
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
mov(ARMEmitter::XReg::x0, STATE);
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
@@ -220,13 +236,20 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
#ifndef _WIN32
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP2, 0);
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
br(TMP1);
}
@@ -235,15 +258,44 @@ void Dispatcher::EmitDispatcher() {
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::SingleUseForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
#endif
SpillStaticRegs(TMP1);
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, TMP3);
mov(ARMEmitter::XReg::x2, RipReg);
}
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
ldr(ARMEmitter::XReg::x0, &l_CTX);
mov(ARMEmitter::XReg::x1, STATE);
@@ -259,13 +311,20 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP1, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP1, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
#ifndef _WIN32
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
// Trigger segfault if any deferred signals are pending
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
str(ARMEmitter::XReg::zr, TMP1, 0);
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
b(&LoopTop);
}
@@ -494,6 +553,8 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
Common.DispatcherLoopTopEnterEC = AbsoluteLoopTopAddressEnterEC;
Common.DispatcherLoopTopEnterECFillSRA = AbsoluteLoopTopAddressEnterECFillSRA;
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
@@ -47,6 +47,8 @@ public:
uint64_t ThreadStopHandlerAddressSpillSRA {};
uint64_t AbsoluteLoopTopAddress {};
uint64_t AbsoluteLoopTopAddressFillSRA {};
uint64_t AbsoluteLoopTopAddressEnterEC {};
uint64_t AbsoluteLoopTopAddressEnterECFillSRA {};
uint64_t ThreadPauseHandlerAddress {};
uint64_t ThreadPauseHandlerAddressSpillSRA {};
uint64_t ExitFunctionLinkerAddress {};
+9 -9
View File
@@ -221,14 +221,16 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
{
// If we have a VSIB byte (as opposed to SIB), then the index register is a vector.
const bool IsIndexVector = (DecodeInst->TableInfo->Flags & InstFlags::FLAGS_VEX_VSIB) != 0;
uint8_t InvalidSIBIndex = 0b100; ///< SIB Index where there is no register encoding.
if (IsIndexVector) {
DecodeInst->Flags |= X86Tables::DecodeFlags::FLAG_VSIB_BYTE;
InvalidSIBIndex = ~0; ///< No Invalid SIB Index with Index Vectors.
}
const uint8_t IndexREX = (DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X) != 0 ? 1 : 0;
const uint8_t BaseREX = (DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B) != 0 ? 1 : 0;
Operand->Data.SIB.Index = MapModRMToReg(IndexREX, SIB.index, false, false, IsIndexVector, false, 0b100);
Operand->Data.SIB.Index = MapModRMToReg(IndexREX, SIB.index, false, false, IsIndexVector, false, InvalidSIBIndex);
Operand->Data.SIB.Base = MapModRMToReg(BaseREX, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
}
@@ -630,7 +632,6 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
uint16_t X87Op = ((Op - 0xD8) << 8) | ModRMByte;
return NormalOp(&X87Ops[X87Op], X87Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
uint16_t map_select = 1;
uint16_t pp = 0;
const uint8_t Byte1 = ReadByte();
@@ -659,6 +660,9 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
}
if (options.w) {
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
}
if (!(map_select >= 1 && map_select <= 3)) {
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
return false;
@@ -673,7 +677,6 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
FEXCore::X86Tables::X86InstInfo* LocalInfo = &VEXTableOps[Op];
if (LocalInfo->Type >= FEXCore::X86Tables::TYPE_VEX_GROUP_12 && LocalInfo->Type <= FEXCore::X86Tables::TYPE_VEX_GROUP_17) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
@@ -941,14 +944,12 @@ void Decoder::BranchTargetInMultiblockRange() {
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
// Target offset is PC + InstSize + Literal
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
break;
}
case 0xE9:
case 0xEB: // Both are unconditional JMP instructions
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
Conditional = false;
break;
case 0xE8: // Call - Immediate target, We don't want to inline calls
@@ -1000,8 +1001,7 @@ bool Decoder::BranchTargetCanContinue(bool FinalInstruction) const {
if (DecodeInst->OP == 0xE8) { // Call - immediate target
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
if (GPRSize == 4) {
// If we are running a 32bit guest then wrap around addresses that go above 32bit
-1
View File
@@ -120,7 +120,6 @@ private:
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
FEXCORE_TELEMETRY_INIT(VEXOpTelem, TYPE_USES_VEX_OPS);
FEXCORE_TELEMETRY_INIT(EVEXOpTelem, TYPE_USES_EVEX_OPS);
};
} // namespace FEXCore::Frontend
@@ -1,252 +0,0 @@
// SPDX-License-Identifier: MIT
#include "Interface/Core/CPUID.h"
#include <FEXCore/Core/HostFeatures.h>
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
#ifdef _M_X86_64
#define XBYAK64
#define XBYAK_NO_EXCEPTION
#include <FEXCore/fextl/list.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/unordered_set.h>
#include <xbyak/xbyak.h>
#include <xbyak/xbyak_util.h>
#endif
namespace FEXCore {
// Data Zero Prohibited flag
// 0b0 = ZVA/GVA/GZVA permitted
// 0b1 = ZVA/GVA/GZVA prohibited
[[maybe_unused]] constexpr uint32_t DCZID_DZP_MASK = 0b1'0000;
// Log2 of the blocksize in 32-bit words
[[maybe_unused]] constexpr uint32_t DCZID_BS_MASK = 0b0'1111;
#ifdef _M_ARM_64
[[maybe_unused]]
static uint32_t GetDCZID() {
uint64_t Result {};
__asm("mrs %[Res], DCZID_EL0" : [Res] "=r"(Result));
return Result;
}
static uint32_t GetFPCR() {
uint64_t Result {};
__asm("mrs %[Res], FPCR" : [Res] "=r"(Result));
return Result;
}
static void SetFPCR(uint64_t Value) {
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
}
#else
static uint32_t GetDCZID() {
// Return unsupported
return DCZID_DZP_MASK;
}
#endif
static void OverrideFeatures(HostFeatures* Features) {
// Override features if the user has specifically called for it.
FEX_CONFIG_OPT(HostFeatures, HOSTFEATURES);
if (!HostFeatures()) {
// Early exit if no features are overriden.
return;
}
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
do { \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
const bool AlreadyEnabled = Features->FeatureName; \
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
Features->FeatureName = Result; \
} while (0)
#define GET_SINGLE_OPTION(name, enum_name) \
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive");
ENABLE_DISABLE_OPTION(SupportsAVX, AVX, AVX);
ENABLE_DISABLE_OPTION(SupportsAVX2, AVX2, AVX2);
ENABLE_DISABLE_OPTION(SupportsSVE, SVE, SVE);
ENABLE_DISABLE_OPTION(SupportsAFP, AFP, AFP);
ENABLE_DISABLE_OPTION(SupportsRCPC, LRCPC, LRCPC);
ENABLE_DISABLE_OPTION(SupportsTSOImm9, LRCPC2, LRCPC2);
ENABLE_DISABLE_OPTION(SupportsCSSC, CSSC, CSSC);
ENABLE_DISABLE_OPTION(SupportsPMULL_128Bit, PMULL128, PMULL128);
ENABLE_DISABLE_OPTION(SupportsRAND, RNG, RNG);
ENABLE_DISABLE_OPTION(SupportsCLZERO, CLZERO, CLZERO);
ENABLE_DISABLE_OPTION(SupportsAtomics, Atomics, ATOMICS);
ENABLE_DISABLE_OPTION(SupportsFCMA, FCMA, FCMA);
ENABLE_DISABLE_OPTION(SupportsFlagM, FlagM, FLAGM);
ENABLE_DISABLE_OPTION(SupportsFlagM2, FlagM2, FLAGM2);
ENABLE_DISABLE_OPTION(SupportsRPRES, RPRES, RPRES);
ENABLE_DISABLE_OPTION(SupportsPreserveAllABI, PRESERVEALLABI, PRESERVEALLABI);
GET_SINGLE_OPTION(Crypto, CRYPTO);
#undef ENABLE_DISABLE_OPTION
#undef GET_SINGLE_OPTION
if (EnableCrypto) {
Features->SupportsAES = true;
Features->SupportsCRC = true;
Features->SupportsSHA = true;
Features->SupportsPMULL_128Bit = true;
} else if (DisableCrypto) {
Features->SupportsAES = false;
Features->SupportsCRC = false;
Features->SupportsSHA = false;
Features->SupportsPMULL_128Bit = false;
}
}
HostFeatures::HostFeatures() {
#ifdef VIXL_SIMULATOR
auto Features = vixl::CPUFeatures::All();
// Vixl simulator doesn't support AFP.
Features.Remove(vixl::CPUFeatures::Feature::kAFP);
// Vixl simulator doesn't support RPRES.
Features.Remove(vixl::CPUFeatures::Feature::kRPRES);
#elif !defined(_WIN32)
auto Features = vixl::CPUFeatures::InferFromOS();
#else
// Need to use ID registers in WINE.
auto Features = vixl::CPUFeatures::InferFromIDRegisters();
#endif
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
SupportsSHA = Features.Has(vixl::CPUFeatures::Feature::kSHA1) && Features.Has(vixl::CPUFeatures::Feature::kSHA2);
SupportsAtomics = Features.Has(vixl::CPUFeatures::Feature::kAtomics);
SupportsRAND = Features.Has(vixl::CPUFeatures::Feature::kRNG);
// Only supported when FEAT_AFP is supported
SupportsAFP = Features.Has(vixl::CPUFeatures::Feature::kAFP);
SupportsRCPC = Features.Has(vixl::CPUFeatures::Feature::kRCpc);
SupportsTSOImm9 = Features.Has(vixl::CPUFeatures::Feature::kRCpcImm);
SupportsPMULL_128Bit = Features.Has(vixl::CPUFeatures::Feature::kPmull1Q);
SupportsCSSC = Features.Has(vixl::CPUFeatures::Feature::kCSSC);
SupportsFCMA = Features.Has(vixl::CPUFeatures::Feature::kFcma);
SupportsFlagM = Features.Has(vixl::CPUFeatures::Feature::kFlagM);
SupportsFlagM2 = Features.Has(vixl::CPUFeatures::Feature::kAXFlag);
SupportsRPRES = Features.Has(vixl::CPUFeatures::Feature::kRPRES);
Supports3DNow = true;
SupportsSSE4A = true;
#ifdef VIXL_SIMULATOR
// Hardcode enable SVE with 256-bit wide registers.
SupportsSVE = true;
SupportsAVX = true;
#else
SupportsSVE = Features.Has(vixl::CPUFeatures::Feature::kSVE);
SupportsAVX = Features.Has(vixl::CPUFeatures::Feature::kSVE2) && vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
#endif
// TODO: AVX2 is currently unsupported. Disable until the remaining features are implemented.
SupportsAVX2 = false;
SupportsBMI1 = true;
SupportsBMI2 = true;
SupportsCLWB = true;
// TODO: AFP is disabled until the scalar usage in the codebase can be audited to be working as expected.
SupportsAFP = false;
// RPRES has a dependency on AFP. Disable it until AFP is enabled.
SupportsRPRES = false;
if (!SupportsAtomics) {
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
}
#ifdef _M_ARM_64
// We need to get the CPU's cache line size
// We expect sane targets that have correct cacheline sizes across clusters
uint64_t CTR;
__asm volatile("mrs %[ctr], ctr_el0" : [ctr] "=r"(CTR));
DCacheLineSize = 4 << ((CTR >> 16) & 0xF);
ICacheLineSize = 4 << (CTR & 0xF);
// Test if this CPU supports float exception trapping by attempting to enable
// On unsupported these bits are architecturally defined as RAZ/WI
constexpr uint32_t ExceptionEnableTraps = (1U << 8) | // Invalid Operation float exception trap enable
(1U << 9) | // Divide by zero float exception trap enable
(1U << 10) | // Overflow float exception trap enable
(1U << 11) | // Underflow float exception trap enable
(1U << 12) | // Inexact float exception trap enable
(1U << 15); // Input Denormal float exception trap enable
uint32_t OriginalFPCR = GetFPCR();
uint32_t FPCR = OriginalFPCR | ExceptionEnableTraps;
SetFPCR(FPCR);
FPCR = GetFPCR();
SupportsFloatExceptions = (FPCR & ExceptionEnableTraps) == ExceptionEnableTraps;
// Set FPCR back to original just in case anything changed
SetFPCR(OriginalFPCR);
#endif
#ifdef VIXL_SIMULATOR
// simulator doesn't support dc(ZVA)
SupportsCLZERO = false;
// Simulator doesn't support SHA
SupportsSHA = false;
#else
// Check if we can support cacheline clears
uint32_t DCZID = GetDCZID();
if ((DCZID & DCZID_DZP_MASK) == 0) {
uint32_t DCZID_Log2 = DCZID & DCZID_BS_MASK;
uint32_t DCZID_Bytes = (1 << DCZID_Log2) * sizeof(uint32_t);
// If the DC ZVA size matches the emulated cache line size
// This means we can use the instruction
SupportsCLZERO = DCZID_Bytes == CPUIDEmu::CACHELINE_SIZE;
}
#endif
#if defined(_M_X86_64)
// Hardcoded cacheline size.
DCacheLineSize = 64U;
ICacheLineSize = 64U;
#if !defined(VIXL_SIMULATOR)
Xbyak::util::Cpu X86Features {};
SupportsAES = X86Features.has(Xbyak::util::Cpu::tAESNI);
SupportsCRC = X86Features.has(Xbyak::util::Cpu::tSSE42);
SupportsRAND = X86Features.has(Xbyak::util::Cpu::tRDRAND) && X86Features.has(Xbyak::util::Cpu::tRDSEED);
SupportsRCPC = true;
SupportsTSOImm9 = true;
Supports3DNow = X86Features.has(Xbyak::util::Cpu::t3DN) && X86Features.has(Xbyak::util::Cpu::tE3DN);
SupportsSSE4A = X86Features.has(Xbyak::util::Cpu::tSSE4a);
SupportsAVX = true;
SupportsAVX2 = true;
SupportsSHA = X86Features.has(Xbyak::util::Cpu::tSHA);
SupportsBMI1 = X86Features.has(Xbyak::util::Cpu::tBMI1);
SupportsBMI2 = X86Features.has(Xbyak::util::Cpu::tBMI2);
SupportsCLWB = X86Features.has(Xbyak::util::Cpu::tCLWB);
SupportsPMULL_128Bit = X86Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
// xbyak doesn't know how to check for CLZero
// First ensure we support a new enough extended CPUID function range
uint32_t data[4];
Xbyak::util::Cpu::getCpuid(0x8000'0000, data);
if (data[0] >= 0x8000'0008U) {
// CLZero defined in 8000_00008_EBX[bit 0]
Xbyak::util::Cpu::getCpuid(0x8000'0008, data);
SupportsCLZERO = data[1] & 1;
}
SupportsAFP = true;
SupportsFloatExceptions = true;
#endif
#endif
SupportsPreserveAllABI = FEXCORE_HAS_PRESERVE_ALL_ATTR;
OverrideFeatures(this);
}
} // namespace FEXCore
@@ -6,54 +6,59 @@
#include "Interface/IR/IR.h"
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static void LoadDeferredFCW(uint16_t NewFCW) {
auto PC = (NewFCW >> 8) & 3;
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW) {
softfloat_state State;
State.detectTininess = softfloat_tininess_afterRounding;
State.exceptionFlags = 0;
auto PC = (FCW >> 8) & 3;
switch (PC) {
case 0: extF80_roundingPrecision = 32; break;
case 2: extF80_roundingPrecision = 64; break;
case 3: extF80_roundingPrecision = 80; break;
case 0: State.roundingPrecision = 32; break;
case 2: State.roundingPrecision = 64; break;
case 3: State.roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
}
auto RC = (NewFCW >> 10) & 3;
auto RC = (FCW >> 10) & 3;
switch (RC) {
case 0: softfloat_roundingMode = softfloat_round_near_even; break;
case 1: softfloat_roundingMode = softfloat_round_min; break;
case 2: softfloat_roundingMode = softfloat_round_max; break;
case 3: softfloat_roundingMode = softfloat_round_minMag; break;
case 0: State.roundingMode = softfloat_round_near_even; break;
case 1: State.roundingMode = softfloat_round_min; break;
case 2: State.roundingMode = softfloat_round_max; break;
case 3: State.roundingMode = softfloat_round_minMag; break;
}
return State;
}
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, float src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, float src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t FCW, double src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
};
template<>
struct OpHandlers<IR::OP_F80CMP> {
template<uint32_t Flags>
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Flags & (1 << IR::FCMP_FLAG_LT) && lt) {
X80SoftFloat::FCMP(&State, Src1, Src2, &eq, &lt, &nan);
if (lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
if (Flags & (1 << IR::FCMP_FLAG_UNORDERED) && nan) {
if (nan) {
ResultFlags |= (1 << IR::FCMP_FLAG_UNORDERED);
}
if (Flags & (1 << IR::FCMP_FLAG_EQ) && eq) {
if (eq) {
ResultFlags |= (1 << IR::FCMP_FLAG_EQ);
}
return ResultFlags;
@@ -62,275 +67,261 @@ struct OpHandlers<IR::OP_F80CMP> {
template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF64(&State);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI16(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI64(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
auto rv = extF80_to_i32(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
auto rv = extF80_to_i32(&State, src, softfloat_round_minMag, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
return INT16_MIN;
} else {
return rv;
}
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i32(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i32(&State, src, softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i64(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i64(&State, src, softfloat_round_minMag, false);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTTOINT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t NewFCW, int16_t src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t FCW, int16_t src) {
return src;
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, int32_t src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, int32_t src) {
return src;
}
};
template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FRNDINT(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FRNDINT(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::F2XM1(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::F2XM1(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FTAN(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FTAN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSQRT(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSQRT(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSIN(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSIN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FCOS(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FCOS(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_EXP> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_EXP(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_SIG> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_SIG(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FADD(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FADD(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSUB(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSUB(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FMUL(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FMUL(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FDIV(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FDIV(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FYL2X(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FYL2X(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FATAN(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FATAN(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM1(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FREM1(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FREM(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSCALE(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSCALE(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return sin(src);
}
};
template<>
struct OpHandlers<IR::OP_F64COS> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return cos(src);
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return tan(src);
}
};
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return exp2(src) - 1.0;
}
};
template<>
struct OpHandlers<IR::OP_F64ATAN> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return atan2(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return fmod(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return remainder(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return src2 * log2(src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
double trunc = (double)(int64_t)(src2); // truncate
return src1 * exp2(trunc);
}
@@ -338,16 +329,16 @@ struct OpHandlers<IR::OP_F64SCALE> {
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(Src1);
Src1 = X80SoftFloat::FRNDINT(&State, Src1);
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1;
uint64_t Tmp = Src1.ToI64(&State);
X80SoftFloat Rv;
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
@@ -379,8 +370,7 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
template<>
struct OpHandlers<IR::OP_F80BCDLOAD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src) {
uint8_t* Src1 = reinterpret_cast<uint8_t*>(&Src);
uint64_t BCD {};
// We walk through each uint8_t and pull out the BCD encoding
@@ -35,14 +35,7 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t* Info) {
Info[Core::OPINDEX_F80CVTINT_TRUNC2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t);
Info[Core::OPINDEX_F80CVTINT_TRUNC4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t);
Info[Core::OPINDEX_F80CVTINT_TRUNC8] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t);
Info[Core::OPINDEX_F80CMP_0] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>);
Info[Core::OPINDEX_F80CMP_1] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>);
Info[Core::OPINDEX_F80CMP_2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>);
Info[Core::OPINDEX_F80CMP_3] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>);
Info[Core::OPINDEX_F80CMP_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>);
Info[Core::OPINDEX_F80CMP_5] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>);
Info[Core::OPINDEX_F80CMP_6] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>);
Info[Core::OPINDEX_F80CMP_7] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>);
Info[Core::OPINDEX_F80CMP] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle);
Info[Core::OPINDEX_F80CVTTOINT_2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
Info[Core::OPINDEX_F80CVTTOINT_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
@@ -154,17 +147,8 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers {
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags),
SupportsPreserveAllABI};
*Info = {FABI_I64_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle,
(Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP), SupportsPreserveAllABI};
return true;
}
@@ -185,22 +169,22 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
break;
}
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
#define COMMON_UNARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
return true; \
}
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
#define COMMON_BINARY_X87_OP(OP) \
case IR::OP_F80##OP: { \
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
return true; \
return true; \
}
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
#define COMMON_F64_OP(OP) \
case IR::OP_F64##OP: { \
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
return true; \
return true; \
}
// Unary
@@ -50,22 +50,22 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
// Bits are arranged as:
// Bit #: 3 2 1 0
// [OF | CF | SF | ZF]
// [SF | ZF | CF | OF]
uint32_t flags = 0;
flags |= (valid_rhs < upper_limit) ? 0b01 : 0b00;
flags |= (valid_lhs < upper_limit) ? 0b10 : 0b00;
flags |= (valid_rhs < upper_limit) ? 0b0100 : 0b0000;
flags |= (valid_lhs < upper_limit) ? 0b1000 : 0b0000;
const uint32_t result = HandlePolarity(aggregation, control, upper_limit, valid_rhs);
if (result != 0) {
flags |= 0b0100;
flags |= 0b0010;
}
if ((result & 1) != 0) {
flags |= 0b1000;
flags |= 0b0001;
}
// We tack the flags on top of the result to avoid needing to handle
// multiple return values in the JITs.
return result | (flags << 16);
// We track the flags in the usual NZCV bit position so we can msr them
// later. Avoids handling flags natively in JIT.
return result | (flags << 28);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetExplicitLength(uint64_t reg, uint16_t control) {
File diff suppressed because it is too large. Load diff
@@ -6,7 +6,6 @@ $end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -69,8 +68,8 @@ DEF_OP(CASPair) {
DEF_OP(CAS) {
auto Op = IROp->C<IR::IROp_CAS>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
// DataSrc = *Src1
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
// This will write to memory! Careful!
@@ -79,13 +78,6 @@ DEF_OP(CAS) {
auto Desired = GetReg(Op->Desired.ID());
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mov(EmitSize, TMP2, Expected);
casal(SubEmitSize, TMP2, Desired, MemSrc);
@@ -96,9 +88,9 @@ DEF_OP(CAS) {
ARMEmitter::SingleUseForwardLabel LoopExpected;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
if (OpSize == 1) {
if (IROp->Size == 1) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTB, 0);
} else if (OpSize == 2) {
} else if (IROp->Size == 2) {
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTH, 0);
} else {
cmp(EmitSize, TMP2, Expected);
@@ -120,19 +112,12 @@ DEF_OP(CAS) {
DEF_OP(AtomicAdd) {
auto Op = IROp->C<IR::IROp_AtomicAdd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
staddl(SubEmitSize, Src, MemSrc);
} else {
@@ -147,19 +132,12 @@ DEF_OP(AtomicAdd) {
DEF_OP(AtomicSub) {
auto Op = IROp->C<IR::IROp_AtomicSub>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
staddl(SubEmitSize, TMP2, MemSrc);
@@ -175,19 +153,12 @@ DEF_OP(AtomicSub) {
DEF_OP(AtomicAnd) {
auto Op = IROp->C<IR::IROp_AtomicAnd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
stclrl(SubEmitSize, TMP2, MemSrc);
@@ -203,19 +174,12 @@ DEF_OP(AtomicAnd) {
DEF_OP(AtomicCLR) {
auto Op = IROp->C<IR::IROp_AtomicCLR>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
stclrl(SubEmitSize, Src, MemSrc);
} else {
@@ -230,19 +194,12 @@ DEF_OP(AtomicCLR) {
DEF_OP(AtomicOr) {
auto Op = IROp->C<IR::IROp_AtomicOr>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
stsetl(SubEmitSize, Src, MemSrc);
} else {
@@ -257,19 +214,12 @@ DEF_OP(AtomicOr) {
DEF_OP(AtomicXor) {
auto Op = IROp->C<IR::IROp_AtomicXor>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
steorl(SubEmitSize, Src, MemSrc);
} else {
@@ -284,18 +234,11 @@ DEF_OP(AtomicXor) {
DEF_OP(AtomicNeg) {
auto Op = IROp->C<IR::IROp_AtomicNeg>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -312,7 +255,7 @@ DEF_OP(AtomicSwap) {
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
@@ -333,19 +276,12 @@ DEF_OP(AtomicSwap) {
DEF_OP(AtomicFetchAdd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldaddal(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
@@ -361,19 +297,12 @@ DEF_OP(AtomicFetchAdd) {
DEF_OP(AtomicFetchSub) {
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
neg(EmitSize, TMP2, Src);
ldaddal(SubEmitSize, TMP2, GetReg(Node), MemSrc);
@@ -390,19 +319,12 @@ DEF_OP(AtomicFetchSub) {
DEF_OP(AtomicFetchAnd) {
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
mvn(EmitSize, TMP2, Src);
ldclral(SubEmitSize, TMP2, GetReg(Node), MemSrc);
@@ -419,19 +341,12 @@ DEF_OP(AtomicFetchAnd) {
DEF_OP(AtomicFetchCLR) {
auto Op = IROp->C<IR::IROp_AtomicFetchCLR>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldclral(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
@@ -447,19 +362,12 @@ DEF_OP(AtomicFetchCLR) {
DEF_OP(AtomicFetchOr) {
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldsetal(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
@@ -475,19 +383,12 @@ DEF_OP(AtomicFetchOr) {
DEF_OP(AtomicFetchXor) {
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
auto Src = GetReg(Op->Value.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
if (CTX->HostFeatures.SupportsAtomics) {
ldeoral(SubEmitSize, Src, GetReg(Node), MemSrc);
} else {
@@ -503,18 +404,11 @@ DEF_OP(AtomicFetchXor) {
DEF_OP(AtomicFetchNeg) {
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
uint8_t OpSize = IROp->Size;
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
const auto EmitSize = ConvertSize(IROp);
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
auto MemSrc = GetReg(Op->Addr.ID());
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
ARMEmitter::BackwardLabel LoopTop;
Bind(&LoopTop);
ldaxr(SubEmitSize, TMP2, MemSrc);
@@ -7,7 +7,6 @@ $end_info$
#include "Interface/Context/Context.h"
#include "FEXCore/IR/IR.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
@@ -55,7 +54,8 @@ DEF_OP(ExitFunction) {
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
#ifdef _M_ARM_64EC
if (RtlIsEcCode(NewRIP)) {
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, NewRIP);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
} else {
@@ -101,39 +101,13 @@ DEF_OP(Jump) {
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
}
static ARMEmitter::Condition MapBranchCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
}
}
DEF_OP(CondJump) {
auto Op = IROp->C<IR::IROp_CondJump>();
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
if (Op->FromNZCV) {
b(MapBranchCC(Op->Cond), TrueTargetLabel);
b(MapCC(Op->Cond), TrueTargetLabel);
} else {
[[maybe_unused]] uint64_t Const;
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
@@ -210,7 +184,7 @@ DEF_OP(Syscall) {
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask);
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -311,7 +285,7 @@ DEF_OP(InlineSyscall) {
if ((Op->Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r1);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -5,7 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
@@ -18,12 +18,7 @@ DEF_OP(VInsGPR) {
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
const auto SubEmitSize = ConvertSubRegSize8(IROp);
const auto ElementsPer128Bit = 16 / ElementSize;
const auto Dst = GetVReg(Node);
@@ -66,7 +61,7 @@ DEF_OP(VInsGPR) {
// Inserts the GPR value into the given V register.
// Also automatically adjusts the index in the case of using the
// moved upper lane.
const auto Insert = [&](const FEXCore::ARMEmitter::VRegister& reg, int index) {
const auto Insert = [&](const ARMEmitter::VRegister& reg, int index) {
if (InUpperLane) {
index -= ElementsPer128Bit;
}
@@ -117,16 +112,7 @@ DEF_OP(VDupFromGPR) {
const auto Src = GetReg(Op->Src.ID());
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto ElementSize = IROp->ElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} element size: {}",
__func__, ElementSize);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ARMEmitter::SubRegSize::i8Bit;
const auto SubEmitSize = ConvertSubRegSize8(IROp);
if (HostSupportsSVE256 && Is256Bit) {
dup(SubEmitSize, Dst.Z(), Src);
@@ -216,14 +202,9 @@ DEF_OP(Vector_SToF) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE256 && Is256Bit) {
@@ -253,14 +234,9 @@ DEF_OP(Vector_FToZS) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE256 && Is256Bit) {
@@ -289,14 +265,8 @@ DEF_OP(Vector_FToS) {
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ARMEmitter::SubRegSize::i16Bit;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -323,15 +293,10 @@ DEF_OP(Vector_FToF) {
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -353,23 +318,23 @@ DEF_OP(Vector_FToF) {
switch (Conv) {
case 0x0402: { // Float <- Half
zip1(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
zip1(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0804: { // Double <- Float
zip1(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
zip1(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
fcvtlt(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
break;
}
case 0x0204: { // Half <- Float
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
fcvtnt(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
uzp2(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
case 0x0408: { // Float <- Double
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
fcvtnt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
uzp2(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
break;
}
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
@@ -391,18 +356,46 @@ DEF_OP(Vector_FToF) {
}
}
DEF_OP(VFCVTL2) {
const auto Op = IROp->C<IR::IROp_VFCVTL2>();
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
fcvtl2(SubEmitSize, Dst.D(), Vector.D());
}
DEF_OP(VFCVTN2) {
const auto Op = IROp->C<IR::IROp_VFCVTN2>();
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Dst = GetVReg(Node);
const auto VectorLower = GetVReg(Op->VectorLower.ID());
const auto VectorUpper = GetVReg(Op->VectorUpper.ID());
auto Lower = VectorLower;
if (Dst != VectorLower) {
mov(VTMP1.Q(), VectorLower.Q());
Lower = VTMP1;
}
fcvtn2(SubEmitSize, Lower.Q(), VectorUpper.Q());
if (Dst != VectorLower) {
mov(Dst.Q(), Lower.Q());
}
}
DEF_OP(Vector_FToI) {
const auto Op = IROp->C<IR::IROp_Vector_FToI>();
const auto OpSize = IROp->Size;
const auto ElementSize = Op->Header.ElementSize;
const auto SubEmitSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ARMEmitter::SubRegSize::i16Bit;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
@@ -425,15 +418,15 @@ DEF_OP(Vector_FToI) {
// frinti having AdvSIMD, AdvSIMD scalar, and an SVE version),
// we can't just use a lambda without some seriously ugly casting.
// This is fairly self-contained otherwise.
#define ROUNDING_FN(name) \
if (ElementSize == 2) { \
name(Dst.H(), Vector.H()); \
#define ROUNDING_FN(name) \
if (ElementSize == 2) { \
name(Dst.H(), Vector.H()); \
} else if (ElementSize == 4) { \
name(Dst.S(), Vector.S()); \
name(Dst.S(), Vector.S()); \
} else if (ElementSize == 8) { \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
name(Dst.D(), Vector.D()); \
} else { \
FEX_UNREACHABLE; \
}
switch (Op->Round) {
@@ -457,5 +450,63 @@ DEF_OP(Vector_FToI) {
}
}
DEF_OP(Vector_F64ToI32) {
const auto Op = IROp->C<IR::IROp_Vector_F64ToI32>();
const auto OpSize = IROp->Size;
const auto Round = Op->Round;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE128 || HostSupportsSVE256) {
const auto Mask = Is256Bit ? PRED_TMP_32B.Merging() : PRED_TMP_16B.Merging();
// First step is to round the f64 values to integrals (frint*)
// Then convert to integers using fcvtzs.
auto CVTReg = Dst.Z();
switch (Round) {
case IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Towards_Zero.Val: CVTReg = Vector.Z(); break;
case IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
}
fcvtzs(Dst.Z(), ARMEmitter::SubRegSize::i32Bit, Mask, CVTReg, ARMEmitter::SubRegSize::i64Bit);
///< Fixup format of register that fcvtzs returns.
uzp1(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
if (Op->EnsureZeroUpperHalf) {
///< Match CVTPD2DQ/CVTTPD2DQ behaviour if necessary by zeroing the upper bits here.
if (Is256Bit) {
mov(Dst.Q(), Dst.Q());
} else {
mov(Dst.D(), Dst.D());
}
}
} else {
// This has a known precision issue that isn't easily resolvable without throwing away performance.
// Doing the conversion in multi-stage steps has an issue that you can lose precision in the f32->i32 step if your source was f64.
// To get around this with ASIMD FEX needs to use fcvtzs (Scalar, Integer, to GPR) for each F64 to be directly converted to i32.
// This is a very costly transform that the SVE path doesn't need to do since it supports f64->i32 directly.
// If this precision issue is necessary then we can add an option for it in the future.
///< Round float to integral depending on rounding mode.
switch (Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
}
// Now narrow from f64 to f32.
fcvtn(ARMEmitter::SubRegSize::i32Bit, Dst.Q(), Dst.Q());
///< Convert the two F32 integrals to real integers.
fcvtzs(ARMEmitter::SubRegSize::i32Bit, Dst.D(), Dst.D());
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -5,9 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "Interface/IR/Passes/RegisterAllocationPass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
@@ -1,18 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
} // namespace FEXCore::CPU
+17 -35
View File
@@ -13,7 +13,6 @@ $end_info$
#include "FEXCore/Utils/Telemetry.h"
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/LookupCache.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
@@ -48,8 +47,8 @@ static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
static int64_t LDIV(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
@@ -60,8 +59,8 @@ static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
static int64_t LREM(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
@@ -469,13 +468,13 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
uintptr_t branch = (uintptr_t)(Record)-8;
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
FEXCore::ARMEmitter::SingleUseForwardLabel l_BranchHost;
ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
ARMEmitter::SingleUseForwardLabel l_BranchHost;
emit.ldr(TMP1, &l_BranchHost);
emit.blr(TMP1);
emit.Bind(&l_BranchHost);
emit.dc64(LinkerAddress);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 8);
ARMEmitter::Emitter::ClearICache((void*)branch, 8);
}
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
@@ -497,12 +496,12 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Fram
uintptr_t branch = (uintptr_t)(Record)-8;
auto offset = HostCode / 4 - branch / 4;
if (vixl::IsInt26(offset)) {
if (ARMEmitter::Emitter::IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
emit.b(offset);
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 4);
ARMEmitter::Emitter::ClearICache((void*)branch, 4);
// Add de-linking handler
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
@@ -522,27 +521,20 @@ void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::NodeID Node) {}
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
, Arm64Emitter(ctx)
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsAVX}
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE128}
, HostSupportsSVE256 {ctx->HostFeatures.SupportsSVE256}
, HostSupportsAVX256 {ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256}
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
, CTX {ctx} {
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
RAPass->AllocateRegisterSet(RegisterClasses);
RAPass->AddRegisters(FEXCore::IR::GPRClass, GeneralRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, StaticRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRClass, GeneralFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, StaticFPRegisters.size());
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, GeneralPairRegisters.size());
RAPass->AddRegisters(FEXCore::IR::ComplexClass, 1);
for (uint32_t i = 0; i < GeneralPairRegisters.size(); ++i) {
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2, FEXCore::IR::GPRPairClass, i);
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2 + 1, FEXCore::IR::GPRPairClass, i);
}
RAPass->PairRegs = PairRegisters;
{
// Set up pointers that the JIT needs to load
@@ -669,7 +661,7 @@ bool Arm64JITCore::IsGPRPair(IR::NodeID Node) const {
}
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
FEXCore::IR::RegisterAllocationData* RAData) {
const FEXCore::IR::RegisterAllocationData* RAData) {
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
JumpTargets.clear();
@@ -732,14 +724,12 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
}
// LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
SpillSlots = RAData->SpillSlots();
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (vixl::aarch64::Assembler::IsImmAddSub(TotalSpillSlotsSize)) {
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
@@ -882,7 +872,7 @@ void Arm64JITCore::ResetStack() {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (vixl::aarch64::Assembler::IsImmAddSub(TotalSpillSlotsSize)) {
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
// Too big to fit in a 12bit immediate
@@ -895,12 +885,4 @@ fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl*
return fextl::make_unique<Arm64JITCore>(ctx, Thread);
}
CPUBackendFeatures GetArm64JITBackendFeatures() {
return CPUBackendFeatures {
.SupportsFlags = true,
.SupportsSaturatingRoundingShifts = true,
.SupportsVTBL2 = true,
};
}
} // namespace FEXCore::CPU
@@ -8,7 +8,6 @@ $end_info$
#pragma once
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/CPUBackend.h"
#include "Interface/Core/Dispatcher/Dispatcher.h"
#include "Interface/IR/IR.h"
@@ -24,6 +23,8 @@ $end_info$
#include <FEXCore/fextl/string.h>
#include <FEXCore/fextl/vector.h>
#include <CodeEmitter/Emitter.h>
#include <array>
#include <cstdint>
#include <utility>
@@ -46,7 +47,7 @@ public:
[[nodiscard]]
CPUBackend::CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
FEXCore::IR::RegisterAllocationData* RAData) override;
const FEXCore::IR::RegisterAllocationData* RAData) override;
[[nodiscard]]
void* MapRegion(void* HostPtr, uint64_t, uint64_t) override {
@@ -71,6 +72,7 @@ private:
const bool HostSupportsSVE128 {};
const bool HostSupportsSVE256 {};
const bool HostSupportsAVX256 {};
const bool HostSupportsRPRES {};
const bool HostSupportsAFP {};
@@ -83,7 +85,7 @@ private:
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
[[nodiscard]]
FEXCore::ARMEmitter::Register GetReg(IR::NodeID Node) const {
ARMEmitter::Register GetReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
@@ -98,7 +100,7 @@ private:
}
[[nodiscard]]
FEXCore::ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
@@ -113,12 +115,12 @@ private:
}
[[nodiscard]]
std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
std::pair<ARMEmitter::Register, ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
const auto Reg = GetPhys(Node);
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRPairClass.Val, "Unexpected Class: {}", Reg.Class);
return GeneralPairRegisters[Reg.Reg];
return std::make_pair(GeneralRegisters[Reg.Reg], GeneralRegisters[Reg.Reg + 1]);
}
[[nodiscard]]
@@ -134,7 +136,7 @@ private:
}
[[nodiscard]]
FEXCore::ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
uint64_t Const;
if (IsInlineConstant(Src, &Const)) {
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
@@ -154,6 +156,99 @@ private:
ARMEmitter::ShiftType::ROR;
}
[[nodiscard]]
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
return Op->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
}
[[nodiscard]]
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->Size == 4 || Op->Size == 8, "Invalid size");
return ConvertSize(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8 || ElementSize == 16, "Invalid size");
return ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
ARMEmitter::SubRegSize::i128Bit;
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
return ConvertSubRegSize16(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(uint8_t ElementSize) {
LOGMAN_THROW_AA_FMT(ElementSize != 16, "Invalid size");
return ConvertSubRegSize16(ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
return ConvertSubRegSize8(Op->ElementSize);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 8, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
return ConvertSubRegSize8(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
return ARMEmitter::ToVectorSizePair(ConvertSubRegSize16(Op));
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 16, "Invalid size");
return ConvertSubRegSizePair16(Op);
}
[[nodiscard]]
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
return ConvertSubRegSizePair8(Op);
}
[[nodiscard]]
ARMEmitter::Condition MapCC(IR::CondClassType Cond) {
switch (Cond.Val) {
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
case FEXCore::IR::COND_VS:
case FEXCore::IR::COND_VC:
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
}
}
[[nodiscard]]
bool IsFPR(IR::NodeID Node) const;
[[nodiscard]]
@@ -162,8 +257,8 @@ private:
bool IsGPRPair(IR::NodeID Node) const;
[[nodiscard]]
FEXCore::ARMEmitter::ExtendedMemOperand GenerateMemOperand(
uint8_t AccessSize, FEXCore::ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
// NOTE: Will use TMP1 as a way to encode immediates that happen to fall outside
// the limits of the scalar plus immediate variant of SVE load/stores.
@@ -171,8 +266,8 @@ private:
// TMP1 is safe to use again once this memory operand is used with its
// equivalent loads or stores that this was called for.
[[nodiscard]]
FEXCore::ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, FEXCore::ARMEmitter::Register Base,
IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
[[nodiscard]]
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
@@ -187,7 +282,7 @@ private:
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
void EmitDetectionString();
IR::RegisterAllocationPass* RAPass;
IR::RegisterAllocationData* RAData;
const IR::RegisterAllocationData* RAData;
FEXCore::Core::DebugData* DebugData;
void ResetStack();
@@ -250,6 +345,10 @@ private:
uint32_t SpillSlots {};
using OpType = void (Arm64JITCore::*)(const IR::IROp_Header* IROp, IR::NodeID Node);
using ScalarFMAOpCaller =
std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3)>;
void VFScalarFMAOperation(uint8_t OpSize, uint8_t ElementSize, ScalarFMAOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2, ARMEmitter::VRegister Addend);
using ScalarBinaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2)>;
void VFScalarOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit,
ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2);
@@ -257,6 +356,10 @@ private:
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
void Emulate128BitGather(size_t Size, size_t ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, size_t VectorIndexSize,
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
File diff suppressed because it is too large. Load diff
@@ -10,7 +10,6 @@ $end_info$
#endif
#include "Interface/Context/Context.h"
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
#include "FEXCore/Debug/InternalThreadState.h"
@@ -28,9 +27,9 @@ DEF_OP(GuestOpcode) {
DEF_OP(Fence) {
auto Op = IROp->C<IR::IROp_Fence>();
switch (Op->Fence) {
case IR::Fence_Load.Val: dmb(FEXCore::ARMEmitter::BarrierScope::LD); break;
case IR::Fence_LoadStore.Val: dmb(FEXCore::ARMEmitter::BarrierScope::SY); break;
case IR::Fence_Store.Val: dmb(FEXCore::ARMEmitter::BarrierScope::ST); break;
case IR::Fence_Load.Val: dmb(ARMEmitter::BarrierScope::LD); break;
case IR::Fence_LoadStore.Val: dmb(ARMEmitter::BarrierScope::SY); break;
case IR::Fence_Store.Val: dmb(ARMEmitter::BarrierScope::ST); break;
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
}
}
@@ -99,6 +98,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg(Op->RoundMode.ID());
auto MXCSR = GetReg(Op->MXCSR.ID());
// As above, setup the rounding flags in [31:30]
rbit(ARMEmitter::Size::i32Bit, TMP2, Src);
@@ -117,10 +117,48 @@ DEF_OP(SetRoundingMode) {
lsr(ARMEmitter::Size::i64Bit, TMP2, Src, 2);
bfi(ARMEmitter::Size::i64Bit, TMP1, TMP2, 24, 1);
if (Op->SetDAZ && HostSupportsAFP) {
// Extract DAZ from MXCSR and insert to in FPCR.FIZ
bfxil(ARMEmitter::Size::i64Bit, TMP1, MXCSR, 6, 1);
}
// Now save the new FPCR
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
}
DEF_OP(PushRoundingMode) {
auto Op = IROp->C<IR::IROp_PushRoundingMode>();
auto Dest = GetReg(Node);
// Save the old rounding mode
mrs(Dest, ARMEmitter::SystemRegister::FPCR);
// vixl simulator doesn't support anything beyond ties-to-even rounding
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
return;
}
// Insert the rounding flags, reversing the mode bits as above
if (Op->RoundMode == 3) {
orr(ARMEmitter::Size::i64Bit, TMP1, Dest, 3 << 22);
} else if (Op->RoundMode == 0) {
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(3 << 22));
} else {
LOGMAN_THROW_AA_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(Op->RoundMode << 22));
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, (Op->RoundMode == 2 ? 1 : 2) << 22);
}
// Now save the new FPCR
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
}
DEF_OP(PopRoundingMode) {
auto Op = IROp->C<IR::IROp_PopRoundingMode>();
msr(ARMEmitter::SystemRegister::FPCR, GetReg(Op->FPCR.ID()));
}
DEF_OP(Print) {
auto Op = IROp->C<IR::IROp_Print>();
@@ -195,7 +233,7 @@ DEF_OP(ProcessorID) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -11,12 +11,14 @@ namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(ExtractElementPair) {
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
LOGMAN_THROW_AA_FMT(Op->Header.Size == 4 || Op->Header.Size == 8, "Invalid size");
const auto EmitSize = Op->Header.Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
const auto Src = GetRegPair(Op->Pair.ID());
const std::array<ARMEmitter::Register, 2> Regs = {Src.first, Src.second};
mov(EmitSize, GetReg(Node), Regs[Op->Element]);
const auto Dst = GetReg(Node);
const auto Pair = GetRegPair(Op->Pair.ID());
const auto Src = Op->Element == 0 ? Pair.first : Pair.second;
if (Dst != Src) {
mov(ConvertSize48(IROp), Dst, Src);
}
}
DEF_OP(CreateElementPair) {
@@ -42,5 +44,25 @@ DEF_OP(CreateElementPair) {
}
}
DEF_OP(Copy) {
auto Op = IROp->C<IR::IROp_Copy>();
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Source.ID()));
}
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A.ID()), B = GetReg(Op->B.ID());
LOGMAN_THROW_A_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
mov(ARMEmitter::Size::i64Bit, A, B);
mov(ARMEmitter::Size::i64Bit, B, TMP1);
}
DEF_OP(Swap2) {
// Implemented above
}
#undef DEF_OP
} // namespace FEXCore::CPU
File diff suppressed because it is too large. Load diff
@@ -17,6 +17,5 @@ class CPUBackend;
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
Loaded 100 of 618 files, more files were not shown because too many files have changed in this diff. Show more