Compare commits

...
367 Commits
Author SHA1 Message Date
Ryan Houdek fb60a8a032 Docs: Update for release FEX-2408 2024-08-12 15:05:37 -07:00
Alyssa Rosenzweig f40bc134df Merge pull request #3945 from Sonicadvance1/config_nonnullable
Config: Little assume non-null check
2024-08-11 15:23:37 -04:00
Ryan Houdek f3811f04bd Config: Little assume non-null check
Removes a simple runtime nullcheck in Config::Layer::Set. Since we never pass a
nullptr to this.
2024-08-11 10:25:53 -07:00
Alyssa Rosenzweig 94bb7eb311 Merge pull request #3937 from Sonicadvance1/fix_script
Scripts: Fix issue in aarch64_fit_native
2024-08-10 20:07:43 -04:00
Ryan Houdek 03ca3e7e68 Merge pull request #3934 from bylaws/wow64-b
WOW64: Support the JIT API as used by Windows
2024-08-10 09:07:58 -07:00
Ryan Houdek 2c3e6cbe65 Merge pull request #3932 from bylaws/arm64-callchk
ARM64EC: Install a custom call checker to bypass NTDLL function patches
2024-08-10 09:07:31 -07:00
Alyssa Rosenzweig f9bdf0bd01 Merge pull request #3938 from Sonicadvance1/fix_vpblend_test
unittests: Fixes vpblend unittest
2024-08-10 11:08:03 -04:00
Ryan Houdek 4afc7adb05 unittests: Fixes vpblend unittest
This typo was causing undefined data to be used in the unittest, showed
up in debug builds.
2024-08-10 07:43:25 -07:00
Ryan Houdek 2152d1b2e9 Scripts: Fix issue in aarch64_fit_native
Apparently I messed this up in testing, is now fixed.
2024-08-09 21:04:26 -07:00
Ryan Houdek 0ecfc651b6 Merge pull request #3931 from Sonicadvance1/move_hostfeatures_init
FEXCore: Pass HostFeatures in to CreateNewContext directly
2024-08-09 20:30:17 -07:00
Ryan Houdek 4a3250ddea Merge pull request #3928 from bylaws/winval
InvalidationTracker: Better match Windows code invalidation behaviour
2024-08-09 20:29:57 -07:00
Alyssa Rosenzweig 633f624a69 Merge pull request #3930 from Sonicadvance1/hostfeatures_only_harnessrunner
HostFeatures: Removes feature flags always supported by FEX
2024-08-09 15:12:04 -04:00
Billy Laws 26a8a2717c WOW64: Mark the CPU area context as dirty initially
After thread creation, the WOW64 CPU area context needs to be flushed
into the FEX state before entering the JIT. Wine explicitly calls
BTCpuSetContext to trigger this but Windows doesn't.
2024-08-09 11:57:09 +00:00
Billy Laws 4c0e6d5779 WOW64: Shift down used TLS slots
Fixes a crash on native Windows.
2024-08-09 11:57:09 +00:00
Billy Laws a350ef5d1b WOW64: Match the Windows function protoypes 2024-08-09 11:57:09 +00:00
Billy Laws 9d9bd750e2 ARM64EC: Install a custom call checker to bypass NTDLL function patches
Some programs will hook the NTDLL exports that FEX depends on, the
regular ARM64EC call checker will detect such patches and invoke the
JIT to run them, which leads to infinite recursion if those same
exports are used during code compilation. Fix this by resolving all
patchable FFSs to their native ARM implementations for all indirect
calls performed by FEX, skipping any x86 patches.
2024-08-09 11:48:18 +00:00
Ryan Houdek 85d1b573ef Merge pull request #3927 from bylaws/winafp
ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit
2024-08-08 22:21:23 -07:00
Ryan Houdek 2f8c5b4820 FEXCore: Pass HostFeatures in to CreateNewContext directly
The class constructor for ContextImpl::CPUID requires HostFeatures to be
available at construction time. Pass the host features struct directly
through during construction time instead, which cleans up the interface
slightly and fixes that issue.
2024-08-08 21:02:41 -07:00
Ryan Houdek a1f55f0b0b HostFeatures: Removes feature flags always supported by FEX
These are only missing if using the hostrunner and the CI machine
doesn't support that particular feature. FEX otherwise always supports
these feature flags so they don't need to exist as options.

Just check the feature bit directly in the HostRunner frontend for these
bits.
2024-08-08 19:05:55 -07:00
Ryan Houdek 7b1d9540b7 Merge pull request #3925 from bylaws/arm64ecrt
ARM64EC: Introduce FEX-side CRT and Windows API replacements
2024-08-08 17:57:25 -07:00
Ryan Houdek 1007f874bf Merge pull request #3926 from bylaws/windef
FEXCore: Drop deferred signal handling on Windows
2024-08-08 17:33:44 -07:00
Ryan Houdek 24ea4b7537 Merge pull request #3924 from bylaws/svc
ARM64EC: Handle direct syscall instructions
2024-08-08 17:32:56 -07:00
Billy Laws 93ab57454a InvalidationTracker: Better match Windows code invalidation behaviour
When given a NULL base address, Windows invalidation callbacks will
ignore the given size and invalidate all code.
2024-08-08 12:46:49 +00:00
Mai 4882f10536 Merge pull request #3888 from Sonicadvance1/avx128_optimize_blends
AVX128: Optimize blends
2024-08-07 17:08:19 -04:00
Billy Laws fe43a2bcb2 ARM64EC: Set appropriate AFP and SVE256 state on JIT entry/exit 2024-08-07 18:34:35 +00:00
Billy Laws 6700511cdf FEXCore: Drop deferred signal handling on Windows
The async signal issues this handles do not exist on Windows.
2024-08-07 18:31:48 +00:00
Billy Laws e836639427 ARM64EC: Manually define the ARM64EC linker structures 2024-08-07 15:49:41 +00:00
Billy Laws 7eb3f33162 Update jemalloc submodule 2024-08-07 15:49:41 +00:00
Billy Laws b35a514ea6 CMake: Don't link CMake's set of extra system libraries on Windows
These are already linked in by default with clang, sp having these set
here only served to prevent the -nostdlib compiler option from having
any effect, as CMake will always explicitly include the libs in the
compiler cmdline.
2024-08-07 15:49:41 +00:00
Billy Laws 8f372821f8 Windows: Use the FEX CRT 2024-08-07 15:49:41 +00:00
Billy Laws 6376d06bed Windows: Introduce a minimal Windows API replacement 2024-08-07 15:49:41 +00:00
Billy Laws 1bb293d4be Windows: Pull in some math and formatting functions from Musl 2024-08-07 15:49:41 +00:00
Billy Laws 0e069f1e97 Windows: Introduce a minimal CRT replacement
It is dangerous for FEX to rely on the system CRT as calls can have side
effects that are also visible to the running application, and if patches
are applied to any CRT exports used during compilation the call checker
would trigger a reentry into the JIT to compile the patch and hence
deadlock. Only functions that FEX actively uses are implemented, with
the rest triggering an abort.
2024-08-07 15:49:41 +00:00
Billy Laws 9074b810a9 CMake: Always enable jemalloc for MinGW builds 2024-08-07 15:49:41 +00:00
Mai c5fe8723d3 Merge pull request #3918 from Sonicadvance1/constexpr_config_maps
Config: Converts two LUT maps over linear scan arrays
2024-08-07 10:40:22 -04:00
Ryan Houdek 434bffac33 Merge pull request #3916 from Sonicadvance1/frontend_hostfeatures
FEX: Moves HostFeatures querying to the frontend
2024-08-07 05:33:55 -07:00
Ryan Houdek e84848b16b FEX: Moves HostFeatures querying to the frontend
This moves the CPU feature querying to the frontend. The primary purpose
here is for the wow64 frontend to not require linux-isms for querying
these features. This is required since non-Linux environments don't have
the "CPUID" feature for reading EL1 MSRs in EL0.

Wiring up the remaining wow64 registry querying is left for a future
exercise.

This also technically removes an xbyak requirement from FEXCore for when
building the x86 Test harness runner, but that doesn't really matter for
regular use cases.
2024-08-07 05:26:02 -07:00
Ryan Houdek 69ed39d49e Merge pull request #3892 from Sonicadvance1/optimize_vpermq
AVX128: Optimize all cases of vpermq
2024-08-06 20:07:28 -07:00
Ryan Houdek 230bde6aef InstcountCI: Adds vpermq coverage 2024-08-06 09:08:30 -07:00
Ryan Houdek c24d7aacba unittests/ASM: Adds vpermq test that covers all immediate encodings
To ensure we cover all tests when optimizing.
2024-08-06 09:08:30 -07:00
Ryan Houdek e613876e9d AVX128: Optimize all cases of vpermq
Started by cherry-picking some cases from the variants that appeared when running
Steam, games, AV1 convolve tests, openssl, ffmpeg, libjpeg-turbo,
openh264, libvpx, gemmlowp, libyuv, and dav1d.

Then turned it around and optimized them all since all variants end up
needing to be split in to two halves, that effectively means we need to
have 16 implementations, plus a couple of special cases for duplicated
results.

Fixes #3795
2024-08-06 09:08:30 -07:00
Mai 1473129a8f Merge pull request #3920 from Sonicadvance1/fix_newline_asm
SpinWaitLock: Fixes missing newline in asm
2024-08-06 12:08:14 -04:00
Alyssa Rosenzweig c42808cb70 Merge pull request #3904 from Sonicadvance1/packaging
Scripts: Workaround deprecated parse_version
2024-08-06 11:52:59 -04:00
Ryan Houdek 054c119e2e Config: Converts two LUT maps over linear scan arrays
These two maps used for environment lookup translations were getting
globally initialized and then registers with atexit handlers.

Switch over to a constexpr array and just do linear scans. This plus
short-circuiting the environment loader so it skips all entries that
don't start with `FEX_` has the side benefit of cutting the CPU time to
1/10th the time.

This plus #3917 removes the global static initializers entirely from
this file.
2024-08-06 07:44:08 -07:00
Alyssa Rosenzweig f75bd2f09b Merge pull request #3922 from bylaws/structs
Windows: Pull in additional method and structure definitions from wine
2024-08-06 09:29:03 -04:00
Alyssa Rosenzweig a7424416d9 Merge pull request #3921 from bylaws/reloadf
Arm64Emitter: Reload STATE before SRA fill on ARM64EC
2024-08-06 09:28:23 -04:00
Alyssa Rosenzweig cadb0a2ddb Merge pull request #3923 from bylaws/except
ARM64EC: Improvements to exception flag handling
2024-08-06 09:27:50 -04:00
Alyssa Rosenzweig 2da819c0f3 Merge pull request #3919 from Sonicadvance1/remove_vestigial_vixl_usage
CodeEmitter: Removes vestigial vixl usage
2024-08-06 09:26:24 -04:00
Alyssa Rosenzweig e0c783de74 Merge pull request #3917 from Sonicadvance1/remove_static_vector
Config: Removes a static vector initializer
2024-08-06 09:26:03 -04:00
Billy Laws 6c003fcb9a Windows: Pull in more method/structure definitions from wine 2024-08-05 19:23:18 +00:00
Billy Laws c4faffc0e2 Windows: Add complete NTDLL export definitions
Generated from wine's ntdll.spec
2024-08-05 19:22:12 +00:00
Billy Laws cc2d21f411 WOW64: Resolve the wine unix call dispatcher at runtime 2024-08-05 19:22:12 +00:00
Billy Laws 59686a6c60 ARM64EC: Clear TF in the exception resumption context after a trap
Matches Windows behaviour.
2024-08-05 17:38:45 +00:00
Billy Laws 0ab864da17 ARM64EC: Reset the CPU area JIT state before handling exceptions
An exception in JIT code acts as a transition to ARM64EC code (in
NTDLL for exception handling etc) as such, much like ExitFunction,
InSimulation must be unset. InSyscallCallback is unset for robustness
against exception in the JIT itself.
2024-08-05 17:38:45 +00:00
Billy Laws 3c32271dd0 ARM64EC: Merge EFlags with the current JIT flags state on a ctx sync
Only NZCV and TF are passed through to BeginSimulation as the rest are
lost when converting to a native context and back on the ntdll side. To
prevent thread suspension from wiping out the rest of the flags, only
copy these specific flags into the current JIT EFlags state.
2024-08-05 17:38:45 +00:00
Billy Laws 507a95b817 ARM64EC: Map TF to PSTATE.SS when reconstructing a native context 2024-08-05 17:38:45 +00:00
Billy Laws 9bb9e954c2 ARM64EC: Spill EFlags when reconstructing state from in the JIT 2024-08-05 17:38:45 +00:00
Billy Laws 7f3582bb23 ARM64EC: Only clear the trap flag when handling an exception
Better matches Windows emulator behaviour.
2024-08-05 17:38:45 +00:00
Billy Laws efe15ce336 ARM64EC: Handle direct syscall instructions
Most syscalls on Windows are done by calling into their NTDLL thunks,
however some DRMs parse out their numbers from NTDLL and directly call
them. Support this by redirecting to their entry thunks in the FEX
syscall handler.
2024-08-05 17:35:32 +00:00
Billy Laws 21b0f35ef4 ARM64EC: Populate a LUT mapping NTDLL FFS exports to their native impls
To prevent FEX from redirecting to x86 code when NTDLL exports it calls
into are patched, a custom call checker will be used that checks this LUT
to redirect calls rather than the FFS itself.
2024-08-05 17:35:32 +00:00
Billy Laws ccf332d48e Arm64Emitter: Reload STATE before SRA fill on ARM64EC
While ARM64EC code cannot use x28, it can be cleared by the kernel
when performing syscalls etc so restore it from the TEB to be safe.
2024-08-05 17:31:01 +00:00
Ryan Houdek 802a32ce8a SpinWaitLock: Fixes missing newline in asm
This would cause the atomic load after the wfe to be dropped,
effectively returning stale data.
2024-08-04 06:57:05 -07:00
Ryan Houdek 70c02d5c58 ARM64Emitter: Removes unused vixl CPU object 2024-08-03 22:26:00 -07:00
Ryan Houdek 2e4fb47848 HostFeatures: Read VL ourselves
Instead of calling out to vixl
2024-08-03 22:26:00 -07:00
Ryan Houdek a4d5302369 Arm64: Adds Int helpers
One more vixl step removed.
2024-08-03 21:40:28 -07:00
Ryan Houdek 6ff3c90af3 CodeEmitter: Removes vestigial vixl usage
- IsImmLogical already existed in our CodeEmitter. We just forgot to
  allow nullptr arguments and to use it.
- Adds an equivalent IsImmAddSub helper and uses it

This gets us closer to removing vixl's global initializers from FEXCore.
2024-08-03 21:04:56 -07:00
Ryan Houdek c114279118 Config: Removes a static vector initializer
Saw this vector was getting initialized at runtime, sticking around, and
installing an atexit handler. This is completely unnecessary, just use
the OPT_BASE handler directly to walk the environment variable names.
2024-08-03 19:13:29 -07:00
Ryan Houdek 7ffd3e55d5 Merge pull request #3915 from bylaws/winbase
ARM64EC: Support the JIT API as is used by Windows
2024-08-02 10:56:46 -07:00
Ryan Houdek 201fe6ee23 Merge pull request #3909 from bylaws/ec-bitmap
Directly use the EC code bitmap for determining page arch
2024-08-02 10:55:43 -07:00
Ryan Houdek 83fedd6c8f Merge pull request #3912 from bylaws/addroverride
Don't apply the address-size flag to segment addresses
2024-08-01 18:35:06 -07:00
Ryan Houdek dedf4a93d0 Merge pull request #3913 from bylaws/logcommon
Commonise logging and fallback to a log file for debug output on Windows
2024-08-01 12:06:35 -07:00
Ryan Houdek 1f59f0e226 Merge pull request #3914 from bylaws/wincfg
Config: Search more locations for the config directory on Windows
2024-08-01 12:05:50 -07:00
Ryan Houdek c3c2b6115d Merge pull request #3910 from bylaws/f80
F80: Drop dependency on state stored in TLS
2024-08-01 12:03:41 -07:00
Billy Laws 115fbb5039 ARM64EC: Implement remaining notification callbacks 2024-08-01 12:06:25 +00:00
Billy Laws 27973d5637 ARM64EC: Fix the exception dispatcher stack layout 2024-08-01 12:06:24 +00:00
Billy Laws f121be649d ARM64EC: Match the Windows BT API function prototypes 2024-08-01 12:06:24 +00:00
Billy Laws af9bcb3efd ARM64EC: Avoid syncing uninitialized context members to the JIT state
Windows can sometimes pass in incomplete contexts to BeginSimulation. So
only sync the valid parts specified in ContextFlags.
2024-08-01 12:06:05 +00:00
Billy Laws 4877bb3f19 ARM64EC: Support directly issuing the NtContinue syscall
This is required for handling SMC with the ResetToConsistentState
arguments as used in Windows, as using the NTDLL exported NtContinue
would wipe out any reserved registers in the ARM64EC ABI.

For Windows the syscall numbers are somewhat stable, and the SVC
instruction can be called directly. Since wine doesn't handle that on
ARM64, hardcode the system call number and manually call into wine
dispatcher. Once wine gains proper syscall thunks, those can be
parsed to get the number and the hardcoding dropped.
2024-08-01 12:06:05 +00:00
Billy Laws e2583249c6 ARM64EC: Allocate the emulator stack ourselves
Actual Windows does not allocate it for us.
2024-08-01 12:06:05 +00:00
Billy Laws f48071dcbe ARM64EC: Switch to the emulator stack in BeginSimulation
Windows calls this function on the guest stack for some reason.
2024-08-01 12:06:05 +00:00
Billy Laws fa72be2ec5 Windows: Don't warn for unknown CPU features
This happens regularly as wine/games will scan all features from 0 to 64.
2024-08-01 12:06:05 +00:00
Billy Laws b7ff6dd9c6 Windows: Disable logging if SilentLog is enabled 2024-08-01 12:04:59 +00:00
Billy Laws cb6d60aa87 Config: Search more locations for the config directory on Windows 2024-08-01 11:48:40 +00:00
Billy Laws 80c9a43eef Windows: Fallback to a log file for debug output on Windows
OutputDebugString etc are exception based and thus don't really work for
FEX's needs as often times logs can happen in places where exceptions
cannot be thrown.
2024-08-01 11:45:07 +00:00
Billy Laws 05155778d4 Windows: Commonise logging code 2024-08-01 11:45:07 +00:00
Tony Wasserka 49b8dae189 Merge pull request #3908 from bylaws/pdb
CMake: Add option to build PDB debug info instead of DWARF
2024-08-01 10:01:04 +02:00
Ryan Houdek 3e59fc0a8c Merge pull request #3911 from bylaws/x80bug
x87StackOptimizationPass: Default initialise StackMemberInfo members
2024-07-31 22:57:16 -07:00
Ryan Houdek f98c010854 Merge pull request #3907 from bylaws/ec-mema
AllocatorHooks: Correct memory API usage on Windows
2024-07-31 17:48:13 -07:00
Ryan Houdek 10ee963f52 Merge pull request #3906 from bylaws/ec-sysbi
FEXCore: Add a generic spill/fill-all syscall ABI and use for Windows
2024-07-31 17:46:55 -07:00
Billy Laws af7462ee6a unittests: Add test using the address-override flag with segment addressing 2024-07-31 20:04:30 +01:00
Billy Laws be4777110c OpcodeDispatcher: Don't apply the address-size flag to segment addresses
The address-size flag only applies to the offset from the segment base,
rather than the segment address itself.
2024-07-31 18:14:32 +00:00
Billy Laws 696503680a F80: Drop dependency on state stored in TLS
Windows cannot support the implicit TLS as was used prior, so introduce
a state structure and pass it in to functions where necessary.
2024-07-31 18:51:42 +01:00
Billy Laws dd4d3bcf38 AllocatorHooks: Correct memory API usage on Windows
These issues end up being tolerated by wine but not actual windows.
2024-07-31 18:36:50 +01:00
Billy Laws f26bb6bf53 x87StackOptimizationPass: Default initialise StackMemberInfo members
Not doing so is UB.
2024-07-31 17:30:29 +00:00
Billy Laws 2c4fd79304 FEXCore: Add a generic spill/fill-all syscall ABI and use for Windows
Also drop the legacy hangover ABI as it has no users.
2024-07-31 17:25:59 +00:00
Billy Laws 910ec4aadd Dispatcher: Directly use the EC code bitmap for determining page arch
The prior approach using the L2 cache was flawed as it assumed L2
page entries had a 1-1 correspondence with actual pages. While the L2
cache could be extended to handle aliases with EC, this could lead to
thrashing etc. The cost of a lookup in the actual EC code bitmap is
cheap enough to perform every time considering the infrequency of calls
to ARM64EC code when compare to X86 L2 hits.
2024-07-31 17:24:50 +00:00
Billy Laws fb7275b3d8 Revert: "LookupCache: Track ARM64EC page state in the code cache"
This reverts the commit 526e3e654f.
2024-07-31 17:24:50 +00:00
Billy Laws 51b4bfc6a6 FEXCore: Move ARM64EC TEB offset constants to Arm64Emitter
These need to be used from outside the dispatcher, and there are already
similar defines for EC registers in the emitter header.
2024-07-31 17:24:50 +00:00
Billy Laws 9f7bc94f9d CMake: Add option to build PDB debug info instead of DWARF 2024-07-31 17:23:24 +00:00
Ryan Houdek 069e2ce62a Merge pull request #3905 from pmatos/FixWarn
Fix nasm warning in Rounding.asm
2024-07-31 08:06:12 -07:00
Paulo Matos 5bbbced1bd Fix nasm warning in Rounding.asm 2024-07-31 16:23:39 +02:00
Alyssa Rosenzweig 941fd9c6ea Merge pull request #3901 from pmatos/TopUsage
Reuse Top in ReconstructFSW_Helper
2024-07-31 08:01:05 -04:00
Paulo Matos 3332220d06 instcountci: Intersperse flag retrieval and FSW insertion 2024-07-31 12:05:17 +02:00
Paulo Matos 5933a59c09 Intersperse flag retrieval and FSW insertion 2024-07-31 12:04:54 +02:00
Paulo Matos aee8c9def2 instcountci: Reuse Top in ReconstructFSW_Helper 2024-07-31 11:57:19 +02:00
Paulo Matos c2136272bf Reuse Top in ReconstructFSW_Helper
This is a non functional. Instead of fetching top again, we use the one
obtained through the fast path calculation.
2024-07-31 11:56:14 +02:00
Ryan Houdek dd26b0c879 Merge pull request #3903 from Sonicadvance1/v6.10_syscalls
Syscalls: Updates for v6.10
2024-07-31 02:22:26 -07:00
Ryan Houdek 3d9114b74b Merge pull request #3902 from Sonicadvance1/man_page_fix
man: Fixes newline issue with strenum
2024-07-31 02:21:47 -07:00
Ryan Houdek 3f3e937967 man: Fixes newline issue with strenum
In the environment section this was causing the next environment
variable line to be merged with the strenum options

Also makes it so strenum options doesn't have a spurious comma at the
end of the list.
2024-07-31 02:02:05 -07:00
Ryan Houdek 4beb29141b Scripts: Workaround deprecated parse_version
Different approach from #3579

Instead of completely ddropping the deprecated path, support the new
path and the old path using python try-except import exceptions.

This allows us to continue using old packages in CI, while supporting
the future API once pkg_resources gets deprecated and removed. Best of
both worlds.
2024-07-31 01:57:31 -07:00
Paulo Matos 9a8e7eaace instcountci: Add instcountci for fnstsw from fast path 2024-07-31 09:59:32 +02:00
Paulo Matos 4227e012aa Add instcountci for fnstsw from fast path 2024-07-31 09:59:28 +02:00
Ryan Houdek 40bbcb7061 Syscalls: Updates for v6.10
Only mseal was added and can be a simple passthrough for us. Which is
nice.
2024-07-30 19:08:27 -07:00
Ryan Houdek 67663b812e Linux: Update syscall defines for v6.10 2024-07-30 19:01:52 -07:00
Ryan Houdek d24d0a95a0 Merge pull request #3894 from pmatos/RoundingModeTests
ASM Tests: X87 Rounding modes
2024-07-29 23:26:52 -07:00
Ryan Houdek 87fbcf754d Vector: Optimize pblendw
Using a brute force solver to add in more optimized code paths

- Adds 12 single VInsElement implementations
- Adds 4 two IR operation implementations

Not adding any of the two or three IR operation implementations that use
VInsElement because SRA interacts badly and becomes worse than the VTBX
implementation.
2024-07-27 19:25:51 -07:00
Ryan Houdek 403fd62b34 Merge pull request #3890 from Sonicadvance1/refactor_frontend_threadmanager
FEXCore: Removes ThreadManager
2024-07-26 13:27:43 -07:00
Ryan Houdek d92b6a9ac4 Merge pull request #3898 from alyssarosenzweig/ir/creative-refs
OpcodeDispatcher/X87: use less creative Refs
2024-07-26 13:26:43 -07:00
Ryan Houdek c2092bfed0 Merge pull request #3893 from pmatos/FNINITFix
Fix call to FNINITF64 and refactor
2024-07-26 13:25:49 -07:00
Ryan Houdek a6cf7fa508 Merge pull request #3896 from pmatos/CTestSkip
Test running scripts tell ctest of skipped tests
2024-07-26 13:25:28 -07:00
Alyssa Rosenzweig 5ff09f5091 OpcodeDispatcher/X87: use less creative Refs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-26 14:30:42 -04:00
Paulo Matos 9af7ee6bd2 instcountci: Fix call to FNINITF64 and refactor 2024-07-26 14:56:10 +02:00
Paulo Matos d1e36f264f Fix call to FNINITF64 and refactor 2024-07-26 14:56:06 +02:00
Paulo Matos b1ec50c7c2 Test running scripts tell ctest of skipped tests
CMake sets 125 as the skipped test exit code that the scripts use.
2024-07-26 14:04:54 +02:00
Mai 93eead243f Merge pull request #3864 from Sonicadvance1/threads_atexit_remove
Threads: Setup the stack tracker to not need global initialization
2024-07-26 06:38:29 -04:00
Paulo Matos ceac38a6ac ASM Tests: X87 Rounding modes 2024-07-26 10:07:58 +02:00
Ryan Houdek 3b2e657fd4 FEXCore: Removes ThreadManager
This has been leaked state to FEXCore for quite a while. FEXCore never
actually needed this information, moves the bits to the frontend that
are necessary.

Minor behaviour change that `RunUntilExit` now just assumes the primary
thread is using it. This behaviour is on the chopping block to get
removed next anyway.
2024-07-25 14:54:10 -07:00
Ryan Houdek 380ba0a014 Merge pull request #3889 from Sonicadvance1/refactor_frontend_exithandler
FEXCore: Refactor ExitHandler slightly
2024-07-25 14:53:25 -07:00
Mai 1fe497d1dd Merge pull request #3891 from Sonicadvance1/remove_cpubackendfeatures
FEXCore: Removes CPUBackendFeatures
2024-07-24 20:53:10 -04:00
Ryan Houdek 7816b150d0 FEXCore: Removes CPUBackendFeatures
We were only ever hardcoding true for TBL2 and Flags now. Get rid of it.
2024-07-24 17:19:30 -07:00
Ryan Houdek ce8bc9d25c FEXCore: Refactor ExitHandler slightly
Instead of passing the TID back to the exit handler, just pass the whole
thread object. This will allow some cleanups with the frontend thread
tracking soon

NFC
2024-07-24 14:39:56 -07:00
Ryan Houdek 4634688aca InstcountCI: Update for AVX128 blends 2024-07-23 19:24:19 -07:00
Ryan Houdek dd3e3ed189 unittests/ASM: Implements a vpblendw test
Runs through all immediate encodings for vpblendw and crcs the results
to ensure correct behaviour. This was just a concern because of the typo
in documentation. But it is also good to have.
2024-07-23 19:24:19 -07:00
Ryan Houdek f8ef6feff9 AVX128: Optimize blends
Optimizes the AVX128 blends by reusing the prior SSE4.1 implementation.
Only difference is the destination register isn't reused as a source
register.

One confusing thing is that Felix Cloutier's documentation has a typo on
the 256-bit VPBLENDW instruction where it had the top 128-bit lane
reusing the destination instead of sources. So I wrote a unittest to
ensure correctness.

Fixes #3796
2024-07-23 19:24:19 -07:00
Ryan Houdek 9201ac5a6b Merge pull request #3882 from Sonicadvance1/scalar_afp_fma
AVX128: Implement support for scalar FMA with AFP
2024-07-22 13:19:59 -07:00
Ryan Houdek 8ebf049fb9 InstcountCI: Update for Scalar FMA with AFP 2024-07-22 12:58:20 -07:00
Ryan Houdek 3c5b59d985 AVX128: Implement support for scalar FMA with AFP
Now that I have AFP supporting hardware I felt better implementing this
since I can run unit tests.

Fixes #3793
2024-07-22 12:58:19 -07:00
Ryan Houdek 6b91e0cb0e Merge pull request #3887 from alyssarosenzweig/ir/prefix
json_ir_generator: stop prefixing arguments
2024-07-22 12:57:09 -07:00
Alyssa Rosenzweig 587b924de9 json_ir_generator: stop prefixing arguments
stop prefixing the arguments when we generate allocate ops (in particular), this
is more convenient and simpler. in exchange we need to prefix Op to avoid a
collision on fcmpscalarinsert which has an argument named Op, but that's a local
change at least.

came up when experimenting with new IR, but I think this is probably a win by
itself.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-22 13:50:21 -04:00
Tony Wasserka d507f4c9b1 Merge pull request #3547 from pmatos/wip_x87_stack
x87 Stack Optimization
2024-07-22 14:42:02 +02:00
Paulo Matos 39bc2a82c1 instcountci: X87 Pass and refactoring 2024-07-22 08:50:01 +02:00
Paulo Matos 774325dcf2 Tests: X87 Refactoring and Pass 2024-07-22 08:44:45 +02:00
Paulo Matos a1378f94ce X87 Code Refactoring and Optimization Pass 2024-07-22 08:44:45 +02:00
Ryan Houdek 77ec950ff2 Merge pull request #3885 from alyssarosenzweig/opt/zero-flag
Optimize zero x87 flags
2024-07-21 13:06:25 -07:00
Alyssa Rosenzweig 592d6cc43f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:50:10 -04:00
Alyssa Rosenzweig 610caf8529 ConstProp: treat StoreContext as zeroable
todo: FPR equivalent.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:09 -04:00
Alyssa Rosenzweig d20b46e46f IR: drop LoadFlag/StoreFlag ops
pointless, we can just load/store the context now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:09 -04:00
Alyssa Rosenzweig 4094aa1b9a DeadStoreElimination: drop flag handling
now that we do everything via NZCV, this is mostly vestigial. DF/x87 flags are
sufficiently rare to be "don't care"s here, and we don't even have multiblock
enabled yet!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-21 15:49:08 -04:00
Ryan Houdek f8c6baae97 Merge pull request #3883 from Sonicadvance1/implement_daz
Arm64: Implements support for DAZ using AFP.FIZ
2024-07-21 10:03:34 -07:00
Ryan Houdek 5c9bb6594c Merge pull request #3884 from Sonicadvance1/remove_vex_telem
Telemetry: Remove VEX flag
2024-07-20 19:44:28 -07:00
Ryan Houdek 56df57e980 InstcountCI: Update 2024-07-20 17:26:27 -07:00
Ryan Houdek f7b4d25803 Telemetry: Remove VEX flag
This is no longer necessary and it also no longer provides us any useful
information. Since we expose the AVX CPUID flag, basically everything
uses VEX encoding now, so it is basically always set.
2024-07-20 17:24:00 -07:00
Ryan Houdek 4fffe68f81 InstcountCI: Update 2024-07-20 15:57:01 -07:00
Ryan Houdek 95b15d788b Arm64: Fix filling static registers
Some locations could end up with SRA registers that only spilled one
register.
Allow passing in temporaries from the call site.
Fixes rpid and syscalls asserting.
2024-07-20 15:57:01 -07:00
Ryan Houdek ae9312bdab unittests: Implements a DAZ test
Specifically does a vector add with and without DAZ enabled and ensures
the value is different when the source values contain a denormal.
2024-07-20 15:34:54 -07:00
Ryan Houdek b78da2e5ad Arm64: Implements support for DAZ using AFP.FIZ
When AFP is supported then we can actually support DAZ. This might also
fix the audio corruption in Animal Well but I can't test it until Steam
is running on Oryon. Requires a bit of plumbing for MXCSR which we were
hacking around before but now we actually want to store the value.

Fixes #3856
2024-07-20 15:34:54 -07:00
Ryan Houdek 54fc8cb0bd TestHarnessRunner: Support querying AFP for features
Also fixes desync of flags
2024-07-20 15:34:54 -07:00
Ryan Houdek 228009c283 Merge pull request #3881 from bylaws/race
FixedSizePooledAllocation: Fix a race when unclaiming disowned buffers
2024-07-20 12:35:35 -07:00
Billy Laws 1f878ce4cd FixedSizePooledAllocation: Fix a race when unclaiming disowned buffers
A disowned buffer could be unclaimed or claimed by a different thread in
the time between the !IsFree check and locking the allocation mutex.
Fix this and prevent such errors in the future by always checking
ownership with the allocator locked before attempting to unclaim
buffers.
2024-07-20 00:12:52 +00:00
Alyssa Rosenzweig e4b7a65a49 Merge pull request #3880 from pmatos/InstCountMemcpy
Add x87 memcpy instcountci tests
2024-07-19 08:53:23 -04:00
Paulo Matos c77a707dbe Add x87 memcpy instcountci tests 2024-07-19 09:09:34 +02:00
Ryan Houdek f81fc4e4f0 Merge pull request #3866 from Sonicadvance1/ArgumentLoader_atexit_remove
ArgumentLoader: Removes static fextl::vector usage
2024-07-18 13:12:03 -07:00
Ryan Houdek d385e496d3 Merge pull request #3879 from Sonicadvance1/cpuid_leafs
EmulatedFiles: Adds a few leaf CPUID flags
2024-07-18 13:10:21 -07:00
Mai f8c4c543e3 Merge pull request #3871 from Sonicadvance1/improve_vpshufd_vpermilps
AVX128: Improve VPERMILPS/PD and VPSHUFD
2024-07-18 15:58:47 -04:00
Ryan Houdek d2f903ae55 EmulatedFiles: Adds a few leaf CPUID flags
We support leaf functions now, so add the few that were calling for it.
We will be gaining support for the xsave ones relatively soon, so its
good to have them supported.

Also deletes a couple of cdt/cqm things that aren't exposed and we won't
be supporting.
2024-07-18 07:02:49 -07:00
Ryan Houdek d1249ec5cf Merge pull request #3878 from neobrain/refactor_fix_format_oops
EmulatedFiles: Fix bad formatting
2024-07-18 06:44:44 -07:00
Tony Wasserka bf9a6d763c EmulatedFiles: Fix bad formatting 2024-07-18 15:06:57 +02:00
Ryan Houdek 0b829d2c46 unittests: Adds a test for full pshufd imm coverage 2024-07-18 04:13:03 -07:00
Ryan Houdek bddb533fa0 InstcountCI: Add some more of the cases 2024-07-18 04:13:03 -07:00
Ryan Houdek 1c35eeffeb Vector: Optimize PSHUFD with brute force search
With a brute force search of methods between 1-3 instructions we cover a
lot more cases more optimally.

There's definitely still more cases (and probably some that can reduce
from 3 instruction to 2), but covering 44 cases is a pretty good margin
already.
2024-07-18 04:10:58 -07:00
Ryan Houdek c7254e31ed InstcountCI: Update for VPERM/VPSHUFD improvements 2024-07-18 04:10:58 -07:00
Ryan Houdek b0bd8a62a2 AVX128: Improve VPERMILPS/PD and VPSHUFD
VPSHUFD and VPERMILPS are aliases of each other.

Reuses the implementation path from the PSHUFD implementation which has
a few swizzles and then a table lookup.

VPERMILPD is a very simple swizzle per 128-bit lane.

Fixes #3797
Fixes #3784
2024-07-18 04:10:58 -07:00
Ryan Houdek 17a55fbb39 Merge pull request #3876 from alyssarosenzweig/json/x87
Autogenerate LoweredX87() query, misc json_ir_generator cleanup in the area
2024-07-18 04:10:21 -07:00
Alyssa Rosenzweig dedec83881 json_ir_generator: autoderive array names
these are purely internal.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:32:51 -04:00
Alyssa Rosenzweig 9fd5c73633 json_ir_generator: generate IsLoweredX87 helper
X87 pass will use this query.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:31:46 -04:00
Alyssa Rosenzweig bdb890a8b0 json_ir_generator: rename X87 -> LoweredX87
to reflect its actual meaning

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:31:14 -04:00
Alyssa Rosenzweig 5043d09771 json_ir_generator: use textwrap.dedent, f-string
for size

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-17 15:28:59 -04:00
Ryan Houdek da51169ba9 Merge pull request #3875 from alyssarosenzweig/ir/gethostflag
IR: garbage collect premature F80Cmp optimizations
2024-07-17 03:05:48 -07:00
Ryan Houdek f72cee480f Merge pull request #3874 from alyssarosenzweig/opt/reconstructftw
X87: save uop in ReconstructFTW
2024-07-17 03:05:37 -07:00
Alyssa Rosenzweig 7546160811 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Alyssa Rosenzweig e7d5a01c5f IR: remove F80Cmp flags
nothing is optimizing around this, it's just adding pointless complexity. if we
want to actually optimize F80Cmp, the right way would be to lift the
implementation into the OpcodeDispatcher or JIT. it wouldn't be terribly
difficult. This kludge doesn't get us closer there.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:53:58 -04:00
Alyssa Rosenzweig 0c3a8d0bc8 IR: remove GetHostFlag
it doesn't get host flags, it's just an extra Bfe used in x87. pointless and
confusing!

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 14:44:34 -04:00
Alyssa Rosenzweig 19e58cac62 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 13:54:28 -04:00
Alyssa Rosenzweig c4ba7eee87 X87: save uop in ReconstructFTW
noticed while reviewing Paulo's work

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 13:54:09 -04:00
Ryan Houdek 09c4a5594a Merge pull request #3870 from Sonicadvance1/enable_more_tests
github: Vixl simulator enable more asm tests
2024-07-16 07:23:22 -07:00
Alyssa Rosenzweig d204155661 Merge pull request #3872 from pmatos/X87AutoMarking
X87 Stack Ops Auto-marking
2024-07-16 09:45:42 -04:00
Tony Wasserka 924b8c10a9 Merge pull request #3873 from pmatos/UnusedFunction
Remove unused function MmapOverride
2024-07-16 12:30:04 +02:00
Paulo Matos 9017cd14c8 Remove unused function MmapOverride 2024-07-16 11:16:07 +02:00
Paulo Matos 8d89adef2e Add IR stack operations
These IR operations deal implicitly with the x87 stack and are removed
by the x87 stack optimization pass.
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 6615b55c12 json_ir_generator: call RecordX87Use when generating ops
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 66865dd177 json_ir_generator: alias X87 to !JITDispatch
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 1e709d1150 OpcodeDispatcher: add RecordX87 helper
calls will be generated.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Alyssa Rosenzweig 476ee0cd7d IR: track whether x87 is used in header
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-16 09:07:35 +02:00
Ryan Houdek 6df51a57b3 Merge pull request #3868 from bylaws/arm64ec-oldnew
ARM64EC frontend
2024-07-15 09:54:56 -07:00
Ryan Houdek e65545a537 Merge pull request #3867 from pmatos/NoDisableTests
Remove Disabled_Tests file
2024-07-15 09:54:08 -07:00
Ryan Houdek b8e864ffdf Merge pull request #3865 from Sonicadvance1/telemetry_atexit
Telemetry: Change how visibility of telemetry values work
2024-07-15 09:53:37 -07:00
Paulo Matos ed87c01470 Simplify Disabled_Tests and remove pr57275 from failures
Disabled_Tests was mostly a copy of Known_Failures. Leave only the race
condition on SIGPROF test (mcount_pic.c).

Also remove pr57275.c from known failures. It passes now that we support
 AVX.

Also if a test is disabled, just skip it
2024-07-13 07:08:02 +02:00
Ryan Houdek 6c21b86a8f github: Vixl simulator enable more asm tests
We were only running SVE256 and SVE128 with AVX disabled.

Enable asm tests with SVE256, SVE128, and ASIMD, all running with AVX
enabled to hit all the tests.
2024-07-12 20:39:01 -07:00
Ryan Houdek d79b7fcc49 Merge pull request #3808 from alyssarosenzweig/rclse/3
Try to delete RCLSE again
2024-07-12 20:38:06 -07:00
Ryan Houdek b9a6caea8d Merge pull request #3844 from Sonicadvance1/fix_vmovq
AVX128: Fixes vmovq loading too much data
2024-07-12 17:07:32 -07:00
Ryan Houdek 9688f5e17d Merge pull request #3863 from Sonicadvance1/remove_static_ioctl_handlers
Ioctl32: Removes static fextl::vector in ioctlemulation
2024-07-12 17:06:16 -07:00
Billy Laws f6f8d26426 Update jemalloc submodule 2024-07-12 19:24:13 +00:00
Billy Laws dba0a1d09e ARM64EC: Initialize x86 control registers on thread start 2024-07-12 18:51:31 +00:00
Billy Laws af3145674e ARM64EC: Fixup exception information for faulting x86 instructions
FEX emulates faulting instructions (e.g. ud2 or int 2d) by jumping to
the dispatcher and filling out a structure with fault details in the
thread context. Parse this out into a windows exception record structure
so the correct fault information can be seen by the guest.
2024-07-12 18:51:31 +00:00
Billy Laws 3c19e634b3 ARM64EC: Rethrow exceptions from within the JIT
As the exception dispatcher is initially invoked on the emulator stack,
control needs to be transferred to the dispatcher on the guest stack
after recovering the x86 RSP to allow for invoking x86 exception
handlers.
2024-07-12 18:41:20 +00:00
Billy Laws f964a5187e ARM64EC: Implement BeginSimulation
This is used by the kernel (or UNIX side of ntdll in wine) to jump into
x86 code with the given context as is necessary when e.g. returning from
an exception.
2024-07-12 18:41:13 +00:00
Billy Laws 8e0fdfc325 ARM64EC: Add a helper to lookup the redirected address of an export
FEX is unable to deal with reentrant compilation of any x64 hotpatches
so they need to be ignored by bypassing FFSs and calling directly into
the native target.
2024-07-12 18:41:08 +00:00
Billy Laws 839f9ecd3b Windows: Add ARM64EC image structures 2024-07-12 18:41:06 +00:00
Billy Laws 95fc69b628 ARM64EC: Handle SMC 2024-07-12 18:41:02 +00:00
Billy Laws b9da95838a ARM64EC: Handle unaligned atomic accesses 2024-07-12 18:40:43 +00:00
Billy Laws 1059279d5d ARM64EC: Handle calls into ARM64EC code with an 8-byte-aligned SP
ARM64 requires that SP is always 16-byte aligned for memory accesses,
but ARM64EC shares the SP between x64 code and ARM64 code, the former
of which doesn't enforce such a restriction. This causes crashes in
programs such as HITMAN 3 that don't correctly follow the Windows ABI
and call into system library functions with SP only 8-byte-aligned.
Fixup stack alignment in such cases by leaving the 8-byte return
address on the stack and returning to a lone 'ret' instruction instead.
2024-07-12 18:30:04 +00:00
Billy Laws 5dc85307a6 Windows: Introduce an initial ARM64EC frontend
This allows for running x64 applications under wine without having to run all
of wine under FEX. The JIT is invoked when ARM64EC code performs an indirect
branch to x64 code, and left whenever the x64 code calls into ARM64EC
code.
2024-07-12 18:07:50 +00:00
Billy Laws 549e06aade CMake: Enable assembly source file support 2024-07-12 18:01:22 +00:00
Billy Laws 3b189f6d7d WOW64: Install into lib
This convention is used by most other projects.
2024-07-12 18:01:22 +00:00
Ryan Houdek a5d3692b53 ArgumentLoader: Removes static fextl::vector usage
Removes a global initializer and atexit registration

Ownership of this data has always been the frontend and the config
system, we just used these static vectors as a side-channel.
2024-07-12 04:48:22 -07:00
Ryan Houdek 97a68cb643 Telemetry: Change how visibility of telemetry values work
Removes global initializer for telemetry values since their address is
visible and PIC relative code loading handles the address fetching for
us.
2024-07-12 03:18:23 -07:00
Ryan Houdek d1b5dfd4b1 Threads: Setup the stack tracker to not need global initialization
Also removes the atexit handler installation
This now gets tracked by an object owned by FEXLoader (and shared with
the pthreads interface)
2024-07-12 03:00:20 -07:00
Ryan Houdek 6cdaea680d Ioctl32: Removes static fextl::vector in ioctlemulation
Removes a global static initializer for the vector and its atexit
handler.

This handler array can be consteval similar to the x86 tables so it can
be generated entirely at compile time.
2024-07-12 02:05:32 -07:00
Ryan Houdek 870e395ac4 Merge pull request #3862 from Sonicadvance1/remove_atexit_logman
LogManager: Removes fextl::vector usage
2024-07-12 02:05:02 -07:00
Ryan Houdek 04592f82f5 Merge pull request #3861 from Sonicadvance1/remove_atexit_vdso
VDSO: Stop using a vector for a static
2024-07-12 02:04:25 -07:00
Ryan Houdek 19e849283f Merge pull request #3860 from Sonicadvance1/force_noinline
OpcodeDispatcher: Force noinline for the function call in the Bind helper
2024-07-12 00:14:04 -07:00
Ryan Houdek b6e1469cd4 Merge pull request #3847 from pmatos/CoverageSupport
Enable coverage configuration for FEX
2024-07-12 00:02:08 -07:00
Ryan Houdek 5ef0db994d VDSO: Stop using a vector for a static
This causes a global initializer that registers an atexit handler.

Be smarter, use an std::array and pass its data around using a span
instead.

Removes the global initializer and removes the atexit installation
2024-07-11 23:53:57 -07:00
Ryan Houdek b523407a3e LogManager: Removes fextl::vector usage
We never use more than one logging method at a time so this was
overengineered for what it is doing.

Instead only allow one handler for messages and throw messages each
which just is a pointer.

Removes a global initializer and an atexit handler being installed
2024-07-11 22:51:56 -07:00
Ryan Houdek 8021dc10a1 OpcodeDispatcher: Force noinline for the function call in the Bind helper
Clang was inlining a few of the functions it was calling. So force it
never to inline since this is supports to be a little shim trampoline
only.
2024-07-11 19:00:42 -07:00
Ryan Houdek 7e8d734e43 AVX256: Initial fixes just to get my unittest working
This is the initial split to decouple AVX256 composed operations from
their MMX/SSE counterparts. This is to work around the subtle
differences with AVX/SSE zext/insert behaviour.
2024-07-11 18:43:31 -07:00
Ryan Houdek 3d90d1ab4f InstcountCI: Update for vmovq fix 2024-07-11 18:34:06 -07:00
Ryan Houdek 3c7318d7c8 AVX128: Fixes vmovq loading too much data
This was doing a 128-bit load from memory and then a 64-bit zero extend
which looked like a spurious move but it was trying to match the
behaviour of vmovq where it needed the zero extend.

Also adds a unit test to ensure that we aren't loading too much data by
loading right up against a page boundary.

Fixes #3787
2024-07-11 18:34:05 -07:00
Ryan Houdek fc0b233046 Merge pull request #3859 from neobrain/refactor_opdispatch_templates
OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility
2024-07-11 18:18:23 -07:00
Mai e25918d846 Merge pull request #3858 from Sonicadvance1/implement_nt_load
Implement support for SSE4.1/AVX NT loads
2024-07-11 14:22:41 -04:00
Alyssa Rosenzweig d78b0ea435 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 3a334c4585 Reapply "IR: drop RCLSE"
This reverts commit 78aee4d96e.
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 8dae4bcd44 OpcodeDispatcher: drop stale comment
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Alyssa Rosenzweig 294f10fdd0 OpcodeDispatcher: reg cache mmx
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-11 13:21:14 -04:00
Tony Wasserka b9829ed316 OpcodeDispatcher: Replace even more hand-written wrapper templates 2024-07-11 16:19:15 +02:00
Tony Wasserka 4ccec17676 OpcodeDispatcher: Replace more hand-written wrapper templates 2024-07-11 16:19:15 +02:00
Tony Wasserka f45082043b OpcodeDispatcher: Replace hand-written wrapper templates with a generic utility 2024-07-11 16:19:14 +02:00
Tony Wasserka 3222f13dde Fix comment formatting 2024-07-11 16:19:14 +02:00
Mai b282620a48 Merge pull request #3857 from Sonicadvance1/sve_bitperm
Arm64: Implement support for SVE bitperm
2024-07-11 05:05:41 -04:00
Ryan Houdek 3ff1ff8f74 InstcountCI: Update for svebitperm 2024-07-11 01:46:35 -07:00
Ryan Houdek e24b01b6cb Arm64: Implement support for SVE bitperm 2024-07-11 01:46:35 -07:00
Tony Wasserka 9a8694c2f3 Merge pull request #3853 from neobrain/refactor_warn_fixes
Fix all the warnings
2024-07-11 10:12:41 +02:00
Tony Wasserka 070a9148aa Merge pull request #3852 from neobrain/refactor_opdispatch_codesize
OpcodeDispatcher: Avoid template monomorphization to reduce FEXLoader binary size
2024-07-11 09:58:49 +02:00
Tony Wasserka f19fe3b6f3 Fix warning about an expression with side effects being passed to __builtin_assume
LOGMAN_THROW_AA_FMT has no benefit over LOGMAN_THROW_A_FMT here, so just use
the latter.
2024-07-11 09:54:31 +02:00
Tony Wasserka 8d2b15665d Fix unused-variable warnings 2024-07-11 09:54:30 +02:00
Tony Wasserka 4dec8f22f8 Fix packed-non-pod warnings 2024-07-11 09:54:30 +02:00
Tony Wasserka a39b3aca78 Fix invalid-offsetof warnings due to JsonAllocator not being standard layout
Inheritance can be used here instead, which allows the JsonAllocator to be
reconstructed using a downcast.
2024-07-11 09:54:30 +02:00
Tony Wasserka 5dc4ab062d Fix invalid-offsetof warnings due to InternalThreadState not being standard layout
See https://github.com/llvm/llvm-project/issues/53021 for more information
about unique_ptr turning non-standard-layout.
2024-07-11 09:54:30 +02:00
Ryan Houdek 31f82c1d96 InstcountCI: Update for SVE NT load support 2024-07-10 23:07:58 -07:00
Ryan Houdek 548fd9daf8 OpcodeDispatcher: Implement support for SSE4.1 NT load 2024-07-10 23:07:37 -07:00
Ryan Houdek f831f5a0e1 AVX128: Implement support for NT Load 2024-07-10 23:07:14 -07:00
Ryan Houdek 4c21aa2604 Arm64: Implement support for NT Loads with ASIMD fallback 2024-07-10 23:06:46 -07:00
Ryan Houdek c9efb75714 CodeEmitter: Implement support for SVE NT loads 2024-07-10 23:06:19 -07:00
Ryan Houdek 5e56bdc0fd InstcountCI: Add support for SVE bitperm 2024-07-10 21:48:37 -07:00
Ryan Houdek 3554d5c2f7 HostFeatures: Check for SVE bit permute extension 2024-07-10 21:45:07 -07:00
Mai 5fe405e1fb Merge pull request #3855 from neobrain/fix_aotir_uniqueptr
AOTIR: Change std::unique_ptr to fextl::unique_ptr
2024-07-10 17:04:12 -04:00
Tony Wasserka 8381d44bbd Merge pull request #3854 from neobrain/fix_default_delete
fextl: Properly handle nullptr arguments in fextl::default_delete
2024-07-10 23:00:52 +02:00
Tony Wasserka 56bb3744a5 AOTIR: Change std::unique_ptr to fextl::unique_ptr 2024-07-10 19:34:24 +02:00
Tony Wasserka 470b435afd fextl: Properly handle nullptr arguments in fextl::default_delete
This reflects behavior of std::default_delete.
2024-07-10 19:17:50 +02:00
Alyssa Rosenzweig a4f8bbff02 OpcodeDispatcher: reg cache avx high
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig cf5ab05b90 OpcodeDispatcher: reg cache AbridgedFTW
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 3a2ce240f9 OpcodeDispatcher: reg cache DF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 1f01dd53f7 OpcodeDispatcher: reg cache fprs
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 72d41d70b6 OpcodeDispatcher: introduce GPR-only reg cache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 42b5b1f64c Core: partially flush register cache per instruction
This will mitigate problems later.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 2949bc211d OpcodeDispatcher: thunk through FlushRegisterCache
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:36:18 -04:00
Alyssa Rosenzweig 5e0952159d unittests: add test for a MMX register cache bug
this failed on an earlier version of the register cache.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-10 11:34:24 -04:00
Tony Wasserka 441187470e OpcodeDispatcher: Avoid monomorphization of some AVX functions 2024-07-10 17:01:30 +02:00
Tony Wasserka 59fd13cc2f OpcodeDispatcher: Avoid monomorphization of even more functions 2024-07-10 17:01:30 +02:00
Tony Wasserka c9e7bfdf16 OpcodeDispatcher: Avoid monomorphization of more functions 2024-07-10 17:01:30 +02:00
Tony Wasserka 2d700c381e OpcodeDispatcher: Avoid monomorphization of large functions 2024-07-10 17:01:30 +02:00
Ryan Houdek 72d6c8ebd6 Merge pull request #3820 from alyssarosenzweig/ir/drop-deferred
Drop deferred flag infrastructure
2024-07-09 17:06:25 -07:00
Ryan Houdek 991c6941c1 Merge pull request #3849 from alyssarosenzweig/ir/drop-parser-2
Scripts: drop remnant of IR parser
2024-07-09 16:48:36 -07:00
Alyssa Rosenzweig f974696e34 Scripts: drop remnant of IR parser
unused.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-09 16:08:38 -04:00
Alyssa Rosenzweig 3ef9ea94e5 Merge pull request #3848 from pmatos/FTSTX87Tests
Tests for X87 FTST
2024-07-09 09:10:29 -04:00
Paulo Matos 381ce23fd7 Tests for X87 FTST 2024-07-09 13:36:16 +02:00
Mai af6a0be832 Merge pull request #3842 from Sonicadvance1/fix_f64_to_i32
VCVT{T,}PD2DQ fixes and optimization
2024-07-09 03:49:31 -04:00
Ryan Houdek 287fe5beac InstcountCI: Update 2024-07-09 00:38:48 -07:00
Ryan Houdek b9c214e6e8 OpcodeDispatcher: Use new IR op for vcvt{t,}pd2dq
Also fixes a bug where it was failing to zero the upper bits of the
destination register in the AVX128 implementation. Which the updated
unit tests now check against.

Fixes a minor precision issue that was reported in #2995. We still don't
return correct values for overflow. x86 always returns maximum negative
int32_t on overflow, ARM will return maximum negative or positive
depending on sign of the double.
2024-07-09 00:38:47 -07:00
Ryan Houdek d3d76aa8ce IR: Adds new F64 -> I32 operation that changes behaviour depending on SVE
SVE added the ability to do F64 -> I32 conversions directly without an
fcvtn inbetween. So maybe sure to support them.
2024-07-09 00:38:47 -07:00
Ryan Houdek 3bea08da5f Merge pull request #3843 from Sonicadvance1/remove_half_moves_fma3
Arm64: Remove one move if possible in FMA operations
2024-07-09 00:25:07 -07:00
Paulo Matos 3d5cacbdc3 Enable coverage configuration for FEX 2024-07-09 08:03:41 +02:00
Mai 7ccb252069 Merge pull request #3837 from Sonicadvance1/optimize_sve_vpgatherdq
AVX128: Extends 32-bit indexes path for 128-bit operations
2024-07-08 22:01:02 -04:00
Ryan Houdek 31547462bb InstcountCI: Update for final SVE AVX128 improvements. 2024-07-08 18:44:07 -07:00
Ryan Houdek b3a7a973a1 AVX128: Extends 32-bit indexes path for 128-bit operations
The codepath from #3826 was only targeting 256-bit sized operations.
This missed the vpgatherdq/vgatherdpd 128-bit operations. By extending
the codepath to understand 128-bit operations, we now hit these
instruction variants.

With this PR, we now have SVE128 codepaths that handle ALL variants of
x86 gather instructions! There are zero ASIMD fallbacks used in this
case!

Of course depending on the instruction, the performance still leaves a
lot to be desired, and there is no way to emulate x86 TSO behaviour
without an ASIMD fallback, which we will likely need to add as a
fallback at some point.

Based on #3836 until that is merged.
2024-07-08 18:44:07 -07:00
Mai 22b26696ba Merge pull request #3836 from Sonicadvance1/optimize_sve_vpgatherdd
AVX128: Optimize the vpgatherdd/vgatherdps cases that would fall back to ASIMD
2024-07-08 21:43:36 -04:00
Ryan Houdek 495241f8ca InstcountCI: Update for wide gather vpgatherdd SVE usage 2024-07-08 18:12:28 -07:00
Ryan Houdek 4afbfcae17 AVX128: Optimize the vpgatherdd/vgatherdps cases that would fall back to ASIMD
With the introduction of the wide gathers in #3828 this has opened new
avenues for optimizing these cases that would typically fall back to
ASIMD. In the cases that 32-bit SVE scaling doesn't fit, we can instead
sign extend the elements in to double-width address registers.

This then feeds naturally in to the SVE path even though we end up
needing to allocate 512-bits worth of address registers. This ends up
being significantly better than the ASIMD path still.

Relies on #3828 to be merged first
Fixes #3829
2024-07-08 18:12:28 -07:00
Mai 3627de4cbc Merge pull request #3828 from Sonicadvance1/optimize_wide_gathers
AVX128: Optimize QPS/QD variant of gather loads!
2024-07-08 21:11:36 -04:00
Ryan Houdek 007c07e612 InstcountCI: Update for wide gathers 2024-07-08 17:19:18 -07:00
Ryan Houdek ec7c8fd922 AVX128: Optimize QPS/QD variant of gather loads!
SVE has a special version of their gather instruction that gets similar
behaviour to x86's VGATHERQPS/VPGATHERQD instructions.

The quirk of these instructions that the previous SVE implementation
didn't handle and required ASIMD fallback, was that most gather
instructions require the data element size and address element size to
match. This x86 instruction uses a 64-bit address size while loading 32-bit
elements. This matches this specific variant of the SVE instruction, but
the data is zero-extended once loaded, requiring us to shuffle the data
after it is loaded.

This isn't the worst but the implementation is different enough that
stuffing it in to the other gather load will cause headaches.

Basically gets 32 instruction variants to use the SVE version!

Fixes #3827
2024-07-08 17:19:18 -07:00
Ryan Houdek c5a0ae7b34 IR: Adds new QPS gather load variant! 2024-07-08 17:19:18 -07:00
Ryan Houdek 4bd207ebf3 Arm64: Moves 128Bit gather ASIMD emulation to its own helper
It is going to get reused.
2024-07-08 17:19:18 -07:00
Tony Wasserka 45011234d9 Merge pull request #3845 from pmatos/TESTJOBCOUNTFix
Use nproc only if TEST_JOB_COUNT not specified
2024-07-08 22:31:23 +02:00
Paulo Matos 24017f379e Use nproc only if TEST_JOB_COUNT not specified 2024-07-08 21:38:56 +02:00
Mai aad7656b38 Merge pull request #3826 from Sonicadvance1/scale_32bit_gather
AVX128: Extend 32-bit address indices when possible
2024-07-08 15:29:44 -04:00
Ryan Houdek 80de890f05 InstcountCI: Update for removed FMA moves 2024-07-08 04:50:49 -07:00
Ryan Houdek 62cec7b6b2 Arm64: Remove one move if possible in FMA operations
If the destination isn't any of the incoming sources then we can avoid
one of the moves at the end. This half works around the problem proposed
in #3794, but doesn't solve the entire problem.

To solve the other half of the moving problem means we need to solve the
SRA allocation problem for this temporary register with addsub/subadd, so it gets allocated
for both the FMA operation and the XOR operation.
2024-07-08 04:44:40 -07:00
Ryan Houdek c9c163cd7b unittests: Update vcv{t,tt}pd2dq tests to ensure upper bits of destination are cleared 2024-07-08 03:30:10 -07:00
Mai 95a9f32bf0 Merge pull request #3840 from Sonicadvance1/extend_vinsert128_tests
unittests: Extends vinsert{i,f}128 tests for garbage data
2024-07-07 13:39:20 -04:00
Mai c4ae761a0e Merge pull request #3841 from Sonicadvance1/add_missing_cpu_names
CPUID: Adds a few missing CPU names for new CPU cores
2024-07-07 13:38:27 -04:00
Ryan Houdek 0653b346e0 CPUID: Adds a few missing CPU names for new CPU cores
These should be making their way to the market sooner rather than later
so make sure we have the descriptor text for them.
2024-07-07 02:40:19 -07:00
Ryan Houdek fa587398bd unittests: Extends vinsert{i,f}128 tests for garbage data
Just to ensure we don't hit an issue with masking the immediate bits.

Fixes #3753
2024-07-07 02:16:21 -07:00
Ryan Houdek 6b67857151 InstcountCI: Adds a missing gather instruction invariant
Oops, must have accidentally deleted this while copying things around.
2024-07-06 18:32:36 -07:00
Ryan Houdek 81165f0c40 InstcountCI: Update for 32-bit gather sign extend optimization 2024-07-06 18:32:35 -07:00
Ryan Houdek df40515087 AVX128: Extend 32-bit address indices when possible
When loading 256-bits of data with only 128-bits of address indices, we
can sign extend the source indices to be 64-bit. Thus falling down the
ideal path for SVE where each 128-bit lane is loading the data to
addresses in a 1:1 element ratio.

This means we use the SVE path more often because of this.

Based on top of #3825 because the prescaling behaviour was introduced
there. This implements its own prescaling when the sign extension occurs
because ARM's SSHLL{,2} instruction gives us that for free.

This additionally fixes a bug where we were accidentally loading the top
128-bit half of the addresses for gathers when it was unnecessary, and
on the AVX256 side it was duplicating and doing some additional work
when it shouldn't have.

It'll be good to walk the commits when looking at this one, as there are
a couple of incremental changes that are easier to follow that way.

Fixes #3806
2024-07-06 18:32:35 -07:00
Ryan Houdek c77922e3e5 InstcountCI: Update for previous fix 2024-07-06 18:32:35 -07:00
Ryan Houdek 0f9abe68b9 AVX128: Fixes accidentally loading high addr register when unnnecessary
Was missing a clamp on the high half when encounting a 128-bit gather
instruction. Was causing us to unconditionally load the top half when it
was unncessary.
2024-07-06 18:32:35 -07:00
Ryan Houdek c168ee6940 Arm64: Implements VSSHLL{,2} IR ops 2024-07-06 18:32:35 -07:00
Ryan Houdek 0d4414fdd0 AVX128: Removes templated AddrElementSize and add as argument
NFC
2024-07-06 18:32:35 -07:00
Ryan Houdek 968d5e0d8f Merge pull request #3774 from bylaws/win-ci
FEXCore ARM64EC CI support
2024-07-06 18:22:57 -07:00
Ryan Houdek 635182b57c Merge pull request #3832 from bylaws/wow64-wine
WOW64: Mark the FEX dll as a wine builtin
2024-07-06 17:58:00 -07:00
Ryan Houdek 9d0b6ce75e Merge pull request #3835 from bylaws/ec-topdown
AllocatorHooks: Allocate from the top down on windows
2024-07-06 17:40:36 -07:00
Ryan Houdek 2fdd80fe3a Merge pull request #3833 from bylaws/common-tso
Windows: Commonise TSOHandlerConfig
2024-07-06 17:38:45 -07:00
Ryan Houdek dbac23b749 Merge pull request #3834 from bylaws/ec-amd64
Windows: Report as an AMD64 processor when targeting ARM64EC
2024-07-06 17:38:13 -07:00
Billy Laws 7fa7061aa5 Windows: Report as an AMD64 processor when targeting ARM64EC 2024-07-06 20:37:15 +00:00
Billy Laws e45e631199 AllocatorHooks: Allocate from the top down on windows
FEX allocations can get in the way of allocations that are 4gb-limited
even in 65-bit mode (i.e. those from LuaJIT), so allocate starting from
the top of the AS to prevent conflicts.
2024-07-06 20:35:38 +00:00
Billy Laws b21e77c1e0 Windows: Commonise TSOHandlerConfig 2024-07-06 19:20:49 +00:00
Billy Laws ba33294225 WOW64: Mark the FEX dll as a wine builtin
Allows it to be automatically picked up by wine during prefix setup,
without a manual dll override.

Thanks to AndreRH for pointing me to this.
2024-07-06 19:19:36 +00:00
Billy Laws 97c21cc3a7 CI: Add ARM64EC build CI 2024-07-06 17:27:41 +01:00
Billy Laws 7d7e6f5326 CMake: Disable WOW64 module for ARM64EC 2024-07-06 17:27:41 +01:00
Billy Laws 5e15bd935e CMake: Disable glibc jemalloc for MinGW builds 2024-07-06 17:27:41 +01:00
Ryan Houdek 9bad09c45f Merge pull request #3823 from alyssarosenzweig/bug/shl-var-small
Fix CF with small shifts
2024-07-06 01:33:57 -07:00
Ryan Houdek 47d077ff22 Merge pull request #3825 from Sonicadvance1/scale_64bit_gather
AVX128: Prescale addresses in gathers if possible
2024-07-05 19:10:43 -07:00
Ryan Houdek bbf8dde3ca Merge pull request #3824 from alyssarosenzweig/bug/rc2
OpcodeDispatcher: Fix 8/16-bit rcr masking
2024-07-05 17:01:16 -07:00
Ryan Houdek 6e8ca3bc6c InstcountCI: Update for gather prescaling 2024-07-05 16:47:11 -07:00
Ryan Houdek 11a494d7b3 AVX128: Prescale addresses in gathers if possible
If the host supports SVE128, if the address element size and data size is 64-bit, and the scale is not one of the two that is supported by SVE; Then prescale the addresses.
64-bit address overflow masks the top bits so is well defined that we
can scale the vector elements and still execute the SVE code path in
that case. Removing the ASIMD code paths from a lot of gathers.

Fixes #3805
2024-07-05 16:47:11 -07:00
Alyssa Rosenzweig 9b570de33f InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:44:21 -04:00
Ryan Houdek b67343fc5a unittests: Adds a test for small shift flags calculation
Currently we calculate CF incorrectly in the case of small shifts with
large offsets.
2024-07-05 18:38:12 -04:00
Alyssa Rosenzweig 5a3c0eb83c OpcodeDispatcher: fix shl with 8/16-bit variable
the special case here lines up with the special case of using a larger shift for
a smaller result, so we can just grab CF from the larger result.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:38:12 -04:00
Alyssa Rosenzweig 10391608a0 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:34:18 -04:00
Ryan Houdek 51c57cc5ae unittests: More rotate with carry unit tests
Looks like we missed some edge cases with small carry rotate. Adds even
more unit tests.
2024-07-05 18:34:18 -04:00
Alyssa Rosenzweig 05e4678e65 OpcodeDispatcher: fix missing masking on smaller RCR
I probably broke this when working on eliminating crossblock liveness.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:34:18 -04:00
Alyssa Rosenzweig 0f0e402db4 OpcodeDispatcher: fix CF with 8/16-bit immediate
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 18:24:34 -04:00
Alyssa Rosenzweig 837bccb1d8 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 17:24:51 -04:00
Alyssa Rosenzweig adc709db2f OpcodeDispatcher: drop remnants of deferred flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 17:22:41 -04:00
Alyssa Rosenzweig 395573720d OpcodeDispatcher: drop pointless flag defers for shifts
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 0e62759d24 OpcodeDispatcher: stop deferring logical
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 926b6c3117 OpcodeDispatcher: don't defer mul flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig c9f9304ba5 OpcodeDispatcher: stop deferring obscure bitwise
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig fabd6be5af OpcodeDispatcher: drop SUB defer
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:54 -04:00
Alyssa Rosenzweig 1bf31d20b6 OpcodeDispatcher: switch to CalculateFlags_SUB
most of these are deferred only to be calculated immediately anyway.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 16:24:53 -04:00
Ryan Houdek 653bf04db0 Merge pull request #3819 from alyssarosenzweig/bug/rcr-smol
Fix 8/16-bit RCR
2024-07-05 12:49:23 -07:00
Ryan Houdek b77a25b21a Merge pull request #3818 from alyssarosenzweig/jit/shiftbymaskstozero
JIT: fix ShiftFlags masking
2024-07-05 12:49:16 -07:00
Alyssa Rosenzweig 9db6931cea InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 10:49:12 -04:00
Ryan Houdek bad5cef52b unittests: Adds rotate with carry test for large rotates
FEX-Emu currently doesn't do large rotates for small data sources
correctly. This will fail CI until fixed in OpcodeDispatcher
2024-07-05 10:49:02 -04:00
Alyssa Rosenzweig 94bd79b2bf OpcodeDispatcher: fix 8/16-bit RCR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 10:49:02 -04:00
Alyssa Rosenzweig b746146f4e InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 09:57:42 -04:00
Ryan Houdek 8ac9bb5c72 unittests: Adds test for flags when shifting by zero 2024-07-05 09:57:42 -04:00
Alyssa Rosenzweig 1b552a6f62 JIT: fix ShiftFlags masking
we don't update flags for a nonzero shift that masks to zero.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-05 09:57:42 -04:00
Alyssa Rosenzweig 97329ccc7a Merge pull request #3812 from Sonicadvance1/fix_rotates_with_zero
OpcodeDispatcher: Fixes rotates with zero not zero extending 32-bit result
2024-07-05 09:48:01 -04:00
Mai f2d1f2de56 Merge pull request #3817 from Sonicadvance1/fix_x87_integer_indefinite
Softfloat: Fixes Integer indefinite return for 16-bit signed values
2024-07-04 23:11:44 -04:00
Ryan Houdek 692c2fae96 Merge pull request #3813 from alyssarosenzweig/bug/fix-sbb
Fix 16-bit SBB
2024-07-04 19:52:37 -07:00
Mai 3d65b701a2 Merge pull request #3816 from Sonicadvance1/fix_long_signed_divide
Arm64: Fixes long signed divide
2024-07-04 21:43:11 -04:00
Ryan Houdek ecaca0fe15 unittests: Adds x87 integer indefinite test
Tests 16-bit, 32-bit, and 64-bit integer conversions
2024-07-04 17:53:28 -07:00
Ryan Houdek 8955f83ef6 Softfloat: Fixes Integer indefinite return for 16-bit signed values
Regardless of positive or negative value, if the converted integer
doesn't fit in to the converted int16_t then it returns INT16_MIN.
2024-07-04 17:43:28 -07:00
Ryan Houdek 1a8aaebd79 unittests: Adds long signed divide test 2024-07-04 16:43:21 -07:00
Ryan Houdek 38a823cc54 Arm64: Fixes long signed divide
The two halves are provided as two uint64_t values that shouldn't be
sign extended between them. Treat them as uint64_t until combined in to
a single int128_t. Fixes long signed divide.
2024-07-04 16:42:23 -07:00
Ryan Houdek 25306cb373 InstcountCI: Update 2024-07-04 14:35:43 -07:00
Ryan Houdek 1084a031e7 unittests: Adds test for previous fix
All of these results would have failed except for the rorx result.
2024-07-04 14:35:43 -07:00
Ryan Houdek f6ec99bede OpcodeDispatcher: Fixes rotates with zero not zero extending 32-bit result
For all the 32-bit rotates (except for RORX) we were failing to zero
extend the 32-bit result to the destination register when the rotate was
masked to zero.

Ensure we do this.
2024-07-04 14:35:42 -07:00
Ryan Houdek 90a6647fa4 Merge pull request #3811 from alyssarosenzweig/ra/fix-lsp
RA: fix interaction between SRA & shuffles
2024-07-04 14:20:46 -07:00
Alyssa Rosenzweig a926bb81a9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig fbf41e3149 unittests: add test for small sbc flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig a38205069b OpcodeDispatcher: fix SBB carry flag
do it the naive way, just applying the x86 definitions of SBB.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 16:58:45 -04:00
Alyssa Rosenzweig 2d75801024 unittests: add tricky RA test
this fails on current main with blocksize=500 due to mentioned RA bug. passes
with blocksize=1.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 13:37:13 -04:00
Alyssa Rosenzweig 504511fe7e RA: fix interaction between SRA & shuffles
missed a Map. tricky case hit by the unit test added in the next commit.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2024-07-04 13:37:13 -04:00
316 changed files with 108686 additions and 150197 deletions

No files matched your search

+7 -2
View File
@@ -16,7 +16,7 @@ jobs:
runs-on: ${{ matrix.arch }}
strategy:
matrix:
arch: [[self-hosted, ARM64, mingw]]
arch: [[self-hosted, ARM64, mingw], [self-hosted, ARM64EC, mingw, ARM64]]
fail-fast: false
steps:
@@ -38,6 +38,11 @@ jobs:
run: |
echo "MINGW_TRIPLE=aarch64-w64-mingw32" >> $GITHUB_ENV
- name: Set CC Arm64EC
if: matrix.arch[1] == 'ARM64EC'
run: |
echo "MINGW_TRIPLE=arm64ec-w64-mingw32" >> $GITHUB_ENV
- name: Set rootfs paths
run: |
echo "FEX_ROOTFS_MOUNT=/mnt/AutoNFS/rootfs/" >> $GITHUB_ENV
@@ -73,7 +78,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+19 -6
View File
@@ -72,23 +72,22 @@ jobs:
# Execute the build. You can specify a specific target with "--target <NAME>"
run: cmake --build . --config $BUILD_TYPE
- name: ASM Tests
- name: ASM Tests - SVE256
working-directory: ${{runner.workspace}}/build
shell: bash
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test Results move
- name: ASM Test SVE256 Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM.log || true
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE256Bit.log || true
- name: ASM Tests 128-bit
- name: ASM Tests - SVE128
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disableavx"
FEX_FORCESVEWIDTH: "128"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
@@ -97,7 +96,21 @@ jobs:
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM128bit.log || true
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_SVE128Bit.log || true
- name: ASM Tests - ASIMD
working-directory: ${{runner.workspace}}/build
shell: bash
env:
FEX_HOSTFEATURES: "disablesve"
# Execute the unit tests
run: cmake --build . --config $BUILD_TYPE --target asm_tests
- name: ASM Test ASIMD Results move
if: ${{ always() }}
shell: bash
working-directory: ${{runner.workspace}}/build
run: mv ${{runner.workspace}}/build/Testing/Temporary/LastTest.log ${{runner.workspace}}/build/Testing/Temporary/LastTest_ASM_ASIMD.log || true
- name: Truncate test results
if: ${{ always() }}
+50 -10
View File
@@ -1,5 +1,5 @@
cmake_minimum_required(VERSION 3.14)
project(FEX)
project(FEX C CXX ASM)
INCLUDE (CheckIncludeFiles)
CHECK_INCLUDE_FILES ("gdb/jit-reader.h" HAVE_GDB_JIT_READER_H)
@@ -15,6 +15,7 @@ option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_COVERAGE "Enables Coverage" FALSE)
option(ENABLE_ASSERTIONS "Enables assertions in build" FALSE)
option(ENABLE_GDB_SYMBOLS "Enables GDBSymbols integration support" ${HAVE_GDB_JIT_READER_H})
option(ENABLE_STRICT_WERROR "Enables stricter -Werror for CI" FALSE)
@@ -31,6 +32,7 @@ option(COMPILE_VIXL_DISASSEMBLER "Compiles the vixl disassembler in to vixl" FAL
option(ENABLE_FEXCORE_PROFILER "Enables use of the FEXCore timeline profiling capabilities" FALSE)
set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend you want to use for the FEXCore profiler")
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
@@ -40,7 +42,8 @@ string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
message (STATUS "Mingw build")
set (MINGW_BUILD TRUE)
set (ENABLE_JEMALLOC FALSE)
set (ENABLE_JEMALLOC TRUE)
set (ENABLE_JEMALLOC_GLIBC_ALLOC FALSE)
endif()
if (NOT MINGW_BUILD)
@@ -137,9 +140,39 @@ endif()
if (CMAKE_SYSTEM_PROCESSOR MATCHES "^arm64ec")
set(_M_ARM_64EC 1)
add_definitions(-D_M_ARM_64EC=1)
endif()
# Required as FEX is not allowed to lock the CRT heap lock during compilation or callbacks
set(ENABLE_JEMALLOC TRUE)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
int Testy(int a, int b, int c, int d, int e, int f) {
return a + b + c + d + e + f;
}
int main() {
return Testy(0, 1, 2, 3, 4, 5);
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
if (MINGW_BUILD)
message(STATUS "Ignoring broken clang::preserve_all support")
set(HAS_CLANG_PRESERVE_ALL FALSE)
else()
message(STATUS "Has clang::preserve_all")
endif()
endif ()
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
add_definitions("-DFEX_PRESERVE_ALL_ATTR=__attribute__((preserve_all))" "-DFEX_HAS_PRESERVE_ALL_ATTR=1")
else()
add_definitions("-DFEX_PRESERVE_ALL_ATTR=" "-DFEX_HAS_PRESERVE_ALL_ATTR=0")
endif()
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
add_definitions(-DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_CCACHE)
@@ -189,13 +222,18 @@ if (ENABLE_TSAN)
link_libraries(-fno-omit-frame-pointer -fsanitize=thread)
endif()
if (ENABLE_COVERAGE)
add_compile_options(-fprofile-instr-generate -fcoverage-mapping)
link_libraries(-fprofile-instr-generate -fcoverage-mapping)
endif()
if (ENABLE_JEMALLOC_GLIBC_ALLOC)
# The glibc jemalloc subproject which hooks the glibc allocator.
# Required for thunks to work.
# All host native libraries will use this allocator, while *most* other FEX internal allocations will use the other jemalloc allocator.
add_definitions(-DENABLE_JEMALLOC_GLIBC=1)
add_subdirectory(External/jemalloc_glibc/)
else()
elseif (NOT MINGW_BUILD)
message (STATUS
" jemalloc glibc allocator disabled!\n"
" This is not a recommended configuration!\n"
@@ -208,7 +246,7 @@ if (ENABLE_JEMALLOC)
add_definitions(-DENABLE_JEMALLOC=1)
add_subdirectory(External/jemalloc/)
include_directories(External/jemalloc/pregen/include/)
else()
elseif (NOT MINGW_BUILD)
message (STATUS
" jemalloc disabled!\n"
" This is not a recommended configuration!\n"
@@ -216,6 +254,11 @@ else()
" Use at your own risk!")
endif()
if (USE_PDB_DEBUGINFO)
add_compile_options(-g -gcodeview)
add_link_options(-g -Wl,--pdb=)
endif()
set (CMAKE_CXX_FLAGS_RELWITHDEBINFO "${CMAKE_CXX_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
set (CMAKE_LINKER_FLAGS_RELWITHDEBINFO "${CMAKE_LINKER_FLAGS_RELWITHDEBINFO} -fno-omit-frame-pointer")
@@ -269,8 +312,6 @@ include_directories(External/json-maker/)
add_subdirectory(External/tiny-json/)
include_directories(External/tiny-json/)
include_directories(External/xbyak/)
include_directories(Source/)
include_directories("${CMAKE_BINARY_DIR}/Source/")
@@ -372,8 +413,7 @@ if (BUILD_TESTS)
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
if (TEST_JOB_COUNT)
message(STATUS "Running tests with ${TEST_JOB_COUNT} jobs")
endif()
if (CMAKE_VERSION VERSION_LESS "3.29")
elseif(CMAKE_VERSION VERSION_LESS "3.29")
execute_process(COMMAND "nproc" OUTPUT_STRIP_TRAILING_WHITESPACE OUTPUT_VARIABLE TEST_JOB_COUNT)
endif()
set(TEST_JOB_FLAG "-j${TEST_JOB_COUNT}")
+12
View File
@@ -602,6 +602,18 @@ constexpr bool AreVectorsSequential(T first, const Args&... args) {
return (fn(first, args) && ...);
}
// Returns if the immediate can fit in to add/sub immediate instruction encodings.
constexpr bool IsImmAddSub(uint64_t imm) {
constexpr uint64_t U12Mask = 0xFFF;
auto FitsWithin12Bits = [](uint64_t imm) {
return (imm & ~U12Mask) == 0;
};
// Can fit in to the instruction encoding:
// - if only bits [11:0] are set.
// - if only bits [23:12] are set.
return FitsWithin12Bits(imm) || (FitsWithin12Bits(imm >> 12) && (imm & U12Mask) == 0);
}
// This is an emitter that is designed around the smallest code bloat as possible.
// Eschewing most developer convenience in order to keep code as small as possible.
+29 -1
View File
@@ -2569,7 +2569,19 @@ public:
}
// SVE contiguous non-temporal load (scalar plus immediate)
// XXX:
void ldnt1b(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b00, zt, pg, rn, Imm);
}
void ldnt1h(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b01, zt, pg, rn, Imm);
}
void ldnt1w(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b10, zt, pg, rn, Imm);
}
void ldnt1d(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
SVEContiguousNontemporalLoad(0b11, zt, pg, rn, Imm);
}
// SVE contiguous non-temporal load (scalar plus scalar)
// XXX:
// SVE load multiple structures (scalar plus immediate)
@@ -4492,6 +4504,22 @@ private:
dc32(Instr);
}
// SVE contiguous non-temporal load (scalar plus immediate)
void SVEContiguousNontemporalLoad(uint32_t msz, ZRegister zt, PRegister pg, Register rn, int32_t imm) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
"Invalid loadstore offset ({}). Must be between [-8, 7]", imm);
const auto imm4 = static_cast<uint32_t>(imm) & 0xF;
uint32_t Instr = 0b1010'0100'0000'0000'1110'0000'0000'0000;
Instr |= msz << 23;
Instr |= imm4 << 16;
Instr |= pg.Idx() << 10;
Instr |= Encode_rn(rn);
Instr |= zt.Idx();
dc32(Instr);
}
// SVE contiguous non-temporal store (scalar plus immediate)
void SVEContiguousNontemporalStore(uint32_t msz, ZRegister zt, PRegister pg, Register rn, int32_t imm) {
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
+43 -3
View File
@@ -36,9 +36,9 @@
// to by n, imm_s and imm_r are undefined.
static bool IsImmLogical(uint64_t value,
unsigned width,
unsigned* n,
unsigned* imm_s,
unsigned* imm_r) {
unsigned* n = nullptr,
unsigned* imm_s = nullptr,
unsigned* imm_r = nullptr) {
[[maybe_unused]] constexpr auto kBRegSize = 8;
[[maybe_unused]] constexpr auto kHRegSize = 16;
[[maybe_unused]] constexpr auto kSRegSize = 32;
@@ -243,6 +243,46 @@ static bool IsImmLogical(uint64_t value,
return true;
}
static inline bool IsIntN(unsigned n, int64_t x) {
if (n == 64) return true;
int64_t limit = INT64_C(1) << (n - 1);
return (-limit <= x) && (x < limit);
}
static inline bool IsUintN(unsigned n, int64_t x) {
// Convert to an unsigned integer to avoid implementation-defined behavior.
return !(static_cast<uint64_t>(x) >> n);
}
// clang-format off
#define INT_1_TO_32_LIST(V) \
V(1) V(2) V(3) V(4) V(5) V(6) V(7) V(8) \
V(9) V(10) V(11) V(12) V(13) V(14) V(15) V(16) \
V(17) V(18) V(19) V(20) V(21) V(22) V(23) V(24) \
V(25) V(26) V(27) V(28) V(29) V(30) V(31) V(32)
#define INT_33_TO_63_LIST(V) \
V(33) V(34) V(35) V(36) V(37) V(38) V(39) V(40) \
V(41) V(42) V(43) V(44) V(45) V(46) V(47) V(48) \
V(49) V(50) V(51) V(52) V(53) V(54) V(55) V(56) \
V(57) V(58) V(59) V(60) V(61) V(62) V(63)
#define INT_1_TO_63_LIST(V) INT_1_TO_32_LIST(V) INT_33_TO_63_LIST(V)
// clang-format on
#define DECLARE_IS_INT_N(N) \
static inline bool IsInt##N(int64_t x) { return IsIntN(N, x); }
#define DECLARE_IS_UINT_N(N) \
static inline bool IsUint##N(int64_t x) { return IsUintN(N, x); }
INT_1_TO_63_LIST(DECLARE_IS_INT_N)
INT_1_TO_63_LIST(DECLARE_IS_UINT_N)
#undef DECLARE_IS_INT_N
#undef DECLARE_IS_UINT_N
private:
template <typename V>
-21
View File
@@ -24,27 +24,6 @@ include(CheckCXXCompilerFlag)
include(CheckIncludeFileCXX)
include(CheckCXXSourceCompiles)
set(CMAKE_REQUIRED_FLAGS "-std=c++11 -Wattributes -Werror=attributes")
check_cxx_source_compiles(
"
__attribute__((preserve_all))
int Testy(int a, int b, int c, int d, int e, int f) {
return a + b + c + d + e + f;
}
int main() {
return Testy(0, 1, 2, 3, 4, 5);
}"
HAS_CLANG_PRESERVE_ALL)
unset(CMAKE_REQUIRED_FLAGS)
if (HAS_CLANG_PRESERVE_ALL)
if (MINGW_BUILD)
message(STATUS "Ignoring broken clang::preserve_all support")
set(HAS_CLANG_PRESERVE_ALL FALSE)
else()
message(STATUS "Has clang::preserve_all")
endif()
endif ()
if (EXISTS ${CMAKE_CURRENT_DIR}/External/vixl/)
# Useful to have for freestanding libFEXCore
add_subdirectory(External/vixl/)
+4 -6
View File
@@ -148,9 +148,8 @@ def print_man_options(options):
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
output_man.write(", ".join(f"{enum_op_val}" for [_, enum_op_val] in Enums.items()))
output_man.write("\n.sp\n")
output_man.write(".El\n")
@@ -179,9 +178,8 @@ def print_man_environment(options):
if (value_type == "strenum"):
Enums = op_vals["Enums"]
output_man.write("\\fBAvailable Options:\\fR\n")
for enum_op_key, enum_op_vals in Enums.items():
output_man.write("{}, ".format(enum_op_vals))
output_man.write("\n")
output_man.write(", ".join(f"{enum_op_val}" for [_, enum_op_val] in Enums.items()))
output_man.write("\n.sp\n")
print_man_environment_tail()
output_man.write(".El\n")
+55 -104
View File
@@ -2,6 +2,7 @@
import json
import sys
from dataclasses import dataclass, field
import textwrap
def ExitError(msg):
print(msg)
@@ -53,6 +54,7 @@ class OpDefinition:
SSAArgNum: int
NonSSAArgNum: int
DynamicDispatch: bool
LoweredX87: bool
JITDispatch: bool
JITDispatchOverride: str
TiedSource: int
@@ -76,6 +78,7 @@ class OpDefinition:
self.SSAArgNum = 0
self.NonSSAArgNum = 0
self.DynamicDispatch = False
self.LoweredX87 = False
self.JITDispatch = True
self.JITDispatchOverride = None
self.TiedSource = -1
@@ -204,7 +207,7 @@ def parse_ops(ops):
(OpArg.Type == "GPR" or
OpArg.Type == "GPRPair" or
OpArg.Type == "FPR")):
OpDef.EmitValidation.append("GetOpRegClass({}) == InvalidClass || WalkFindRegClass({}) == {}Class".format(NameWithPrefix, NameWithPrefix, OpArg.Type))
OpDef.EmitValidation.append(f"GetOpRegClass({ArgName}) == InvalidClass || WalkFindRegClass({ArgName}) == {OpArg.Type}Class")
OpArg.Name = ArgName
OpArg.NameWithPrefix = NameWithPrefix
@@ -250,6 +253,13 @@ def parse_ops(ops):
if "JITDispatchOverride" in op_val:
OpDef.JITDispatchOverride = op_val["JITDispatchOverride"]
if "X87" in op_val:
OpDef.LoweredX87 = op_val["X87"]
# X87 implies !JITDispatch
assert("JITDispatch" not in op_val)
OpDef.JITDispatch = False
if "TiedSource" in op_val:
OpDef.TiedSource = op_val["TiedSource"]
@@ -258,12 +268,8 @@ def parse_ops(ops):
for i in range(len(OpDef.EmitValidation)):
# Patch up all the argument names
for Arg in OpDef.Arguments:
if Arg.Temporary:
# Temporary ops just replace all instances no prefix variant
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, Arg.Name)
else:
# All other ops replace $ with _ variant for argument passed in
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, "_{}".format(Arg.Name))
# Temporary ops just replace all instances no prefix variant
OpDef.EmitValidation[i] = OpDef.EmitValidation[i].replace(Arg.NameWithPrefix, Arg.Name)
#OpDef.print()
@@ -368,42 +374,28 @@ def print_ir_sizes():
if op.Name == "Last":
output_file.write("\t-1ULL,\n")
else:
output_file.write("\tsizeof(IROp_{}),\n".format(op.Name))
output_file.write(f"\tsizeof(IROp_{op.Name}),\n")
output_file.write("};\n\n")
output_file.write(textwrap.dedent("""
};
output_file.write("// Make sure our array maps directly to the IROps enum\n")
output_file.write("static_assert(IRSizes[IROps::OP_LAST] == -1ULL);\n\n")
// Make sure our array maps directly to the IROps enum
static_assert(IRSizes[IROps::OP_LAST] == -1ULL);
output_file.write("[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }
[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] bool LoweredX87(IROps Op);
[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);\n'
)
output_file.write(
'[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);\n'
)
output_file.write("#undef IROP_SIZES\n")
output_file.write("#endif\n\n")
#undef IROP_SIZES
#endif
"""))
def print_ir_reg_classes():
output_file.write("#ifdef IROP_REG_CLASSES_IMPL\n")
@@ -493,13 +485,14 @@ def print_ir_getraargs():
def print_ir_hassideeffects():
output_file.write("#ifdef IROP_HASSIDEEFFECTS_IMPL\n")
for array, prop, T in [
("SideEffects", "HasSideEffects", "bool"),
("ImplicitFlagClobbers", "ImplicitFlagClobber", "bool"),
("TiedSources", "TiedSource", "int8_t"),
for prop, T in [
("HasSideEffects", "bool"),
("ImplicitFlagClobber", "bool"),
("LoweredX87", "bool"),
("TiedSource", "int8_t"),
]:
output_file.write(
f"constexpr std::array<{'uint8_t' if T == 'bool' else T}, OP_LAST + 1> {array} = {{\n"
f"constexpr std::array<{'uint8_t' if T == 'bool' else T}, OP_LAST + 1> {prop}_ = {{\n"
)
for op in IROps:
if T == "bool":
@@ -512,7 +505,7 @@ def print_ir_hassideeffects():
output_file.write("};\n\n")
output_file.write(f"{T} {prop}(IROps Op) {{\n")
output_file.write(f" return {array}[Op];\n")
output_file.write(f" return {prop}_[Op];\n")
output_file.write("}\n")
output_file.write("#undef IROP_HASSIDEEFFECTS_IMPL\n")
@@ -671,11 +664,11 @@ def print_ir_allocator_helpers():
output_file.write("{} {}".format(CType, arg.Name));
elif arg.IsSSA:
# SSA value
output_file.write("OrderedNode *_{}".format(arg.Name))
output_file.write("OrderedNode *{}".format(arg.Name))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("{} _{}".format(CType, arg.Name));
output_file.write("{} {}".format(CType, arg.Name));
if arg.DefaultInitializer != None:
output_file.write(" = {}".format(arg.DefaultInitializer))
@@ -689,23 +682,28 @@ def print_ir_allocator_helpers():
if op.ImplicitFlagClobber:
output_file.write("\t\tSaveNZCV(IROps::OP_{});".format(op.Name.upper()))
output_file.write("\t\tauto Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
# We gather the "has x87?" flag as we go. This saves the user from
# having to keep track of whether they emitted any x87.
if op.LoweredX87:
output_file.write("\t\tRecordX87Use();\n")
output_file.write("\t\tauto _Op = AllocateOp<IROp_{}, IROps::OP_{}>();\n".format(op.Name, op.Name.upper()))
if op.SSAArgNum != 0:
output_file.write("\t\tauto ListDataBegin = DualListData.ListBegin();\n")
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tOp.first->{} = _{}->Wrapped(ListDataBegin);\n".format(arg.Name, arg.Name))
output_file.write("\t\t_Op.first->{} = {}->Wrapped(ListDataBegin);\n".format(arg.Name, arg.Name))
if op.SSAArgNum != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\t_{}->AddUse();\n".format(arg.Name))
output_file.write("\t\t{}->AddUse();\n".format(arg.Name))
if len(op.Arguments) != 0:
for arg in op.Arguments:
if not arg.Temporary and not arg.IsSSA:
output_file.write("\t\tOp.first->{} = _{};\n".format(arg.Name, arg.Name))
output_file.write("\t\t_Op.first->{} = {};\n".format(arg.Name, arg.Name))
if (op.HasDest):
# We can only infer a size if we have arguments
@@ -715,22 +713,22 @@ def print_ir_allocator_helpers():
if len(op.Arguments) != 0:
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tuint8_t Size{} = GetOpSize(_{});\n".format(arg.Name, arg.Name))
output_file.write("\t\tuint8_t Size{} = GetOpSize({});\n".format(arg.Name, arg.Name))
for arg in op.Arguments:
if arg.IsSSA:
output_file.write("\t\tInferSize = std::max(InferSize, Size{});\n".format(arg.Name))
output_file.write("\t\tOp.first->Header.Size = InferSize;\n")
output_file.write("\t\t_Op.first->Header.Size = InferSize;\n")
# Some ops without a destination still need an operating size
# Effectively reusing the destination size value for operation size
if op.DestSize != None:
output_file.write("\t\tOp.first->Header.Size = {};\n".format(op.DestSize))
output_file.write("\t\t_Op.first->Header.Size = {};\n".format(op.DestSize))
if op.NumElements == None:
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / ({});\n".format(1))
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(1))
else:
output_file.write("\t\tOp.first->Header.ElementSize = Op.first->Header.Size / ({});\n".format(op.NumElements))
output_file.write("\t\t_Op.first->Header.ElementSize = _Op.first->Header.Size / ({});\n".format(op.NumElements))
# Insert validation here
if op.EmitValidation != None:
@@ -741,58 +739,12 @@ def print_ir_allocator_helpers():
output_file.write("\tLOGMAN_THROW_A_FMT({}, \"{}\");\n".format(Validation, Sanitized))
output_file.write("\t\t#endif\n")
output_file.write("\t\treturn Op;\n")
output_file.write("\t\treturn _Op;\n")
output_file.write("\t}\n\n")
output_file.write("#undef IROP_ALLOCATE_HELPERS\n")
output_file.write("#endif\n")
def print_ir_parser_switch_helper():
output_file.write("#ifdef IROP_PARSER_SWITCH_HELPERS\n")
for op in IROps:
if op.Name != "Last" and op.SwitchGen:
output_file.write("\tcase FEXCore::IR::IROps::OP_%s: {\n" % (op.Name.upper()))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
if arg.Temporary:
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("\t\tauto arg{} = DecodeValue<{}>(Def.Args[{}]);\n".format(i, CType, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
elif arg.IsSSA:
# SSA value
output_file.write("\t\tauto arg{} = DecodeValue<OrderedNode*>(Def.Args[{}]);\n".format(i, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
else:
# User defined op that is stored
CType = IRTypesToCXX[arg.Type].CXXName
output_file.write("\t\tauto arg{} = DecodeValue<{}>(Def.Args[{}]);\n".format(i, CType, i))
output_file.write("\t\tif (!CheckPrintErrorArg(Def, arg{}.first, {})) return false;\n".format(i, i))
output_file.write("\t\tDef.Node = _{}(\n".format(op.Name))
for i in range(0, len(op.Arguments)):
arg = op.Arguments[i]
LastArg = len(op.Arguments) - i - 1 == 0
output_file.write("\t\t\targ{}.second".format(i))
if not LastArg:
output_file.write(",\n")
else:
output_file.write("\n")
output_file.write("\t\t);\n")
output_file.write("\t\tSSANameMapper[Def.Definition] = Def.Node;\n")
output_file.write("\t\tbreak;\n")
output_file.write("\t}\n")
output_file.write("#undef IROP_PARSER_SWITCH_HELPERS\n")
output_file.write("#endif\n")
def print_ir_dispatcher_defs():
output_dispatch_file.write("#ifdef IROP_DISPATCH_DEFS\n")
for op in IROps:
@@ -851,7 +803,6 @@ print_ir_hassideeffects()
print_ir_gethasdest()
print_ir_arg_printer()
print_ir_allocator_helpers()
print_ir_parser_switch_helper()
output_file.close()
+2 -10
View File
@@ -67,7 +67,6 @@ set (SRCS
Common/SoftFloat-3e/s_approxRecipSqrt32_1.c
Common/SoftFloat-3e/s_approxRecipSqrt_1Ks.c
Common/SoftFloat-3e/softfloat_raiseFlags.c
Common/SoftFloat-3e/softfloat_state.c
Common/SoftFloat-3e/f64_to_extF80.c
Common/SoftFloat-3e/s_commonNaNToExtF80UI.c
Common/SoftFloat-3e/s_normSubnormalF64Sig.c
@@ -91,7 +90,6 @@ set (SRCS
Interface/Core/CPUBackend.cpp
Interface/Core/CPUID.cpp
Interface/Core/Frontend.cpp
Interface/Core/HostFeatures.cpp
Interface/Core/ObjectCache/JobHandling.cpp
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
Interface/Core/ObjectCache/ObjectCacheService.cpp
@@ -113,7 +111,6 @@ set (SRCS
Interface/Core/JIT/Arm64/BranchOps.cpp
Interface/Core/JIT/Arm64/ConversionOps.cpp
Interface/Core/JIT/Arm64/EncryptionOps.cpp
Interface/Core/JIT/Arm64/FlagOps.cpp
Interface/Core/JIT/Arm64/MemoryOps.cpp
Interface/Core/JIT/Arm64/MiscOps.cpp
Interface/Core/JIT/Arm64/MoveOps.cpp
@@ -138,13 +135,13 @@ set (SRCS
Interface/IR/IREmitter.cpp
Interface/IR/PassManager.cpp
Interface/IR/Passes/ConstProp.cpp
Interface/IR/Passes/DeadContextStoreElimination.cpp
Interface/IR/Passes/IRDumperPass.cpp
Interface/IR/Passes/IRValidation.cpp
Interface/IR/Passes/RAValidation.cpp
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
Interface/IR/Passes/DeadStoreElimination.cpp
Interface/IR/Passes/RegisterAllocationPass.cpp
Interface/IR/Passes/x87StackOptimizationPass.cpp
Utils/Telemetry.cpp
Utils/Threads.cpp
Utils/Profiler.cpp
@@ -161,7 +158,7 @@ if (ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT)
Utils/AllocatorOverride.cpp)
endif()
set(DEFINES -DTHREAD_LOCAL=_Thread_local -DJIT_ARM64)
set(DEFINES -DJIT_ARM64)
if (_M_X86_64)
list(APPEND DEFINES -D_M_X86_64=1)
@@ -171,11 +168,6 @@ if (_M_ARM_64)
list(APPEND DEFINES -D_M_ARM_64=1)
endif()
if (ENABLE_VIXL_SIMULATOR)
# We can run the simulator on both x86-64 or AArch64 hosts
list(APPEND DEFINES -DVIXL_SIMULATOR=1 -DVIXL_INCLUDE_SIMULATOR_AARCH64=1)
endif()
if (ENABLE_VIXL_DISASSEMBLER)
list(APPEND DEFINES -DVIXL_DISASSEMBLER=1)
endif()
@@ -41,7 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_add( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -53,7 +53,7 @@ extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
bool signB;
extFloat80_t
(*magsFuncPtr)(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
uA.f = a;
uiA64 = uA.s.signExp;
@@ -65,6 +65,6 @@ extFloat80_t extF80_add( extFloat80_t a, extFloat80_t b )
signB = signExtF80UI64( uiB64 );
magsFuncPtr =
(signA == signB) ? softfloat_addMagsExtF80 : softfloat_subMagsExtF80;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
return (*magsFuncPtr)( state, uiA64, uiA0, uiB64, uiB0, signA );
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_div( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -107,7 +107,7 @@ extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
if ( ! (sigB & UINT64_C( 0x8000000000000000 )) ) {
if ( ! sigB ) {
if ( ! sigA ) goto invalid;
softfloat_raiseFlags( softfloat_flag_infinite );
softfloat_raiseFlags( state, softfloat_flag_infinite );
goto infinity;
}
normExpSig = softfloat_normSubnormalExtF80Sig( sigB );
@@ -169,18 +169,18 @@ extFloat80_t extF80_div( extFloat80_t a, extFloat80_t b )
sigZExtra = (uint64_t) ((uint_fast64_t) q<<41);
return
softfloat_roundPackToExtF80(
signZ, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, signZ, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t a, extFloat80_t b )
bool extF80_eq( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -62,7 +62,7 @@ bool extF80_eq( extFloat80_t a, extFloat80_t b )
softfloat_isSigNaNExtF80UI( uiA64, uiA0 )
|| softfloat_isSigNaNExtF80UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t a, extFloat80_t b )
bool extF80_lt( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -59,7 +59,7 @@ bool extF80_lt( extFloat80_t a, extFloat80_t b )
uiB64 = uB.s.signExp;
uiB0 = uB.s.signif;
if ( isNaNExtF80UI( uiA64, uiA0 ) || isNaNExtF80UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signExtF80UI64( uiA64 );
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_mul( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -125,11 +125,11 @@ extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
}
return
softfloat_roundPackToExtF80(
signZ, expZ, sig128Z.v64, sig128Z.v0, extF80_roundingPrecision );
state, signZ, expZ, sig128Z.v64, sig128Z.v0, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
@@ -137,7 +137,7 @@ extFloat80_t extF80_mul( extFloat80_t a, extFloat80_t b )
*------------------------------------------------------------------------*/
infArg:
if ( ! magBits ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
} else {
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_rem( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -193,18 +193,18 @@ extFloat80_t extF80_rem( extFloat80_t a, extFloat80_t b )
}
return
softfloat_normRoundPackToExtF80(
signRem, rem.v64 | rem.v0 ? expB + 32 : 0, rem.v64, rem.v0, 80 );
state, signRem, rem.v64 | rem.v0 ? expB + 32 : 0, rem.v64, rem.v0, 80 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
extF80_roundToInt( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_roundToInt( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64, signUI64;
@@ -80,7 +80,7 @@ extFloat80_t
if ( 0x403E <= exp ) {
if ( exp == 0x7FFF ) {
if ( sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
uiZ = softfloat_propagateNaNExtF80UI( uiA64, sigA, 0, 0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, sigA, 0, 0 );
uiZ64 = uiZ.v64;
sigZ = uiZ.v0;
goto uiZ;
@@ -93,7 +93,7 @@ extFloat80_t
goto uiZ;
}
if ( exp <= 0x3FFE ) {
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
switch ( roundingMode ) {
case softfloat_round_near_even:
if ( !(sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF )) ) break;
@@ -145,7 +145,7 @@ extFloat80_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) sigZ |= lastBitMask;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
uiZ:
uZ.s.signExp = uiZ64;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t a )
extFloat80_t extF80_sqrt( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -74,7 +74,7 @@ extFloat80_t extF80_sqrt( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, 0, 0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, 0, 0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
goto uiZ;
@@ -155,11 +155,11 @@ extFloat80_t extF80_sqrt( extFloat80_t a )
}
return
softfloat_roundPackToExtF80(
0, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, 0, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -41,7 +41,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
extFloat80_t extF80_sub( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -54,7 +54,7 @@ extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
extFloat80_t
(*magsFuncPtr)(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
#endif
uA.f = a;
@@ -67,14 +67,14 @@ extFloat80_t extF80_sub( extFloat80_t a, extFloat80_t b )
signB = signExtF80UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_subMagsExtF80( uiA64, uiA0, uiB64, uiB0, signA );
return softfloat_subMagsExtF80( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_addMagsExtF80( uiA64, uiA0, uiB64, uiB0, signA );
return softfloat_addMagsExtF80( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_subMagsExtF80 : softfloat_addMagsExtF80;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
return (*magsFuncPtr)( state, uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t a )
float128_t extF80_to_f128( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -61,7 +61,7 @@ float128_t extF80_to_f128( extFloat80_t a )
exp = expExtF80UI64( uiA64 );
frac = uiA0 & UINT64_C( 0x7FFFFFFFFFFFFFFF );
if ( (exp == 0x7FFF) && frac ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF128UI( &commonNaN );
} else {
sign = signExtF80UI64( uiA64 );
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t a )
float32_t extF80_to_f32( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -66,7 +66,7 @@ float32_t extF80_to_f32( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF32UI( &commonNaN );
} else {
uiZ = packToF32UI( sign, 0xFF, 0 );
@@ -86,7 +86,7 @@ float32_t extF80_to_f32( extFloat80_t a )
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF32( sign, exp, sig32 );
return softfloat_roundPackToF32( state, sign, exp, sig32 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t a )
float64_t extF80_to_f64( struct softfloat_state *state, extFloat80_t a )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -72,7 +72,7 @@ float64_t extF80_to_f64( extFloat80_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
softfloat_extF80UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_extF80UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF64UI( &commonNaN );
} else {
uiZ = packToF64UI( sign, 0x7FF, 0 );
@@ -86,7 +86,7 @@ float64_t extF80_to_f64( extFloat80_t a )
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF64( sign, exp, sig );
return softfloat_roundPackToF64( state, sign, exp, sig );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
extF80_to_i32( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_i32( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -68,7 +68,7 @@ int_fast32_t
#elif (i32_fromNaN == i32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return i32_fromNaN;
#endif
}
@@ -78,7 +78,7 @@ int_fast32_t
shiftDist = 0x4032 - exp;
if ( shiftDist <= 0 ) shiftDist = 1;
sig = softfloat_shiftRightJam64( sig, shiftDist );
return softfloat_roundToI32( sign, sig, roundingMode, exact );
return softfloat_roundToI32( state, sign, sig, roundingMode, exact );
}
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
extF80_to_i64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_i64( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -68,7 +68,7 @@ int_fast64_t
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
? i64_fromNaN
@@ -84,7 +84,7 @@ int_fast64_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
}
return softfloat_roundToI64( sign, sig, sigExtra, roundingMode, exact );
return softfloat_roundToI64( state, sign, sig, sigExtra, roundingMode, exact );
}
@@ -43,7 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t
extF80_to_ui64( extFloat80_t a, uint_fast8_t roundingMode, bool exact )
extF80_to_ui64( struct softfloat_state *state, extFloat80_t a, uint_fast8_t roundingMode, bool exact )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
@@ -65,7 +65,7 @@ uint_fast64_t
*------------------------------------------------------------------------*/
shiftDist = 0x403E - exp;
if ( shiftDist < 0 ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
? ui64_fromNaN
@@ -79,7 +79,7 @@ uint_fast64_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
}
return softfloat_roundToUI64( sign, sig, sigExtra, roundingMode, exact );
return softfloat_roundToUI64( state, sign, sig, sigExtra, roundingMode, exact );
}
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t a )
extFloat80_t f128_to_extF80( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
@@ -70,7 +70,7 @@ extFloat80_t f128_to_extF80( float128_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 | frac0 ) {
softfloat_f128UIToCommonNaN( uiA64, uiA0, &commonNaN );
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -98,7 +98,7 @@ extFloat80_t f128_to_extF80( float128_t a )
sig128 =
softfloat_shortShiftLeft128(
frac64 | UINT64_C( 0x0001000000000000 ), frac0, 15 );
return softfloat_roundPackToExtF80( sign, exp, sig128.v64, sig128.v0, 80 );
return softfloat_roundPackToExtF80( state, sign, exp, sig128.v64, sig128.v0, 80 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t a )
extFloat80_t f32_to_extF80( struct softfloat_state *state, float32_t a )
{
union ui32_f32 uA;
uint_fast32_t uiA;
@@ -67,7 +67,7 @@ extFloat80_t f32_to_extF80( float32_t a )
*------------------------------------------------------------------------*/
if ( exp == 0xFF ) {
if ( frac ) {
softfloat_f32UIToCommonNaN( uiA, &commonNaN );
softfloat_f32UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include "softfloat.h"
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t a )
extFloat80_t f64_to_extF80( struct softfloat_state *state, float64_t a )
{
union ui64_f64 uA;
uint_fast64_t uiA;
@@ -67,7 +67,7 @@ extFloat80_t f64_to_extF80( float64_t a )
*------------------------------------------------------------------------*/
if ( exp == 0x7FF ) {
if ( frac ) {
softfloat_f64UIToCommonNaN( uiA, &commonNaN );
softfloat_f64UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToExtF80UI( &commonNaN );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
@@ -63,19 +63,19 @@ uint_fast32_t softfloat_roundToUI32( bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
uint_fast64_t
softfloat_roundToUI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
struct softfloat_state *, bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
#else
uint_fast64_t softfloat_roundMToUI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t softfloat_roundToI32( bool, uint_fast64_t, uint_fast8_t, bool );
int_fast32_t softfloat_roundToI32( struct softfloat_state *, bool, uint_fast64_t, uint_fast8_t, bool );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
struct softfloat_state *, bool, uint_fast64_t, uint_fast64_t, uint_fast8_t, bool );
#else
int_fast64_t softfloat_roundMToI64( bool, uint32_t *, uint_fast8_t, bool );
#endif
@@ -115,7 +115,7 @@ FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig32 softfloat_normSubnormalF32Sig( uint_fast32_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t softfloat_roundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_roundPackToF32( struct softfloat_state *, bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_normRoundPackToF32( bool, int_fast16_t, uint_fast32_t );
float32_t softfloat_addMagsF32( uint_fast32_t, uint_fast32_t );
@@ -138,7 +138,7 @@ FEXCORE_PRESERVE_ALL_ATTR
struct exp16_sig64 softfloat_normSubnormalF64Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t softfloat_roundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_roundPackToF64( struct softfloat_state *, bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_normRoundPackToF64( bool, int_fast16_t, uint_fast64_t );
float64_t softfloat_addMagsF64( uint_fast64_t, uint_fast64_t, bool );
@@ -167,18 +167,18 @@ struct exp32_sig64 softfloat_normSubnormalExtF80Sig( uint_fast64_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
struct softfloat_state *, bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
struct softfloat_state *, bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast8_t );
extFloat80_t
softfloat_addMagsExtF80(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
extFloat80_t
softfloat_subMagsExtF80(
uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
struct softfloat_state *, uint_fast16_t, uint_fast64_t, uint_fast16_t, uint_fast64_t, bool );
/*----------------------------------------------------------------------------
*----------------------------------------------------------------------------*/
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extFloat80_t
softfloat_addMagsExtF80(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -140,11 +141,11 @@ extFloat80_t
roundAndPack:
return
softfloat_roundPackToExtF80(
signZ, expZ, sigZ, sigZExtra, extF80_roundingPrecision );
state, signZ, expZ, sigZ, sigZExtra, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
uiZ:
@@ -49,11 +49,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
struct softfloat_state *state, uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNExtF80UI( uiA64, uiA0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA64>>15;
zPtr->v64 = uiA0<<1;
@@ -50,12 +50,12 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
struct softfloat_state *state, uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr )
{
struct uint128 NaNSig;
if ( softfloat_isSigNaNF128UI( uiA64, uiA0 ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
NaNSig = softfloat_shortShiftLeft128( uiA64, uiA0, 16 );
zPtr->sign = uiA64>>63;
@@ -46,11 +46,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr )
void softfloat_f32UIToCommonNaN( struct softfloat_state *state, uint_fast32_t uiA, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNF32UI( uiA ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA>>31;
zPtr->v64 = (uint_fast64_t) uiA<<41;
@@ -46,11 +46,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr )
void softfloat_f64UIToCommonNaN( struct softfloat_state *state, uint_fast64_t uiA, struct commonNaN *zPtr )
{
if ( softfloat_isSigNaNF64UI( uiA ) ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
zPtr->sign = uiA>>63;
zPtr->v64 = uiA<<12;
@@ -42,6 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_normRoundPackToExtF80(
struct softfloat_state *state,
bool sign,
int_fast32_t exp,
uint_fast64_t sig,
@@ -66,7 +67,7 @@ extFloat80_t
}
return
softfloat_roundPackToExtF80(
sign, exp, sig, sigExtra, roundingPrecision );
state, sign, exp, sig, sigExtra, roundingPrecision );
}
@@ -53,6 +53,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -76,7 +77,7 @@ struct uint128
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( isSigNaNA | isSigNaNB ) {
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
if ( isSigNaNA ) {
if ( isSigNaNB ) goto returnLargerMag;
if ( isNaNExtF80UI( uiB64, uiB0 ) ) goto returnB;
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t
softfloat_roundPackToExtF80(
struct softfloat_state *state,
bool sign,
int_fast32_t exp,
uint_fast64_t sig,
@@ -59,7 +60,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
if ( roundingPrecision == 80 ) goto precision80;
if ( roundingPrecision == 64 ) {
@@ -87,15 +88,15 @@ extFloat80_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess
(state->detectTininess
== softfloat_tininess_beforeRounding)
|| (exp < 0)
|| (sig <= (uint64_t) (sig + roundIncrement));
sig = softfloat_shiftRightJam64( sig, 1 - exp );
roundBits = sig & roundMask;
if ( roundBits ) {
if ( isTiny ) softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( isTiny ) softfloat_raiseFlags( state, softfloat_flag_underflow );
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= roundMask + 1;
@@ -121,7 +122,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig = (sig & ~roundMask) | (roundMask + 1);
@@ -157,7 +158,7 @@ extFloat80_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess
(state->detectTininess
== softfloat_tininess_beforeRounding)
|| (exp < 0)
|| ! doIncrement
@@ -168,8 +169,8 @@ extFloat80_t
sig = sig64Extra.v;
sigExtra = sig64Extra.extra;
if ( sigExtra ) {
if ( isTiny ) softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( isTiny ) softfloat_raiseFlags( state, softfloat_flag_underflow );
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -207,7 +208,7 @@ extFloat80_t
roundMask = 0;
overflow:
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
if (
roundNearEven
|| (roundingMode == softfloat_round_near_maxMag)
@@ -226,7 +227,7 @@ extFloat80_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( sigExtra ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
float32_t
softfloat_roundPackToF32( bool sign, int_fast16_t exp, uint_fast32_t sig )
softfloat_roundPackToF32( struct softfloat_state *state, bool sign, int_fast16_t exp, uint_fast32_t sig )
{
uint_fast8_t roundingMode;
bool roundNearEven;
@@ -53,7 +53,7 @@ float32_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
roundIncrement = 0x40;
if ( ! roundNearEven && (roundingMode != softfloat_round_near_maxMag) ) {
@@ -71,19 +71,19 @@ float32_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess == softfloat_tininess_beforeRounding)
(state->detectTininess == softfloat_tininess_beforeRounding)
|| (exp < -1) || (sig + roundIncrement < 0x80000000);
sig = softfloat_shiftRightJam32( sig, -exp );
exp = 0;
roundBits = sig & 0x7F;
if ( isTiny && roundBits ) {
softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_raiseFlags( state, softfloat_flag_underflow );
}
} else if ( (0xFD < exp) || (0x80000000 <= sig + roundIncrement) ) {
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
uiZ = packToF32UI( sign, 0xFF, 0 ) - ! roundIncrement;
goto uiZ;
}
@@ -92,7 +92,7 @@ float32_t
*------------------------------------------------------------------------*/
sig = (sig + roundIncrement)>>7;
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -42,7 +42,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
float64_t
softfloat_roundPackToF64( bool sign, int_fast16_t exp, uint_fast64_t sig )
softfloat_roundPackToF64( struct softfloat_state *state, bool sign, int_fast16_t exp, uint_fast64_t sig )
{
uint_fast8_t roundingMode;
bool roundNearEven;
@@ -53,7 +53,7 @@ float64_t
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
roundingMode = softfloat_roundingMode;
roundingMode = state->roundingMode;
roundNearEven = (roundingMode == softfloat_round_near_even);
roundIncrement = 0x200;
if ( ! roundNearEven && (roundingMode != softfloat_round_near_maxMag) ) {
@@ -71,14 +71,14 @@ float64_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
isTiny =
(softfloat_detectTininess == softfloat_tininess_beforeRounding)
(state->detectTininess == softfloat_tininess_beforeRounding)
|| (exp < -1)
|| (sig + roundIncrement < UINT64_C( 0x8000000000000000 ));
sig = softfloat_shiftRightJam64( sig, -exp );
exp = 0;
roundBits = sig & 0x3FF;
if ( isTiny && roundBits ) {
softfloat_raiseFlags( softfloat_flag_underflow );
softfloat_raiseFlags( state, softfloat_flag_underflow );
}
} else if (
(0x7FD < exp)
@@ -87,7 +87,7 @@ float64_t
/*----------------------------------------------------------------
*----------------------------------------------------------------*/
softfloat_raiseFlags(
softfloat_flag_overflow | softfloat_flag_inexact );
state, softfloat_flag_overflow | softfloat_flag_inexact );
uiZ = packToF64UI( sign, 0x7FF, 0 ) - ! roundIncrement;
goto uiZ;
}
@@ -96,7 +96,7 @@ float64_t
*------------------------------------------------------------------------*/
sig = (sig + roundIncrement)>>10;
if ( roundBits ) {
softfloat_exceptionFlags |= softfloat_flag_inexact;
state->exceptionFlags |= softfloat_flag_inexact;
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) {
sig |= 1;
@@ -44,7 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t
softfloat_roundToI32(
bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
struct softfloat_state *state, bool sign, uint_fast64_t sig, uint_fast8_t roundingMode, bool exact )
{
uint_fast16_t roundIncrement, roundBits;
uint_fast32_t sig32;
@@ -86,13 +86,13 @@ int_fast32_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) z |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? i32_fromNegOverflow : i32_fromPosOverflow;
}
@@ -44,6 +44,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t
softfloat_roundToI64(
struct softfloat_state *state,
bool sign,
uint_fast64_t sig,
uint_fast64_t sigExtra,
@@ -89,13 +90,13 @@ int_fast64_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) z |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? i64_fromNegOverflow : i64_fromPosOverflow;
}
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
uint_fast64_t
softfloat_roundToUI64(
struct softfloat_state *state,
bool sign,
uint_fast64_t sig,
uint_fast64_t sigExtra,
@@ -84,13 +85,13 @@ uint_fast64_t
#ifdef SOFTFLOAT_ROUND_ODD
if ( roundingMode == softfloat_round_odd ) sig |= 1;
#endif
if ( exact ) softfloat_exceptionFlags |= softfloat_flag_inexact;
if ( exact ) state->exceptionFlags |= softfloat_flag_inexact;
}
return sig;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
return sign ? ui64_fromNegOverflow : ui64_fromPosOverflow;
}
@@ -43,6 +43,7 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
extFloat80_t
softfloat_subMagsExtF80(
struct softfloat_state *state,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -77,7 +78,7 @@ extFloat80_t
if ( (sigA | sigB) & UINT64_C( 0x7FFFFFFFFFFFFFFF ) ) {
goto propagateNaN;
}
softfloat_raiseFlags( softfloat_flag_invalid );
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ64 = defaultNaNExtF80UI64;
uiZ0 = defaultNaNExtF80UI0;
goto uiZ;
@@ -90,7 +91,7 @@ extFloat80_t
if ( sigB < sigA ) goto aBigger;
if ( sigA < sigB ) goto bBigger;
uiZ64 =
packToExtF80UI64( (softfloat_roundingMode == softfloat_round_min), 0 );
packToExtF80UI64( (state->roundingMode == softfloat_round_min), 0 );
uiZ0 = 0;
goto uiZ;
/*------------------------------------------------------------------------
@@ -142,11 +143,11 @@ extFloat80_t
normRoundPack:
return
softfloat_normRoundPackToExtF80(
signZ, expZ, sig128.v64, sig128.v0, extF80_roundingPrecision );
state, signZ, expZ, sig128.v64, sig128.v0, state->roundingPrecision );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNExtF80UI( uiA64, uiA0, uiB64, uiB0 );
uiZ = softfloat_propagateNaNExtF80UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ64 = uiZ.v64;
uiZ0 = uiZ.v0;
uiZ:
+19 -64
View File
@@ -50,50 +50,11 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
#include <stdint.h>
#include "softfloat_types.h"
#ifndef THREAD_LOCAL
#define THREAD_LOCAL
#endif
/*----------------------------------------------------------------------------
| Software floating-point underflow tininess-detection mode.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_detectTininess;
enum {
softfloat_tininess_beforeRounding = 0,
softfloat_tininess_afterRounding = 1
};
/*----------------------------------------------------------------------------
| Software floating-point rounding mode. (Mode "odd" is supported only if
| SoftFloat is compiled with macro 'SOFTFLOAT_ROUND_ODD' defined.)
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_roundingMode;
enum {
softfloat_round_near_even = 0,
softfloat_round_minMag = 1,
softfloat_round_min = 2,
softfloat_round_max = 3,
softfloat_round_near_maxMag = 4,
softfloat_round_odd = 6
};
/*----------------------------------------------------------------------------
| Software floating-point exception flags.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t softfloat_exceptionFlags;
enum {
softfloat_flag_inexact = 1,
softfloat_flag_underflow = 2,
softfloat_flag_overflow = 4,
softfloat_flag_infinite = 8,
softfloat_flag_invalid = 16
};
/*----------------------------------------------------------------------------
| Routine to raise any or all of the software floating-point exception flags.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t );
void softfloat_raiseFlags( struct softfloat_state *, uint_fast8_t );
/*----------------------------------------------------------------------------
| Integer-to-floating-point conversion routines.
@@ -187,7 +148,7 @@ float16_t f32_to_f16( float32_t );
float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( float32_t );
extFloat80_t f32_to_extF80( struct softfloat_state *, float32_t );
float128_t f32_to_f128( float32_t );
#endif
void f32_to_extF80M( float32_t, extFloat80_t * );
@@ -223,7 +184,7 @@ float16_t f64_to_f16( float64_t );
float32_t f64_to_f32( float64_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f64_to_extF80( float64_t );
extFloat80_t f64_to_extF80( struct softfloat_state *, float64_t );
float128_t f64_to_f128( float64_t );
#endif
void f64_to_extF80M( float64_t, extFloat80_t * );
@@ -244,53 +205,47 @@ bool f64_le_quiet( float64_t, float64_t );
bool f64_lt_quiet( float64_t, float64_t );
bool f64_isSignalingNaN( float64_t );
/*----------------------------------------------------------------------------
| Rounding precision for 80-bit extended double-precision floating-point.
| Valid values are 32, 64, and 80.
*----------------------------------------------------------------------------*/
extern THREAD_LOCAL uint_fast8_t extF80_roundingPrecision;
/*----------------------------------------------------------------------------
| 80-bit extended double-precision floating-point operations.
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t extF80_to_ui32( extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
uint_fast64_t extF80_to_ui64( extFloat80_t, uint_fast8_t, bool );
uint_fast64_t extF80_to_ui64( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast32_t extF80_to_i32( extFloat80_t, uint_fast8_t, bool );
int_fast32_t extF80_to_i32( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
int_fast64_t extF80_to_i64( extFloat80_t, uint_fast8_t, bool );
int_fast64_t extF80_to_i64( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
uint_fast32_t extF80_to_ui32_r_minMag( extFloat80_t, bool );
uint_fast64_t extF80_to_ui64_r_minMag( extFloat80_t, bool );
int_fast32_t extF80_to_i32_r_minMag( extFloat80_t, bool );
int_fast64_t extF80_to_i64_r_minMag( extFloat80_t, bool );
float16_t extF80_to_f16( extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float32_t extF80_to_f32( extFloat80_t );
float32_t extF80_to_f32( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float64_t extF80_to_f64( extFloat80_t );
float64_t extF80_to_f64( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
float128_t extF80_to_f128( extFloat80_t );
float128_t extF80_to_f128( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_roundToInt( extFloat80_t, uint_fast8_t, bool );
extFloat80_t extF80_roundToInt( struct softfloat_state *, extFloat80_t, uint_fast8_t, bool );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_add( extFloat80_t, extFloat80_t );
extFloat80_t extF80_add( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sub( extFloat80_t, extFloat80_t );
extFloat80_t extF80_sub( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_mul( extFloat80_t, extFloat80_t );
extFloat80_t extF80_mul( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_div( extFloat80_t, extFloat80_t );
extFloat80_t extF80_div( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_rem( extFloat80_t, extFloat80_t );
extFloat80_t extF80_rem( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( extFloat80_t );
extFloat80_t extF80_sqrt( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( extFloat80_t, extFloat80_t );
bool extF80_eq( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( extFloat80_t, extFloat80_t );
bool extF80_lt( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
bool extF80_lt_quiet( extFloat80_t, extFloat80_t );
@@ -341,7 +296,7 @@ float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( float128_t );
extFloat80_t f128_to_extF80( struct softfloat_state *, float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
float128_t f128_sub( float128_t, float128_t );
@@ -44,10 +44,10 @@ SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
| should be simply `softfloat_exceptionFlags |= flags;'.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_raiseFlags( uint_fast8_t flags )
void softfloat_raiseFlags( struct softfloat_state *state, uint_fast8_t flags )
{
softfloat_exceptionFlags |= flags;
state->exceptionFlags |= flags;
}
@@ -1,52 +0,0 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All Rights Reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
#ifndef THREAD_LOCAL
#define THREAD_LOCAL
#endif
THREAD_LOCAL uint_fast8_t softfloat_roundingMode = softfloat_round_near_even;
THREAD_LOCAL uint_fast8_t softfloat_detectTininess = init_detectTininess;
THREAD_LOCAL uint_fast8_t softfloat_exceptionFlags = 0;
THREAD_LOCAL uint_fast8_t extF80_roundingPrecision = 80;
@@ -77,5 +77,50 @@ struct extFloat80M { uint16_t signExp; uint64_t signif; };
*----------------------------------------------------------------------------*/
typedef struct extFloat80M extFloat80_t;
enum {
softfloat_tininess_beforeRounding = 0,
softfloat_tininess_afterRounding = 1
};
enum {
softfloat_round_near_even = 0,
softfloat_round_minMag = 1,
softfloat_round_min = 2,
softfloat_round_max = 3,
softfloat_round_near_maxMag = 4,
softfloat_round_odd = 6
};
enum {
softfloat_flag_inexact = 1,
softfloat_flag_underflow = 2,
softfloat_flag_overflow = 4,
softfloat_flag_infinite = 8,
softfloat_flag_invalid = 16
};
struct softfloat_state {
/*----------------------------------------------------------------------------
| Software floating-point underflow tininess-detection mode.
*----------------------------------------------------------------------------*/
uint8_t detectTininess; /* = init_detectTininess */
/*----------------------------------------------------------------------------
| Software floating-point rounding mode. (Mode "odd" is supported only if
| SoftFloat is compiled with macro 'SOFTFLOAT_ROUND_ODD' defined.)
*----------------------------------------------------------------------------*/
uint8_t roundingMode; /* = softfloat_round_near_even */
/*----------------------------------------------------------------------------
| Software floating-point exception flags.
*----------------------------------------------------------------------------*/
uint8_t exceptionFlags; /* = 0 */
/*----------------------------------------------------------------------------
| Rounding precision for 80-bit extended double-precision floating-point.
| Valid values are 32, 64, and 80.
*----------------------------------------------------------------------------*/
uint8_t roundingPrecision; /* = 80 */
};
#endif
@@ -136,7 +136,7 @@ uint_fast16_t
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f32UIToCommonNaN( uint_fast32_t uiA, struct commonNaN *zPtr );
void softfloat_f32UIToCommonNaN( struct softfloat_state *, uint_fast32_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 32-bit floating-point
@@ -173,7 +173,7 @@ uint_fast32_t
| exception is raised.
*----------------------------------------------------------------------------*/
FEXCORE_PRESERVE_ALL_ATTR
void softfloat_f64UIToCommonNaN( uint_fast64_t uiA, struct commonNaN *zPtr );
void softfloat_f64UIToCommonNaN( struct softfloat_state *, uint_fast64_t uiA, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 64-bit floating-point
@@ -222,7 +222,7 @@ uint_fast64_t
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_extF80UIToCommonNaN(
uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
struct softfloat_state *, uint_fast16_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into an 80-bit extended
@@ -244,6 +244,7 @@ struct uint128 softfloat_commonNaNToExtF80UI( const struct commonNaN *aPtr );
FEXCORE_PRESERVE_ALL_ATTR
struct uint128
softfloat_propagateNaNExtF80UI(
struct softfloat_state *,
uint_fast16_t uiA64,
uint_fast64_t uiA0,
uint_fast16_t uiB64,
@@ -274,7 +275,7 @@ struct uint128
FEXCORE_PRESERVE_ALL_ATTR
void
softfloat_f128UIToCommonNaN(
uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
struct softfloat_state *, uint_fast64_t uiA64, uint_fast64_t uiA0, struct commonNaN *zPtr );
/*----------------------------------------------------------------------------
| Converts the common NaN pointed to by 'aPtr' into a 128-bit floating-point
+73 -79
View File
@@ -63,7 +63,7 @@ struct FEX_PACKED X80SoftFloat {
}
// Ops
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FADD(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FADD(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -79,11 +79,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_add(lhs, rhs);
return extF80_add(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSUB(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSUB(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -99,11 +99,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_sub(lhs, rhs);
return extF80_sub(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FMUL(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FMUL(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -119,11 +119,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_mul(lhs, rhs);
return extF80_mul(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FDIV(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FDIV(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -139,11 +139,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_div(lhs, rhs);
return extF80_div(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm(R"(
@@ -160,11 +160,11 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_rem(lhs, rhs);
return extF80_rem(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM1(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FREM1(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
#if defined(DEBUG_X86_FLOAT)
BIGFLOAT Result;
asm(R"(
@@ -181,16 +181,16 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_rem(lhs, rhs);
return extF80_rem(state, lhs, rhs);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs) {
return extF80_roundToInt(lhs, softfloat_roundingMode, false);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(softfloat_state* state, const X80SoftFloat& lhs) {
return extF80_roundToInt(state, lhs, state->roundingMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(const X80SoftFloat& lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(lhs, RoundMode, false);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FRNDINT(softfloat_state* state, const X80SoftFloat& lhs, uint_fast8_t RoundMode) {
return extF80_roundToInt(state, lhs, RoundMode, false);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FXTRACT_SIG(const X80SoftFloat& lhs) {
@@ -237,13 +237,14 @@ struct FEX_PACKED X80SoftFloat {
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static void FCMP(const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(lhs, rhs);
*lt = extF80_lt(lhs, rhs);
FEXCORE_PRESERVE_ALL_ATTR static void
FCMP(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs, bool* eq, bool* lt, bool* nan) {
*eq = extF80_eq(state, lhs, rhs);
*lt = extF80_lt(state, lhs, rhs);
*nan = IsNan(lhs) || IsNan(rhs);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSCALE(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FSCALE which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -261,16 +262,16 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
X80SoftFloat Int = FRNDINT(rhs, softfloat_round_minMag);
LIBRARY_PRECISION Src2_d = Int;
X80SoftFloat Int = FRNDINT(state, rhs, softfloat_round_minMag);
LIBRARY_PRECISION Src2_d = Int.ToFMax(state);
Src2_d = exp2l(Src2_d);
X80SoftFloat Src2_X80 = Src2_d;
X80SoftFloat Result = extF80_mul(lhs, Src2_X80);
X80SoftFloat Src2_X80(state, Src2_d);
X80SoftFloat Result = extF80_mul(state, lhs, Src2_X80);
return Result;
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat F2XM1(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat F2XM1(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used F2XM1 which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -286,14 +287,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Result = exp2l(Src1_d);
Result -= 1.0;
return Result;
return X80SoftFloat(state, Result);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FYL2X(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FYL2X(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FYL2X which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -310,14 +311,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Src2_d = rhs.ToFMax(state);
LIBRARY_PRECISION Tmp = Src2_d * log2l(Src1_d);
return Tmp;
return X80SoftFloat(state, Tmp);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FATAN(const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FATAN(softfloat_state* state, const X80SoftFloat& lhs, const X80SoftFloat& rhs) {
WARN_ONCE_FMT("x87: Application used FATAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -334,14 +335,14 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src1_d = lhs;
LIBRARY_PRECISION Src2_d = rhs;
LIBRARY_PRECISION Src1_d = lhs.ToFMax(state);
LIBRARY_PRECISION Src2_d = rhs.ToFMax(state);
LIBRARY_PRECISION Tmp = atan2l(Src1_d, Src2_d);
return Tmp;
return X80SoftFloat(state, Tmp);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FTAN(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FTAN(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FTAN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -358,13 +359,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = tanl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSIN(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSIN(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FSIN which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -380,13 +381,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = sinl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FCOS(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FCOS(softfloat_state* state, const X80SoftFloat& lhs) {
WARN_ONCE_FMT("x87: Application used FCOS which may have accuracy problems");
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
@@ -402,13 +403,13 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
LIBRARY_PRECISION Src_d = lhs;
LIBRARY_PRECISION Src_d = lhs.ToFMax(state);
Src_d = cosl(Src_d);
return Src_d;
return X80SoftFloat(state, Src_d);
#endif
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSQRT(const X80SoftFloat& lhs) {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat FSQRT(softfloat_state* state, const X80SoftFloat& lhs) {
#ifdef DEBUG_X86_FLOAT
BIGFLOAT Result;
asm(R"(
@@ -423,62 +424,55 @@ struct FEX_PACKED X80SoftFloat {
return Result;
#else
return extF80_sqrt(lhs);
return extF80_sqrt(state, lhs);
#endif
}
operator float() const {
const float32_t Result = extF80_to_f32(*this);
float ToF32(softfloat_state* state) const {
const float32_t Result = extF80_to_f32(state, *this);
return FEXCore::BitCast<float>(Result);
}
operator double() const {
const float64_t Result = extF80_to_f64(*this);
double ToF64(softfloat_state* state) const {
const float64_t Result = extF80_to_f64(state, *this);
return FEXCore::BitCast<double>(Result);
}
#ifndef _WIN32
operator BIGFLOAT() const {
LIBRARY_PRECISION ToFMax(softfloat_state* state) const {
#ifdef _WIN32
return ToF64(state);
#else
#if BIGFLOATSIZE == 16
const float128_t Result = extF80_to_f128(*this);
const float128_t Result = extF80_to_f128(state, *this);
return FEXCore::BitCast<BIGFLOAT>(Result);
#else
BIGFLOAT result {};
memcpy(&result, this, sizeof(result));
return result;
#endif
}
#endif
}
operator int16_t() const {
auto rv = extF80_to_i32(*this, softfloat_roundingMode, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
int16_t ToI16(softfloat_state* state) const {
auto rv = extF80_to_i32(state, *this, state->roundingMode, false);
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
return INT16_MIN;
} else {
return rv;
}
}
operator int32_t() const {
return extF80_to_i32(*this, softfloat_roundingMode, false);
int32_t ToI32(softfloat_state* state) const {
return extF80_to_i32(state, *this, state->roundingMode, false);
}
operator int64_t() const {
return extF80_to_i64(*this, softfloat_roundingMode, false);
int64_t ToI64(softfloat_state* state) const {
return extF80_to_i64(state, *this, state->roundingMode, false);
}
operator uint64_t() const {
return extF80_to_ui64(*this, softfloat_roundingMode, false);
}
void operator=(const float rhs) {
*this = f32_to_extF80(FEXCore::BitCast<float32_t>(rhs));
}
void operator=(const double rhs) {
*this = f64_to_extF80(FEXCore::BitCast<float64_t>(rhs));
uint64_t ToUI64(softfloat_state* state) const {
return extF80_to_ui64(state, *this, state->roundingMode, false);
}
void operator=(const int16_t rhs) {
@@ -509,18 +503,18 @@ struct FEX_PACKED X80SoftFloat {
Sign = rhs.signExp >> 15;
}
X80SoftFloat(const float rhs) {
*this = f32_to_extF80(FEXCore::BitCast<float32_t>(rhs));
X80SoftFloat(softfloat_state* state, const float rhs) {
*this = f32_to_extF80(state, FEXCore::BitCast<float32_t>(rhs));
}
X80SoftFloat(const double rhs) {
*this = f64_to_extF80(FEXCore::BitCast<float64_t>(rhs));
X80SoftFloat(softfloat_state* state, const double rhs) {
*this = f64_to_extF80(state, FEXCore::BitCast<float64_t>(rhs));
}
#ifndef _WIN32
X80SoftFloat(BIGFLOAT rhs) {
X80SoftFloat(softfloat_state* state, BIGFLOAT rhs) {
#if BIGFLOATSIZE == 16
*this = f128_to_extF80(FEXCore::BitCast<float128_t>(rhs));
*this = f128_to_extF80(state, FEXCore::BitCast<float128_t>(rhs));
#else
*this = FEXCore::BitCast<long double>(rhs);
#endif
@@ -76,6 +76,8 @@
"DISABLECRYPTO": "disablecrypto",
"ENABLERPRES": "enablerpres",
"DISABLERPRES": "disablerpres",
"ENABLESVEBITPERM": "enablesvebitperm",
"DISABLESVEBITPERM": "disablesvebitperm",
"ENABLEPRESERVEALLABI": "enablepreserveallabi",
"DISABLEPRESERVEALLABI": "disablepreserveallabi"
},
@@ -97,6 +99,7 @@
"\t{enable,disable}flagm2: Will force enable or disable flagm2 even if the host doesn't support it",
"\t{enable,disable}crypto: Will force enable or disable crypto extensions even if the host doesn't support it",
"\t{enable,disable}rpres: Will force enable or disable rpres even if the host doesn't support it",
"\t{enable,disable}svebitperm: Will force enable or disable svebitperm even if the host doesn't support it",
"\t{enable,disable}preserveallabi: Will force enable or disable preserve_all abi even if the host doesn't support it"
]
},
+3 -6
View File
@@ -6,6 +6,7 @@
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/Context.h>
#include <FEXCore/Core/CPUID.h>
#include <FEXCore/Core/HostFeatures.h>
#include <FEXCore/Core/SignalDelegator.h>
#include "FEXCore/Debug/InternalThreadState.h"
@@ -18,8 +19,8 @@ void InitializeStaticTables(OperatingMode Mode) {
IR::InstallOpcodeHandlers(Mode);
}
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext() {
return fextl::make_unique<FEXCore::Context::ContextImpl>();
fextl::unique_ptr<FEXCore::Context::Context> FEXCore::Context::Context::CreateNewContext(const FEXCore::HostFeatures& Features) {
return fextl::make_unique<FEXCore::Context::ContextImpl>(Features);
}
void FEXCore::Context::ContextImpl::SetExitHandler(ExitHandler handler) {
@@ -42,10 +43,6 @@ void FEXCore::Context::ContextImpl::SetCustomCPUBackendFactory(CustomCPUFactoryT
CustomCPUFactory = std::move(Factory);
}
HostFeatures FEXCore::Context::ContextImpl::GetHostFeatures() const {
return HostFeatures;
}
void FEXCore::Context::ContextImpl::SetSignalDelegator(FEXCore::SignalDelegator* _SignalDelegation) {
SignalDelegation = _SignalDelegation;
}
+2 -9
View File
@@ -96,8 +96,6 @@ public:
void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) override;
HostFeatures GetHostFeatures() const override;
void HandleCallback(FEXCore::Core::InternalThreadState* Thread, uint64_t RIP) override;
uint64_t RestoreRIPFromHostPC(FEXCore::Core::InternalThreadState* Thread, uint64_t HostPC) override;
@@ -195,7 +193,7 @@ public:
// returns false if a handler was already registered
CustomIRResult AddCustomIREntrypoint(uintptr_t Entrypoint, CustomIREntrypointHandler Handler, void* Creator = nullptr, void* Data = nullptr);
void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) override;
void AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) override;
public:
friend class FEXCore::HLE::SyscallHandler;
@@ -264,7 +262,7 @@ public:
SignalDelegator* SignalDelegation {};
X86GeneratedCode X86CodeGen;
ContextImpl();
ContextImpl(const FEXCore::HostFeatures& Features);
~ContextImpl();
static void ThreadRemoveCodeEntry(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
@@ -283,9 +281,6 @@ public:
// Must be called from owning thread
static void ThreadRemoveCodeEntryFromJit(FEXCore::Core::CpuStateFrame* Frame, uint64_t GuestRIP) {
auto Thread = Frame->Thread;
LOGMAN_THROW_A_FMT(Thread->ThreadManager.GetTID() == FHU::Syscalls::gettid(), "Must be called from owning thread {}, not {}",
Thread->ThreadManager.GetTID(), FHU::Syscalls::gettid());
auto lk = GuardSignalDeferringSection(static_cast<ContextImpl*>(Thread->CTX)->CodeInvalidationMutex, Thread);
ThreadRemoveCodeEntry(Thread, GuestRIP);
@@ -378,8 +373,6 @@ public:
return ExitOnHLT;
}
FEXCore::CPU::CPUBackendFeatures BackendFeatures;
protected:
void UpdateAtomicTSOEmulationConfig() {
if (SupportsHardwareTSO) {
@@ -14,10 +14,12 @@
#include <CodeEmitter/Emitter.h>
#include <CodeEmitter/Registers.h>
#ifdef VIXL_DISASSEMBLER
#include <aarch64/cpu-aarch64.h>
#include <aarch64/instructions-aarch64.h>
#include <cpu-features.h>
#include <utils-vixl.h>
#endif
#include <array>
#include <tuple>
@@ -349,8 +351,6 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr
}
#endif
CPU.SetUp();
// Number of register available is dependent on what operating mode the proccess is in.
if (EmitterCTX->Config.Is64BitMode()) {
StaticRegisters = x64::SRA;
@@ -421,7 +421,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
if (RequiredMoveSegments > 1) {
// Only try to use this path if the number of segments is > 1.
// `movz` is better than `orr` since hardware will rename or merge if possible when `movz` is used.
const auto IsImm = vixl::aarch64::Assembler::IsImmLogical(Constant, RegSizeInBits(s));
const auto IsImm = ARMEmitter::Emitter::IsImmLogical(Constant, RegSizeInBits(s));
if (IsImm) {
orr(s, Reg, ARMEmitter::Reg::zr, Constant);
if (NOPPad) {
@@ -458,7 +458,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
// If the aligned offset is within the 4GB window then we can use ADRP+ADD
// and the number of move segments more than 1
if (RequiredMoveSegments > 1 && vixl::IsInt32(AlignedOffset)) {
if (RequiredMoveSegments > 1 && ARMEmitter::Emitter::IsInt32(AlignedOffset)) {
// If this is 4k page aligned then we only need ADRP
if ((AlignedOffset & 0xFFF) == 0) {
adrp(Reg, AlignedOffset >> 12);
@@ -466,7 +466,7 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
// If the constant is within 1MB of PC then we can still use ADR to load in a single instruction
// 21-bit signed integer here
int64_t SmallOffset = static_cast<int64_t>(Constant) - static_cast<int64_t>(PC);
if (vixl::IsInt21(SmallOffset)) {
if (ARMEmitter::Emitter::IsInt21(SmallOffset)) {
adr(Reg, SmallOffset);
} else {
// Need to use ADRP + ADD
@@ -570,12 +570,53 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
}
}
void Arm64Emitter::FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
// Enable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
orr(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
if (SetFIZ) {
// Insert MXCSR.DAZ in to FIZ
ldr(TmpReg2.W(), STATE.R(), offsetof(FEXCore::Core::CPUState, mxcsr));
bfxil(ARMEmitter::Size::i64Bit, TmpReg, TmpReg2, 6, 1);
}
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
if (SetPredRegs) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
if (EmitterCTX->HostFeatures.SupportsSVE256) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
}
if (EmitterCTX->HostFeatures.SupportsSVE128) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
}
}
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Disable AFP features when spilling registers.
//
// Disable FPCR.NEP and FPCR.AH
// Disable FPCR.NEP and FPCR.AH and FPCR.FIZ
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
// Also interacts with RPRES to change reciprocal/rsqrt precision from 8-bit mantissa to 12-bit.
@@ -585,7 +626,8 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
bic(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
(1U << 1) | // AH
(1U << 0)); // FIZ
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
}
#endif
@@ -663,37 +705,37 @@ void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint3
}
}
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
ARMEmitter::Register TmpReg = ARMEmitter::Reg::r0;
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
[[maybe_unused]] bool FoundRegister {};
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & GPRFillMask)) {
TmpReg = Reg;
FoundRegister = true;
break;
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask, std::optional<ARMEmitter::Register> OptionalReg,
std::optional<ARMEmitter::Register> OptionalReg2) {
auto FindTempReg = [this](uint32_t* GPRFillMask) -> std::optional<ARMEmitter::Register> {
for (auto Reg : StaticRegisters) {
if (((1U << Reg.Idx()) & *GPRFillMask)) {
*GPRFillMask &= ~(1U << Reg.Idx());
return std::make_optional(Reg);
}
}
return std::nullopt;
};
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 2 GPRs for a temp");
uint32_t TempGPRFillMask = GPRFillMask;
if (!OptionalReg.has_value()) {
OptionalReg = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(FoundRegister, "Didn't have an SRA register to use as a temporary while spilling!");
#ifndef VIXL_SIMULATOR
if (EmitterCTX->HostFeatures.SupportsAFP) {
// Enable AFP features when filling JIT state.
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
mrs(TmpReg, ARMEmitter::SystemRegister::FPCR);
// Enable FPCR.NEP and FPCR.AH
// NEP(2): Changes ASIMD scalar instructions to insert in to the lower bits of the destination.
// AH(1): Changes NaN behaviour in some instructions. Specifically fmin, fmax.
//
// Additional interesting AFP bits:
// FIZ(0): Flush Inputs to Zero
orr(ARMEmitter::Size::i64Bit, TmpReg, TmpReg,
(1U << 2) | // NEP
(1U << 1)); // AH
msr(ARMEmitter::SystemRegister::FPCR, TmpReg);
if (!OptionalReg2.has_value()) {
OptionalReg2 = FindTempReg(&TempGPRFillMask);
}
LOGMAN_THROW_A_FMT(OptionalReg.has_value() && OptionalReg2.has_value(), "Didn't have an SRA register to use as a temporary while "
"spilling!");
auto TmpReg = *OptionalReg;
auto TmpReg2 = *OptionalReg2;
#ifdef _M_ARM_64EC
// Load STATE in from the CPU area as x28 is not callee saved in the ARM64EC ABI.
ldr(TmpReg.X(), ARMEmitter::Reg::r18, TEB_CPU_AREA_OFFSET);
ldr(STATE, TmpReg, CPU_AREA_EMULATOR_DATA_OFFSET);
#endif
// Regardless of what GPRs/FPRs we're filling, we need to fill NZCV since it
@@ -704,19 +746,9 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
ldr(TmpReg.W(), STATE.R(), offsetof(FEXCore::Core::CpuStateFrame, State.flags[24]));
msr(ARMEmitter::SystemRegister::NZCV, TmpReg);
FillSpecialRegs(TmpReg, TmpReg2, true, FPRs);
if (FPRs) {
// Set up predicate registers.
// We don't bother spilling these in SpillStaticRegs,
// since all that matters is we restore them on a fill.
// It's not a concern if they get trounced by something else.
if (EmitterCTX->HostFeatures.SupportsSVE128) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
if (EmitterCTX->HostFeatures.SupportsSVE256) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
}
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
const auto Reg = StaticFPRegisters[i];
@@ -70,6 +70,14 @@ constexpr auto VTMP2 = ARMEmitter::VReg::v17;
// Entry/Exit ABI
constexpr auto EC_CALL_CHECKER_PC_REG = ARMEmitter::XReg::x9;
constexpr auto EC_ENTRY_CPUAREA_REG = ARMEmitter::XReg::x17;
// These structures are not included in the standard Windows headers, define the offsets of members we care about for EC here.
constexpr size_t TEB_CPU_AREA_OFFSET = 0x1788;
constexpr size_t TEB_PEB_OFFSET = 0x60;
constexpr size_t PEB_EC_CODE_BITMAP_OFFSET = 0x368;
constexpr size_t CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET = 0x1;
constexpr size_t CPU_AREA_EMULATOR_STACK_BASE_OFFSET = 0x8;
constexpr size_t CPU_AREA_EMULATOR_DATA_OFFSET = 0x30;
#endif
// Predicate register temporaries (used when AVX support is enabled)
@@ -86,7 +94,6 @@ protected:
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
FEXCore::Context::ContextImpl* EmitterCTX;
vixl::aarch64::CPU CPU;
std::span<const ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
std::span<const ARMEmitter::Register> StaticRegisters {};
@@ -97,12 +104,15 @@ protected:
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
void FillSpecialRegs(ARMEmitter::Register TmpReg, ARMEmitter::Register TmpReg2, bool SetFIZ, bool SetPredRegs);
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
// TMP4 is left alone.
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U,
std::optional<ARMEmitter::Register> OptionalReg = std::nullopt,
std::optional<ARMEmitter::Register> OptionalReg2 = std::nullopt);
// Register 0-18 + 29 + 30 are caller saved
static constexpr uint32_t CALLER_GPR_MASK = 0b0110'0000'0000'0111'1111'1111'1111'1111U;
@@ -34,11 +34,6 @@ namespace CodeSerialize {
}
namespace CPU {
struct CPUBackendFeatures {
bool SupportsFlags = false;
bool SupportsVTBL2 = false;
};
class CPUBackend {
public:
struct CodeBuffer {
+12 -2
View File
@@ -34,6 +34,8 @@ namespace ProductNames {
static const char ARM_A76AE[] = "Cortex-A76AE";
static const char ARM_V1[] = "Neoverse V1";
static const char ARM_V2[] = "Neoverse V2";
static const char ARM_V3[] = "Neoverse V3";
static const char ARM_V3AE[] = "Neoverse V3AE";
static const char ARM_A77[] = "Cortex-A77";
static const char ARM_A78[] = "Cortex-A78";
static const char ARM_A78AE[] = "Cortex-A78AE";
@@ -41,13 +43,16 @@ namespace ProductNames {
static const char ARM_A710[] = "Cortex-A710";
static const char ARM_A715[] = "Cortex-A715";
static const char ARM_A720[] = "Cortex-A720";
static const char ARM_A725[] = "Cortex-A725";
static const char ARM_X1[] = "Cortex-X1";
static const char ARM_X1C[] = "Cortex-X1C";
static const char ARM_X2[] = "Cortex-X2";
static const char ARM_X3[] = "Cortex-X3";
static const char ARM_X4[] = "Cortex-X4";
static const char ARM_X925[] = "Cortex-X925";
static const char ARM_N1[] = "Neoverse N1";
static const char ARM_N2[] = "Neoverse N2";
static const char ARM_N3[] = "Neoverse N3";
static const char ARM_E1[] = "Neoverse E1";
static const char ARM_A35[] = "Cortex-A35";
static const char ARM_A53[] = "Cortex-A53";
@@ -142,12 +147,17 @@ void CPUIDEmu::SetupHostHybridFlag() {
// CPU priority order
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
// Relative list so things they will commonly end up in big.little configurations sort of relate
static constexpr std::array<CPUMIDR, 43> CPUMIDRs = {{
static constexpr std::array<CPUMIDR, 48> CPUMIDRs = {{
// Typically big CPU cores
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
{0x41, 0xd85, 1, ProductNames::ARM_X925}, // X925
{0x41, 0xd87, 1, ProductNames::ARM_A725}, // A725
{0x41, 0xd84, 1, ProductNames::ARM_V3}, // V3
{0x41, 0xd83, 1, ProductNames::ARM_V3AE}, // V3AE
{0x41, 0xd8e, 1, ProductNames::ARM_N3}, // N3
{0x41, 0xd82, 1, ProductNames::ARM_X4}, // X4
{0x41, 0xd81, 1, ProductNames::ARM_A720}, // A720
{0x41, 0xd4e, 1, ProductNames::ARM_X3}, // X3
@@ -628,7 +638,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
(0 << 21) | // Reserved
(0 << 22) | // Reserved
(1 << 23) | // CLFLUSHOPT instruction
(CTX->HostFeatures.SupportsCLWB << 24) | // CLWB instruction
(1 << 24) | // CLWB instruction
(0 << 25) | // Intel processor trace
(0 << 26) | // Reserved
(0 << 27) | // Reserved
+25 -39
View File
@@ -38,7 +38,6 @@ $end_info$
#include <FEXCore/Debug/InternalThreadState.h>
#include <FEXCore/HLE/SyscallHandler.h>
#include <FEXCore/HLE/SourcecodeResolver.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/Allocator.h>
#include <FEXCore/Utils/Event.h>
#include <FEXCore/Utils/File.h>
@@ -75,8 +74,9 @@ $end_info$
#include <xxhash.h>
namespace FEXCore::Context {
ContextImpl::ContextImpl()
: CPUID {this}
ContextImpl::ContextImpl(const FEXCore::HostFeatures& Features)
: HostFeatures {Features}
, CPUID {this}
, IRCaptureCache {this} {
#ifdef BLOCKSTATS
BlockData = std::make_unique<FEXCore::BlockSamplingData>();
@@ -309,14 +309,6 @@ void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState
bool ContextImpl::InitCore() {
// Initialize the CPU core signal handlers & DispatcherConfig
switch (Config.Core) {
case FEXCore::Config::CONFIG_IRJIT: BackendFeatures = FEXCore::CPU::GetArm64JITBackendFeatures(); break;
case FEXCore::Config::CONFIG_CUSTOM:
// Do nothing
break;
default: LogMan::Msg::EFmt("Unknown core configuration"); return false;
}
Dispatcher = FEXCore::CPU::Dispatcher::Create(this);
// Set up the SignalDelegator config since core is initialized.
@@ -373,14 +365,15 @@ void ContextImpl::HandleCallback(FEXCore::Core::InternalThreadState* Thread, uin
FEXCore::Context::ExitReason ContextImpl::RunUntilExit(FEXCore::Core::InternalThreadState* Thread) {
ExecutionThread(Thread);
while (true) {
auto reason = Thread->ExitReason;
// Don't return if a custom exit handling the exit
if (!CustomExitHandler || reason == ExitReason::EXIT_SHUTDOWN) {
return reason;
}
CoreShuttingDown.store(true);
if (CustomExitHandler) {
CustomExitHandler(Thread, FEXCore::Context::ExitReason::EXIT_SHUTDOWN);
return Thread->ExitReason;
}
return FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
}
void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
@@ -390,9 +383,6 @@ void ContextImpl::ExecuteThread(FEXCore::Core::InternalThreadState* Thread) {
void ContextImpl::InitializeThreadTLSData(FEXCore::Core::InternalThreadState* Thread) {
// Let's do some initial bookkeeping here
Thread->ThreadManager.TID = FHU::Syscalls::gettid();
Thread->ThreadManager.PID = ::getpid();
if (ThunkHandler) {
ThunkHandler->RegisterTLSState(Thread);
}
@@ -446,7 +436,6 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, FEXCore::C
}
// Set up the thread manager state
Thread->ThreadManager.parent_tid = ParentTID;
Thread->CurrentFrame->Thread = Thread;
InitializeCompiler(Thread);
@@ -614,6 +603,20 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
DecodedInfo = &Block.DecodedInstructions[i];
bool IsLocked = DecodedInfo->Flags & FEXCore::X86Tables::DecodeFlags::FLAG_LOCK;
// Do a partial register cache flush before every instruction. This
// prevents cross-instruction static register caching, while allowing
// context load/stores to be optimized within a block. Theoretically,
// this flush is not required for correctness, all mandatory flushes are
// included in instruction-specific handlers. Instead, this is a blunt
// heuristic to make the register cache less aggressive, as the current
// RA generates bad code in common cases with tied registers otherwise.
//
// However, it makes our exception handling behaviour more predictable.
// It is potentially correctness bearing in that sense, but that is a
// side effect here and (if that behaviour is required) we should handle
// that more explicitly later.
Thread->OpDispatcher->FlushRegisterCache(true);
if (ExtendedDebugInfo || Thread->OpDispatcher->CanHaveSideEffects(TableInfo, DecodedInfo)) {
Thread->OpDispatcher->_GuestOpcode(Block.Entry + BlockInstructionsLength - GuestRIP);
}
@@ -804,14 +807,6 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
FEXCORE_PROFILE_SCOPED("CompileBlock");
auto Thread = Frame->Thread;
#ifdef _M_ARM_64EC
// If the target PC is EC code, mark it in the L2 and return straight to the dispatcher
// so it can handle the call/return.
if (Thread->LookupCache->CheckPageEC(GuestRIP)) {
return GuestRIP;
}
#endif
// Invalidate might take a unique lock on this, to guarantee that during invalidation no code gets compiled
auto lk = GuardSignalDeferringSection<std::shared_lock>(CodeInvalidationMutex, Thread);
@@ -918,15 +913,6 @@ void ContextImpl::ExecutionThread(FEXCore::Core::InternalThreadState* Thread) {
// If it is the parent thread that died then just leave
FEX_TODO("This doesn't make sense when the parent thread doesn't outlive its children");
if (Thread->ThreadManager.parent_tid == 0) {
CoreShuttingDown.store(true);
Thread->ExitReason = FEXCore::Context::ExitReason::EXIT_SHUTDOWN;
if (CustomExitHandler) {
CustomExitHandler(Thread->ThreadManager.TID, Thread->ExitReason);
}
}
#ifndef _WIN32
Alloc::OSAllocator::UninstallTLSData(Thread);
#endif
@@ -1021,7 +1007,7 @@ void ContextImpl::UnloadAOTIRCacheEntry(IR::AOTIRCacheEntry* Entry) {
IRCaptureCache.UnloadAOTIRCacheEntry(Entry);
}
void ContextImpl::AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) {
void ContextImpl::AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) {
if (ThunkHandler) {
ThunkHandler->AppendThunkDefinitions(Definitions);
}
@@ -62,14 +62,6 @@ void Dispatcher::EmitDispatcher() {
ARMEmitter::ForwardLabel l_CTX;
ARMEmitter::SingleUseForwardLabel l_Sleep;
#ifdef _M_ARM_64EC
// These structures are not included in the standard Windows headers, define them here
static constexpr size_t TEBCPUAreaOffset = 0x1788;
static constexpr size_t CPUAreaInSyscallCallbackOffset = 0x1;
static constexpr size_t CPUAreaEmulatorStackLimitOffset = 0x8;
static constexpr size_t CPUAreaEmulatorDataOffset = 0x30;
ARMEmitter::SingleUseForwardLabel ExitEC;
#endif
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
// Push all the register we need to save
@@ -94,7 +86,7 @@ void Dispatcher::EmitDispatcher() {
b(&LoopTop);
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
FillStaticRegs();
// Enter JIT
@@ -102,17 +94,15 @@ void Dispatcher::EmitDispatcher() {
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
// Load ThreadState and write the target PC there
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_DATA_OFFSET);
str(EC_CALL_CHECKER_PC_REG, STATE_PTR(CpuStateFrame, State.rip));
// Swap stacks to the emulator stack
ldr(TMP1, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorStackLimitOffset);
ldr(TMP1, EC_ENTRY_CPUAREA_REG, CPU_AREA_EMULATOR_STACK_BASE_OFFSET);
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
if (EmitterCTX->HostFeatures.SupportsSVE128) {
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
}
FillSpecialRegs(TMP1, TMP2, false, true);
// Enter JIT
#endif
@@ -168,10 +158,6 @@ void Dispatcher::EmitDispatcher() {
// If page pointer is zero then we have no block
cbz(ARMEmitter::Size::i64Bit, TMP1, &NoBlock);
#ifdef _M_ARM_64EC
// The LSB of an L2 page entry indicates if this page contains EC code
tbnz(TMP1, 0, &ExitEC);
#endif
// Steal the page offset
and_(ARMEmitter::Size::i64Bit, TMP2, TMP4, 0x0FFF);
@@ -198,23 +184,13 @@ void Dispatcher::EmitDispatcher() {
and_(ARMEmitter::Size::i64Bit, TMP2, RipReg.R(), LookupCache::L1_ENTRIES_MASK);
add(TMP1, TMP1, TMP2, ARMEmitter::ShiftType::LSL, 4);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, TMP3, TMP1);
stp<ARMEmitter::IndexType::OFFSET>(TMP4, RipReg, TMP1);
// Jump to the block
br(TMP4);
}
}
#ifdef _M_ARM_64EC
{
Bind(&ExitEC);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
}
#endif
{
ThreadStopHandlerAddressSpillSRA = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
@@ -232,14 +208,16 @@ void Dispatcher::EmitDispatcher() {
ExitFunctionLinkerAddress = GetCursorAddress<uint64_t>();
SpillStaticRegs(TMP1);
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
mov(ARMEmitter::XReg::x0, STATE);
@@ -259,10 +237,11 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP2, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
strb(ARMEmitter::WReg::zr, TMP2, CPUAreaInSyscallCallbackOffset);
ldr(TMP2, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP2, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
#ifndef _WIN32
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
@@ -270,6 +249,7 @@ void Dispatcher::EmitDispatcher() {
// Trigger segfault if any deferred signals are pending
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
br(TMP1);
}
@@ -278,20 +258,43 @@ void Dispatcher::EmitDispatcher() {
{
Bind(&NoBlock);
#ifdef _M_ARM_64EC
// Check the EC code bitmap incase we need to exit the JIT to call into native code.
ARMEmitter::SingleUseForwardLabel l_NotECCode;
ldr(TMP1, ARMEmitter::XReg::x18, TEB_PEB_OFFSET);
ldr(TMP1, TMP1, PEB_EC_CODE_BITMAP_OFFSET);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 15);
and_(ARMEmitter::Size::i64Bit, TMP2, TMP2, 0x1fffffffffff8);
ldr(TMP1, TMP1, TMP2, ARMEmitter::ExtendedType::LSL_64, 0);
lsr(ARMEmitter::Size::i64Bit, TMP2, RipReg, 12);
lsrv(ARMEmitter::Size::i64Bit, TMP1, TMP1, TMP2);
tbz(TMP1, 0, &l_NotECCode);
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
mov(EC_CALL_CHECKER_PC_REG, RipReg);
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
br(TMP2);
Bind(&l_NotECCode);
#endif
SpillStaticRegs(TMP1);
if (!TMP_ABIARGS) {
mov(ARMEmitter::XReg::x2, TMP3);
mov(ARMEmitter::XReg::x2, RipReg);
}
#ifndef _WIN32
ldr(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
#endif
#ifdef _M_ARM_64EC
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
ldr(ARMEmitter::XReg::x0, &l_CTX);
@@ -309,10 +312,11 @@ void Dispatcher::EmitDispatcher() {
FillStaticRegs();
#ifdef _M_ARM_64EC
ldr(TMP1, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
strb(ARMEmitter::WReg::zr, TMP1, CPUAreaInSyscallCallbackOffset);
ldr(TMP1, ARMEmitter::XReg::x18, TEB_CPU_AREA_OFFSET);
strb(ARMEmitter::WReg::zr, TMP1, CPU_AREA_IN_SYSCALL_CALLBACK_OFFSET);
#endif
#ifndef _WIN32
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
@@ -320,6 +324,7 @@ void Dispatcher::EmitDispatcher() {
// Trigger segfault if any deferred signals are pending
strb(ARMEmitter::XReg::zr, STATE,
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
#endif
b(&LoopTop);
}
@@ -632,7 +632,6 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
uint16_t X87Op = ((Op - 0xD8) << 8) | ModRMByte;
return NormalOp(&X87Ops[X87Op], X87Op);
} else if (Info->Type == FEXCore::X86Tables::TYPE_VEX_TABLE_PREFIX) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
uint16_t map_select = 1;
uint16_t pp = 0;
const uint8_t Byte1 = ReadByte();
@@ -678,7 +677,6 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
FEXCore::X86Tables::X86InstInfo* LocalInfo = &VEXTableOps[Op];
if (LocalInfo->Type >= FEXCore::X86Tables::TYPE_VEX_GROUP_12 && LocalInfo->Type <= FEXCore::X86Tables::TYPE_VEX_GROUP_17) {
FEXCORE_TELEMETRY_SET(VEXOpTelem, 1);
// We have ModRM
uint8_t ModRMByte = ReadByte();
DecodeInst->ModRM = ModRMByte;
-1
View File
@@ -120,7 +120,6 @@ private:
const uint8_t* AdjustAddrForSpecialRegion(const uint8_t* _InstStream, uint64_t EntryPoint, uint64_t RIP);
FEXCORE_TELEMETRY_INIT(VEXOpTelem, TYPE_USES_VEX_OPS);
FEXCORE_TELEMETRY_INIT(EVEXOpTelem, TYPE_USES_EVEX_OPS);
};
} // namespace FEXCore::Frontend
@@ -6,54 +6,59 @@
#include "Interface/IR/IR.h"
namespace FEXCore::CPU {
FEXCORE_PRESERVE_ALL_ATTR static void LoadDeferredFCW(uint16_t NewFCW) {
auto PC = (NewFCW >> 8) & 3;
FEXCORE_PRESERVE_ALL_ATTR static softfloat_state SoftFloatStateFromFCW(uint16_t FCW) {
softfloat_state State;
State.detectTininess = softfloat_tininess_afterRounding;
State.exceptionFlags = 0;
auto PC = (FCW >> 8) & 3;
switch (PC) {
case 0: extF80_roundingPrecision = 32; break;
case 2: extF80_roundingPrecision = 64; break;
case 3: extF80_roundingPrecision = 80; break;
case 0: State.roundingPrecision = 32; break;
case 2: State.roundingPrecision = 64; break;
case 3: State.roundingPrecision = 80; break;
case 1: LOGMAN_MSG_A_FMT("Invalid x87 precision mode, {}", PC);
}
auto RC = (NewFCW >> 10) & 3;
auto RC = (FCW >> 10) & 3;
switch (RC) {
case 0: softfloat_roundingMode = softfloat_round_near_even; break;
case 1: softfloat_roundingMode = softfloat_round_min; break;
case 2: softfloat_roundingMode = softfloat_round_max; break;
case 3: softfloat_roundingMode = softfloat_round_minMag; break;
case 0: State.roundingMode = softfloat_round_near_even; break;
case 1: State.roundingMode = softfloat_round_min; break;
case 2: State.roundingMode = softfloat_round_max; break;
case 3: State.roundingMode = softfloat_round_minMag; break;
}
return State;
}
template<>
struct OpHandlers<IR::OP_F80CVTTO> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, float src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, float src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle8(uint16_t FCW, double src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat(&State, src);
}
};
template<>
struct OpHandlers<IR::OP_F80CMP> {
template<uint32_t Flags>
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static uint64_t handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool eq, lt, nan;
uint64_t ResultFlags = 0;
X80SoftFloat::FCMP(Src1, Src2, &eq, &lt, &nan);
if (Flags & (1 << IR::FCMP_FLAG_LT) && lt) {
X80SoftFloat::FCMP(&State, Src1, Src2, &eq, &lt, &nan);
if (lt) {
ResultFlags |= (1 << IR::FCMP_FLAG_LT);
}
if (Flags & (1 << IR::FCMP_FLAG_UNORDERED) && nan) {
if (nan) {
ResultFlags |= (1 << IR::FCMP_FLAG_UNORDERED);
}
if (Flags & (1 << IR::FCMP_FLAG_EQ) && eq) {
if (eq) {
ResultFlags |= (1 << IR::FCMP_FLAG_EQ);
}
return ResultFlags;
@@ -62,275 +67,261 @@ struct OpHandlers<IR::OP_F80CMP> {
template<>
struct OpHandlers<IR::OP_F80CVT> {
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static float handle4(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static double handle8(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToF64(&State);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTINT> {
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI16(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI32(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return src;
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return src.ToI64(&State);
}
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
auto rv = extF80_to_i32(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int16_t handle2t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
auto rv = extF80_to_i32(&State, src, softfloat_round_minMag, false);
if (rv > INT16_MAX) {
return INT16_MAX;
} else if (rv < INT16_MIN) {
if (rv > INT16_MAX || rv < INT16_MIN) {
///< Indefinite value for 16-bit conversions.
return INT16_MIN;
} else {
return rv;
}
}
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i32(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int32_t handle4t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i32(&State, src, softfloat_round_minMag, false);
}
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t NewFCW, X80SoftFloat src) {
LoadDeferredFCW(NewFCW);
return extF80_to_i64(src, softfloat_round_minMag, false);
FEXCORE_PRESERVE_ALL_ATTR static int64_t handle8t(uint16_t FCW, X80SoftFloat src) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return extF80_to_i64(&State, src, softfloat_round_minMag, false);
}
};
template<>
struct OpHandlers<IR::OP_F80CVTTOINT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t NewFCW, int16_t src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle2(uint16_t FCW, int16_t src) {
return src;
}
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t NewFCW, int32_t src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle4(uint16_t FCW, int32_t src) {
return src;
}
};
template<>
struct OpHandlers<IR::OP_F80ROUND> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FRNDINT(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FRNDINT(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80F2XM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::F2XM1(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::F2XM1(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80TAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FTAN(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FTAN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SQRT> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSQRT(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSQRT(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80SIN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSIN(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSIN(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80COS> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FCOS(Src1);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FCOS(&State, Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_EXP> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_EXP(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80XTRACT_SIG> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
return X80SoftFloat::FXTRACT_SIG(Src1);
}
};
template<>
struct OpHandlers<IR::OP_F80ADD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FADD(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FADD(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SUB> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSUB(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSUB(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80MUL> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FMUL(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FMUL(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80DIV> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FDIV(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FDIV(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FYL2X> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FYL2X(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FYL2X(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80ATAN> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FATAN(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FATAN(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM1> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM1(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FREM1(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80FPREM> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FREM(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FREM(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F80SCALE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1, X80SoftFloat Src2) {
LoadDeferredFCW(NewFCW);
return X80SoftFloat::FSCALE(Src1, Src2);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1, X80SoftFloat Src2) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
return X80SoftFloat::FSCALE(&State, Src1, Src2);
}
};
template<>
struct OpHandlers<IR::OP_F64SIN> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return sin(src);
}
};
template<>
struct OpHandlers<IR::OP_F64COS> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return cos(src);
}
};
template<>
struct OpHandlers<IR::OP_F64TAN> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return tan(src);
}
};
template<>
struct OpHandlers<IR::OP_F64F2XM1> {
static double handle(uint16_t NewFCW, double src) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src) {
return exp2(src) - 1.0;
}
};
template<>
struct OpHandlers<IR::OP_F64ATAN> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return atan2(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return fmod(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FPREM1> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return remainder(src1, src2);
}
};
template<>
struct OpHandlers<IR::OP_F64FYL2X> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
return src2 * log2(src1);
}
};
template<>
struct OpHandlers<IR::OP_F64SCALE> {
static double handle(uint16_t NewFCW, double src1, double src2) {
LoadDeferredFCW(NewFCW);
static double handle(uint16_t FCW, double src1, double src2) {
double trunc = (double)(int64_t)(src2); // truncate
return src1 * exp2(trunc);
}
@@ -338,16 +329,16 @@ struct OpHandlers<IR::OP_F64SCALE> {
template<>
struct OpHandlers<IR::OP_F80BCDSTORE> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src1) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src1) {
softfloat_state State = SoftFloatStateFromFCW(FCW);
bool Negative = Src1.Sign;
Src1 = X80SoftFloat::FRNDINT(Src1);
Src1 = X80SoftFloat::FRNDINT(&State, Src1);
// Clear the Sign bit
Src1.Sign = 0;
uint64_t Tmp = Src1;
uint64_t Tmp = Src1.ToI64(&State);
X80SoftFloat Rv;
uint8_t* BCD = reinterpret_cast<uint8_t*>(&Rv);
memset(BCD, 0, 10);
@@ -379,8 +370,7 @@ struct OpHandlers<IR::OP_F80BCDSTORE> {
template<>
struct OpHandlers<IR::OP_F80BCDLOAD> {
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t NewFCW, X80SoftFloat Src) {
LoadDeferredFCW(NewFCW);
FEXCORE_PRESERVE_ALL_ATTR static X80SoftFloat handle(uint16_t FCW, X80SoftFloat Src) {
uint8_t* Src1 = reinterpret_cast<uint8_t*>(&Src);
uint64_t BCD {};
// We walk through each uint8_t and pull out the BCD encoding
@@ -35,14 +35,7 @@ void InterpreterOps::FillFallbackIndexPointers(uint64_t* Info) {
Info[Core::OPINDEX_F80CVTINT_TRUNC2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle2t);
Info[Core::OPINDEX_F80CVTINT_TRUNC4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle4t);
Info[Core::OPINDEX_F80CVTINT_TRUNC8] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTINT>::handle8t);
Info[Core::OPINDEX_F80CMP_0] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>);
Info[Core::OPINDEX_F80CMP_1] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>);
Info[Core::OPINDEX_F80CMP_2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>);
Info[Core::OPINDEX_F80CMP_3] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>);
Info[Core::OPINDEX_F80CMP_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>);
Info[Core::OPINDEX_F80CMP_5] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>);
Info[Core::OPINDEX_F80CMP_6] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>);
Info[Core::OPINDEX_F80CMP_7] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>);
Info[Core::OPINDEX_F80CMP] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle);
Info[Core::OPINDEX_F80CVTTOINT_2] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle2);
Info[Core::OPINDEX_F80CVTTOINT_4] = reinterpret_cast<uint64_t>(&FEXCore::CPU::OpHandlers<IR::OP_F80CVTTOINT>::handle4);
@@ -154,17 +147,8 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
break;
}
case IR::OP_F80CMP: {
auto Op = IROp->C<IR::IROp_F80Cmp>();
static constexpr std::array handlers {
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<0>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<1>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<2>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<3>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<4>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<5>,
&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<6>, &FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle<7>,
};
*Info = {FABI_I64_I16_F80_F80, (void*)handlers[Op->Flags], (Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP_0 + Op->Flags),
SupportsPreserveAllABI};
*Info = {FABI_I64_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80CMP>::handle,
(Core::FallbackHandlerIndex)(Core::OPINDEX_F80CMP), SupportsPreserveAllABI};
return true;
}
@@ -625,13 +625,17 @@ DEF_OP(ShiftFlags) {
// Set the output outside the branch to avoid needing an extra leg of the
// branch. We specifically do not hardcode the PF register anywhere (relying
// on a tied SRA register instead) to avoid fighting with RA/RCLSE.
// on a tied SRA register instead) to avoid fighting with RA.
if (PFTemp != PFInput) {
mov(ARMEmitter::Size::i64Bit, PFTemp, PFInput);
}
// We need to mask the source before comparing it. We don't just skip flag
// updates for Src2=0 but anything that masks to zero.
and_(ARMEmitter::Size::i32Bit, TMP1, Src2, OpSize == 8 ? 0x3f : 0x1f);
ARMEmitter::SingleUseForwardLabel Done;
cbz(EmitSize, Src2, &Done);
cbz(EmitSize, TMP1, &Done);
{
// PF/SF/ZF/OF
if (OpSize >= 4) {
@@ -642,20 +646,23 @@ DEF_OP(ShiftFlags) {
mov(ARMEmitter::Size::i64Bit, PFTemp, Dst);
}
auto CFWord = TMP1;
unsigned CFBit = 0;
// Extract the last bit shifted in to CF
if (Op->Shift == IR::ShiftType::LSL) {
if (OpSize >= 4) {
neg(EmitSize, TMP1, Src2);
neg(EmitSize, CFWord, Src2);
lsrv(EmitSize, CFWord, Src1, CFWord);
} else {
mov(EmitSize, TMP1, OpSize * 8);
sub(EmitSize, TMP1, TMP1, Src2);
CFWord = Dst.X();
CFBit = (OpSize * 8);
}
} else {
sub(ARMEmitter::Size::i64Bit, TMP1, Src2, 1);
sub(ARMEmitter::Size::i64Bit, CFWord, Src2, 1);
lsrv(EmitSize, CFWord, Src1, CFWord);
}
lsrv(EmitSize, TMP1, Src1, TMP1);
bool SetOF = Op->Shift != IR::ShiftType::ASR;
if (SetOF) {
// Only defined when Shift is 1 else undefined
@@ -664,14 +671,20 @@ DEF_OP(ShiftFlags) {
}
if (CTX->HostFeatures.SupportsFlagM) {
rmif(TMP1, 63, (1 << 1) /* C */);
rmif(CFWord, (CFBit - 1) % 64, (1 << 1) /* C */);
if (SetOF) {
rmif(TMP3, OpSize * 8 - 1, (1 << 0) /* V */);
}
} else {
mrs(TMP2, ARMEmitter::SystemRegister::NZCV);
bfi(ARMEmitter::Size::i32Bit, TMP2, TMP1, 29 /* C */, 1);
if (CFBit != 0) {
lsr(ARMEmitter::Size::i64Bit, TMP1, CFWord, CFBit);
CFWord = TMP1;
}
bfi(ARMEmitter::Size::i32Bit, TMP2, CFWord, 29 /* C */, 1);
if (SetOF) {
lsr(EmitSize, TMP3, TMP3, OpSize * 8 - 1);
@@ -704,59 +717,74 @@ DEF_OP(PDep) {
const auto Dest = GetReg(Node);
// PDep implementation follows the ideas from
// http://0x80.pl/articles/pdep-soft-emu.html ... Basically, iterate the *set*
// bits only, which will be faster than the naive implementation as long as
// there are enough holes in the mask.
//
// The specific arm64 assembly used is based on the sequence that clang
// generates for the C code, giving context to the scheduling yielding better
// ILP than I would do by hand. The registers are allocated by hand however,
// to fit within the tight constraints we have here withot spilling. Also, we
// use cbz/cbnz for conditional branching to avoid clobbering NZCV.
// We can't clobber these
const auto OrigInput = GetReg(Op->Input.ID());
const auto OrigMask = GetReg(Op->Mask.ID());
// So we have shadow as temporaries
const auto Input = TMP1.R();
const auto Mask = TMP2.R();
if (CTX->HostFeatures.SupportsSVEBitPerm) {
// SVE added support for PDEP but it needs to be done in a vector register.
if (EmitSize == ARMEmitter::Size::i32Bit) {
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), OrigInput.W());
fmov(ARMEmitter::Size::i32Bit, VTMP2.S(), OrigMask.W());
bdep(ARMEmitter::SubRegSize::i32Bit, VTMP1.Z(), VTMP1.Z(), VTMP2.Z());
umov<ARMEmitter::SubRegSize::i32Bit>(Dest, VTMP1, 0);
} else {
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), OrigInput.X());
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), OrigMask.X());
bdep(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), VTMP1.Z(), VTMP2.Z());
umov<ARMEmitter::SubRegSize::i64Bit>(Dest, VTMP1, 0);
}
} else {
// PDep implementation follows the ideas from
// http://0x80.pl/articles/pdep-soft-emu.html ... Basically, iterate the *set*
// bits only, which will be faster than the naive implementation as long as
// there are enough holes in the mask.
//
// The specific arm64 assembly used is based on the sequence that clang
// generates for the C code, giving context to the scheduling yielding better
// ILP than I would do by hand. The registers are allocated by hand however,
// to fit within the tight constraints we have here withot spilling. Also, we
// use cbz/cbnz for conditional branching to avoid clobbering NZCV.
// these get used variously as scratch
const auto T0 = TMP3.R();
const auto T1 = TMP4.R();
// So we have shadow as temporaries
const auto Input = TMP1.R();
const auto Mask = TMP2.R();
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
// these get used variously as scratch
const auto T0 = TMP3.R();
const auto T1 = TMP4.R();
// First, copy the input/mask, since we'll be clobbering. Copy as 64-bit to
// make this 0-uop on Firestorm.
mov(ARMEmitter::Size::i64Bit, Input, OrigInput);
mov(ARMEmitter::Size::i64Bit, Mask, OrigMask);
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
cbz(EmitSize, OrigMask, &Done);
// First, copy the input/mask, since we'll be clobbering. Copy as 64-bit to
// make this 0-uop on Firestorm.
mov(ARMEmitter::Size::i64Bit, Input, OrigInput);
mov(ARMEmitter::Size::i64Bit, Mask, OrigMask);
// Setup for first iteration
neg(EmitSize, T0, Mask);
and_(EmitSize, T0, T0, Mask);
// Now, they're copied, so we can start setting Dest (even if it overlaps with
// one of them). Handle early exit case
mov(EmitSize, Dest, 0);
cbz(EmitSize, OrigMask, &Done);
// Main loop
Bind(&NextBit);
sbfx(EmitSize, T1, Input, 0, 1);
eor(EmitSize, Mask, Mask, T0);
and_(EmitSize, T0, T1, T0);
neg(EmitSize, T1, Mask);
orr(EmitSize, Dest, Dest, T0);
lsr(EmitSize, Input, Input, 1);
and_(EmitSize, T0, Mask, T1);
cbnz(EmitSize, T0, &NextBit);
// Setup for first iteration
neg(EmitSize, T0, Mask);
and_(EmitSize, T0, T0, Mask);
// All done with nothing to do.
Bind(&Done);
// Main loop
Bind(&NextBit);
sbfx(EmitSize, T1, Input, 0, 1);
eor(EmitSize, Mask, Mask, T0);
and_(EmitSize, T0, T1, T0);
neg(EmitSize, T1, Mask);
orr(EmitSize, Dest, Dest, T0);
lsr(EmitSize, Input, Input, 1);
and_(EmitSize, T0, Mask, T1);
cbnz(EmitSize, T0, &NextBit);
// All done with nothing to do.
Bind(&Done);
}
}
DEF_OP(PExt) {
@@ -769,35 +797,50 @@ DEF_OP(PExt) {
const auto Mask = GetReg(Op->Mask.ID());
const auto Dest = GetReg(Node);
const auto MaskReg = TMP1;
const auto BitReg = TMP2;
const auto ValueReg = TMP3;
if (CTX->HostFeatures.SupportsSVEBitPerm) {
// SVE added support for PEXT but it needs to be done in a vector register.
if (EmitSize == ARMEmitter::Size::i32Bit) {
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Input.W());
fmov(ARMEmitter::Size::i32Bit, VTMP2.S(), Mask.W());
bext(ARMEmitter::SubRegSize::i32Bit, VTMP1.Z(), VTMP1.Z(), VTMP2.Z());
umov<ARMEmitter::SubRegSize::i32Bit>(Dest, VTMP1, 0);
} else {
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), Input.X());
fmov(ARMEmitter::Size::i64Bit, VTMP2.D(), Mask.X());
bext(ARMEmitter::SubRegSize::i64Bit, VTMP1.Z(), VTMP1.Z(), VTMP2.Z());
umov<ARMEmitter::SubRegSize::i64Bit>(Dest, VTMP1, 0);
}
} else {
const auto MaskReg = TMP1;
const auto BitReg = TMP2;
const auto ValueReg = TMP3;
ARMEmitter::SingleUseForwardLabel EarlyExit;
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
ARMEmitter::SingleUseForwardLabel EarlyExit;
ARMEmitter::BackwardLabel NextBit;
ARMEmitter::SingleUseForwardLabel Done;
cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, MaskReg, Mask);
mov(EmitSize, ValueReg, Input);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
cbz(EmitSize, Mask, &EarlyExit);
mov(EmitSize, MaskReg, Mask);
mov(EmitSize, ValueReg, Input);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
// Main loop
Bind(&NextBit);
cbz(EmitSize, MaskReg, &Done);
clz(EmitSize, BitReg, MaskReg);
lslv(EmitSize, ValueReg, ValueReg, BitReg);
lslv(EmitSize, MaskReg, MaskReg, BitReg);
extr(EmitSize, Dest, Dest, ValueReg, OpSizeBitsM1);
bfc(EmitSize, MaskReg, OpSizeBitsM1, 1);
b(&NextBit);
// Main loop
Bind(&NextBit);
cbz(EmitSize, MaskReg, &Done);
clz(EmitSize, BitReg, MaskReg);
lslv(EmitSize, ValueReg, ValueReg, BitReg);
lslv(EmitSize, MaskReg, MaskReg, BitReg);
extr(EmitSize, Dest, Dest, ValueReg, OpSizeBitsM1);
bfc(EmitSize, MaskReg, OpSizeBitsM1, 1);
b(&NextBit);
// Early exit
Bind(&EarlyExit);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
// Early exit
Bind(&EarlyExit);
mov(EmitSize, Dest, ARMEmitter::Reg::zr);
// All done with nothing to do.
Bind(&Done);
// All done with nothing to do.
Bind(&Done);
}
}
DEF_OP(LDiv) {
@@ -184,7 +184,7 @@ DEF_OP(Syscall) {
if ((Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Result is now in x0
// Fix the stack and any values that were stepped on
FillStaticRegs(true, GPRSpillMask, FPRSpillMask);
FillStaticRegs(true, GPRSpillMask, FPRSpillMask, ARMEmitter::Reg::r1, ARMEmitter::Reg::r2);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -285,7 +285,7 @@ DEF_OP(InlineSyscall) {
if ((Op->Flags & FEXCore::IR::SyscallFlags::NORETURN) != FEXCore::IR::SyscallFlags::NORETURN) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r1);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -5,6 +5,7 @@ tags: backend|arm64
$end_info$
*/
#include "Interface/Context/Context.h"
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
@@ -449,5 +450,63 @@ DEF_OP(Vector_FToI) {
}
}
DEF_OP(Vector_F64ToI32) {
const auto Op = IROp->C<IR::IROp_Vector_F64ToI32>();
const auto OpSize = IROp->Size;
const auto Round = Op->Round;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
if (HostSupportsSVE128 || HostSupportsSVE256) {
const auto Mask = Is256Bit ? PRED_TMP_32B.Merging() : PRED_TMP_16B.Merging();
// First step is to round the f64 values to integrals (frint*)
// Then convert to integers using fcvtzs.
auto CVTReg = Dst.Z();
switch (Round) {
case IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
case IR::Round_Towards_Zero.Val: CVTReg = Vector.Z(); break;
case IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Vector.Z()); break;
}
fcvtzs(Dst.Z(), ARMEmitter::SubRegSize::i32Bit, Mask, CVTReg, ARMEmitter::SubRegSize::i64Bit);
///< Fixup format of register that fcvtzs returns.
uzp1(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
if (Op->EnsureZeroUpperHalf) {
///< Match CVTPD2DQ/CVTTPD2DQ behaviour if necessary by zeroing the upper bits here.
if (Is256Bit) {
mov(Dst.Q(), Dst.Q());
} else {
mov(Dst.D(), Dst.D());
}
}
} else {
// This has a known precision issue that isn't easily resolvable without throwing away performance.
// Doing the conversion in multi-stage steps has an issue that you can lose precision in the f32->i32 step if your source was f64.
// To get around this with ASIMD FEX needs to use fcvtzs (Scalar, Integer, to GPR) for each F64 to be directly converted to i32.
// This is a very costly transform that the SVE path doesn't need to do since it supports f64->i32 directly.
// If this precision issue is necessary then we can add an option for it in the future.
///< Round float to integral depending on rounding mode.
switch (Round) {
case FEXCore::IR::Round_Nearest.Val: frintn(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Negative_Infinity.Val: frintm(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Positive_Infinity.Val: frintp(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Towards_Zero.Val: frintz(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
case FEXCore::IR::Round_Host.Val: frinti(ARMEmitter::SubRegSize::i64Bit, Dst.Q(), Vector.Q()); break;
}
// Now narrow from f64 to f32.
fcvtn(ARMEmitter::SubRegSize::i32Bit, Dst.Q(), Dst.Q());
///< Convert the two F32 integrals to real integers.
fcvtzs(ARMEmitter::SubRegSize::i32Bit, Dst.D(), Dst.D());
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -1,18 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: backend|arm64
$end_info$
*/
#include "Interface/Core/JIT/Arm64/JITClass.h"
namespace FEXCore::CPU {
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
DEF_OP(GetHostFlag) {
auto Op = IROp->C<IR::IROp_GetHostFlag>();
ubfx(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Value.ID()), Op->Flag, 1);
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -47,8 +47,8 @@ static uint64_t LUDIV(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
return Res;
}
static int64_t LDIV(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
static int64_t LDIV(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source / Divisor;
return Res;
}
@@ -59,8 +59,8 @@ static uint64_t LUREM(uint64_t SrcHigh, uint64_t SrcLow, uint64_t Divisor) {
return Res;
}
static int64_t LREM(int64_t SrcHigh, int64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__int128_t>(SrcHigh) << 64) | SrcLow;
static int64_t LREM(uint64_t SrcHigh, uint64_t SrcLow, int64_t Divisor) {
__int128_t Source = (static_cast<__uint128_t>(SrcHigh) << 64) | SrcLow;
__int128_t Res = Source % Divisor;
return Res;
}
@@ -496,7 +496,7 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Fram
uintptr_t branch = (uintptr_t)(Record)-8;
auto offset = HostCode / 4 - branch / 4;
if (vixl::IsInt26(offset)) {
if (ARMEmitter::Emitter::IsInt26(offset)) {
// optimal case - can branch directly
// patch the code
ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
@@ -729,7 +729,7 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
if (SpillSlots) {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (vixl::aarch64::Assembler::IsImmAddSub(TotalSpillSlotsSize)) {
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
sub(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
LoadConstant(ARMEmitter::Size::i64Bit, TMP1, TotalSpillSlotsSize);
@@ -872,7 +872,7 @@ void Arm64JITCore::ResetStack() {
const auto TotalSpillSlotsSize = SpillSlots * MaxSpillSlotSize;
if (vixl::aarch64::Assembler::IsImmAddSub(TotalSpillSlotsSize)) {
if (ARMEmitter::IsImmAddSub(TotalSpillSlotsSize)) {
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, ARMEmitter::Reg::rsp, TotalSpillSlotsSize);
} else {
// Too big to fit in a 12bit immediate
@@ -885,11 +885,4 @@ fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl*
return fextl::make_unique<Arm64JITCore>(ctx, Thread);
}
CPUBackendFeatures GetArm64JITBackendFeatures() {
return CPUBackendFeatures {
.SupportsFlags = true,
.SupportsVTBL2 = true,
};
}
} // namespace FEXCore::CPU
@@ -345,6 +345,10 @@ private:
uint32_t SpillSlots {};
using OpType = void (Arm64JITCore::*)(const IR::IROp_Header* IROp, IR::NodeID Node);
using ScalarFMAOpCaller =
std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3)>;
void VFScalarFMAOperation(uint8_t OpSize, uint8_t ElementSize, ScalarFMAOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2, ARMEmitter::VRegister Addend);
using ScalarBinaryOpCaller = std::function<void(ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2)>;
void VFScalarOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarBinaryOpCaller ScalarEmit,
ARMEmitter::VRegister Dst, ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2);
@@ -352,6 +356,10 @@ private:
void VFScalarUnaryOperation(uint8_t OpSize, uint8_t ElementSize, bool ZeroUpperBits, ScalarUnaryOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, std::variant<ARMEmitter::VRegister, ARMEmitter::Register> Vector2);
void Emulate128BitGather(size_t Size, size_t ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow,
std::optional<ARMEmitter::VRegister> VectorIndexHigh, ARMEmitter::VRegister MaskReg, size_t VectorIndexSize,
size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale);
// Runtime selection;
// Load and store TSO memory style
OpType RT_LoadMemTSO;
@@ -52,7 +52,7 @@ DEF_OP(StoreContext) {
const auto OpSize = IROp->Size;
if (Op->Class == FEXCore::IR::GPRClass) {
auto Src = GetReg(Op->Value.ID());
auto Src = GetZeroableReg(Op->Value);
switch (OpSize) {
case 1: strb(Src, STATE, Op->Offset); break;
@@ -99,7 +99,7 @@ DEF_OP(LoadRegister) {
}
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsAVX256 ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "out of range reg");
LOGMAN_THROW_A_FMT(OpSize == regSize, "expected sized");
@@ -120,8 +120,6 @@ DEF_OP(LoadRegister) {
DEF_OP(StoreRegister) {
const auto Op = IROp->C<IR::IROp_StoreRegister>();
const auto OpSize = IROp->Size;
if (Op->Class == IR::GPRClass) {
unsigned Reg = Op->Reg == Core::CPUState::PF_AS_GREG ? (StaticRegisters.size() - 2) :
@@ -137,9 +135,9 @@ DEF_OP(StoreRegister) {
mov(ARMEmitter::Size::i64Bit, reg, Src);
}
} else if (Op->Class == IR::FPRClass) {
const auto regSize = HostSupportsAVX256 ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto regSize = HostSupportsAVX256 ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
LOGMAN_THROW_A_FMT(Op->Reg < StaticFPRegisters.size(), "reg out of range");
LOGMAN_THROW_A_FMT(OpSize == regSize, "expected sized");
LOGMAN_THROW_A_FMT(IROp->Size == regSize, "expected sized");
const auto guest = StaticFPRegisters[Op->Reg];
const auto host = GetVReg(Op->Value.ID());
@@ -531,25 +529,6 @@ DEF_OP(LoadDF) {
ldrsb(Dst.X(), STATE, offsetof(FEXCore::Core::CPUState, flags[Flag]));
}
DEF_OP(LoadFlag) {
auto Op = IROp->C<IR::IROp_LoadFlag>();
auto Dst = GetReg(Node);
LOGMAN_THROW_A_FMT(Op->Flag != X86State::RFLAG_PF_RAW_LOC && Op->Flag != X86State::RFLAG_AF_RAW_LOC, "PF/AF must be accessed as "
"registers");
ldrb(Dst, STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag);
}
DEF_OP(StoreFlag) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
LOGMAN_THROW_A_FMT(Op->Flag != X86State::RFLAG_PF_RAW_LOC && Op->Flag != X86State::RFLAG_AF_RAW_LOC, "PF/AF must be accessed as "
"registers");
strb(GetReg(Op->Value.ID()), STATE, offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag);
}
ARMEmitter::ExtendedMemOperand Arm64JITCore::GenerateMemOperand(
uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale) {
if (Offset.IsInvalid()) {
@@ -940,6 +919,117 @@ DEF_OP(VStoreVectorMasked) {
}
}
void Arm64JITCore::Emulate128BitGather(
size_t Size, size_t ElementSize, ARMEmitter::VRegister Dst, ARMEmitter::VRegister IncomingDst,
std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndexLow, std::optional<ARMEmitter::VRegister> VectorIndexHigh,
ARMEmitter::VRegister MaskReg, size_t VectorIndexSize, size_t DataElementOffsetStart, size_t IndexElementOffsetStart, uint8_t OffsetScale) {
const auto PerformSMove = [this](size_t ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
switch (ElementSize) {
case 1: smov<ARMEmitter::SubRegSize::i8Bit>(Dst.X(), Vector, index); break;
case 2: smov<ARMEmitter::SubRegSize::i16Bit>(Dst.X(), Vector, index); break;
case 4: smov<ARMEmitter::SubRegSize::i32Bit>(Dst.X(), Vector, index); break;
case 8: umov<ARMEmitter::SubRegSize::i64Bit>(Dst.X(), Vector, index); break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", ElementSize); break;
}
};
const auto PerformMove = [this](size_t ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
switch (ElementSize) {
case 1: umov<ARMEmitter::SubRegSize::i8Bit>(Dst, Vector, index); break;
case 2: umov<ARMEmitter::SubRegSize::i16Bit>(Dst, Vector, index); break;
case 4: umov<ARMEmitter::SubRegSize::i32Bit>(Dst, Vector, index); break;
case 8: umov<ARMEmitter::SubRegSize::i64Bit>(Dst, Vector, index); break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", ElementSize); break;
}
};
// FEX needs to use a temporary destination vector register in a couple of instances.
// When Dst overlaps MaskReg, VectorIndexLow, or VectorIndexHigh
// Due to x86 gather instruction limitations, it is highly likely that a destination temporary isn't required.
const bool NeedsDestTmp = Dst == MaskReg || Dst == VectorIndexLow || (VectorIndexHigh.has_value() && Dst == *VectorIndexHigh);
// If the incoming destination isn't the destination then we need to move.
const bool NeedsIncomingDestMove = Dst != IncomingDst || NeedsDestTmp;
///< Adventurers beware, emulated ASIMD style gather masked load operation.
// Number of elements to load is calculated by the number of index elements available.
size_t NumAddrElements = (VectorIndexHigh.has_value() ? 32 : 16) / VectorIndexSize;
// The number of elements is clamped by the resulting register size.
size_t NumDataElements = std::min<size_t>(Size / ElementSize, NumAddrElements);
size_t IndexElementsSizeBytes = NumAddrElements * VectorIndexSize;
if (IndexElementsSizeBytes > 16) {
// We must have a high register in this case.
LOGMAN_THROW_A_FMT(VectorIndexHigh.has_value(), "Need High vector index register!");
}
auto ResultReg = Dst;
if (NeedsDestTmp) {
// Use VTMP1 as the temporary destination
ResultReg = VTMP1;
}
auto WorkingReg = TMP1;
auto TempMemReg = TMP2;
const uint64_t ElementSizeInBits = ElementSize * 8;
if (NeedsIncomingDestMove) {
mov(ResultReg.Q(), IncomingDst.Q());
}
for (size_t i = DataElementOffsetStart, IndexElement = IndexElementOffsetStart; i < NumDataElements; ++i, ++IndexElement) {
ARMEmitter::SingleUseForwardLabel Skip {};
// Extract mask element
PerformMove(ElementSize, WorkingReg, MaskReg, i);
// Skip if the mask's sign bit isn't set
tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
// Extract Index Element
if ((IndexElement * VectorIndexSize) >= 16) {
// Fetch from the high index register.
PerformSMove(VectorIndexSize, WorkingReg, *VectorIndexHigh, IndexElement - (16 / VectorIndexSize));
} else {
// Fetch from the low index register.
PerformSMove(VectorIndexSize, WorkingReg, VectorIndexLow, IndexElement);
}
// Calculate memory position for this gather load
if (BaseAddr.has_value()) {
if (VectorIndexSize == 4) {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
} else {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
}
} else {
///< In this case we have no base address, All addresses come from the vector register itself
if (VectorIndexSize == 4) {
// Sign extend and shift in to the 64-bit register
sbfiz(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
} else {
lsl(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
}
}
// Now that the address is calculated. Do the load.
switch (ElementSize) {
case 1: ld1<ARMEmitter::SubRegSize::i8Bit>(ResultReg.Q(), i, TempMemReg); break;
case 2: ld1<ARMEmitter::SubRegSize::i16Bit>(ResultReg.Q(), i, TempMemReg); break;
case 4: ld1<ARMEmitter::SubRegSize::i32Bit>(ResultReg.Q(), i, TempMemReg); break;
case 8: ld1<ARMEmitter::SubRegSize::i64Bit>(ResultReg.Q(), i, TempMemReg); break;
case 16: ldr(ResultReg.Q(), TempMemReg, 0); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, ElementSize); FEX_UNREACHABLE;
}
Bind(&Skip);
}
if (NeedsDestTmp) {
// Move result.
mov(Dst.Q(), ResultReg.Q());
}
}
DEF_OP(VLoadVectorGatherMasked) {
const auto Op = IROp->C<IR::IROp_VLoadVectorGatherMasked>();
const auto OpSize = IROp->Size;
@@ -977,31 +1067,11 @@ DEF_OP(VLoadVectorGatherMasked) {
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
const bool SupportsSVELoad = (HostSupportsSVE128 || HostSupportsSVE256) && (OffsetScale == 1 || OffsetScale == VectorIndexSize) &&
(VectorIndexSize == IROp->ElementSize);
const auto PerformSMove = [this](size_t ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
switch (ElementSize) {
case 1: smov<ARMEmitter::SubRegSize::i8Bit>(Dst.X(), Vector, index); break;
case 2: smov<ARMEmitter::SubRegSize::i16Bit>(Dst.X(), Vector, index); break;
case 4: smov<ARMEmitter::SubRegSize::i32Bit>(Dst.X(), Vector, index); break;
case 8: umov<ARMEmitter::SubRegSize::i64Bit>(Dst.X(), Vector, index); break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", ElementSize); break;
}
};
const auto PerformMove = [this](size_t ElementSize, const ARMEmitter::Register Dst, const ARMEmitter::VRegister Vector, int index) {
switch (ElementSize) {
case 1: umov<ARMEmitter::SubRegSize::i8Bit>(Dst, Vector, index); break;
case 2: umov<ARMEmitter::SubRegSize::i16Bit>(Dst, Vector, index); break;
case 4: umov<ARMEmitter::SubRegSize::i32Bit>(Dst, Vector, index); break;
case 8: umov<ARMEmitter::SubRegSize::i64Bit>(Dst, Vector, index); break;
default: LOGMAN_MSG_A_FMT("Unhandled ExtractElementSize: {}", ElementSize); break;
}
};
VectorIndexSize == IROp->ElementSize;
if (SupportsSVELoad) {
ARMEmitter::SVEModType ModType = ARMEmitter::SVEModType::MOD_NONE;
uint8_t SVEScale = FEXCore::ilog2(OffsetScale);
ARMEmitter::SVEModType ModType = ARMEmitter::SVEModType::MOD_NONE;
if (VectorIndexSize == 4) {
ModType = ARMEmitter::SVEModType::MOD_SXTW;
} else if (VectorIndexSize == 8 && OffsetScale != 1) {
@@ -1054,91 +1124,86 @@ DEF_OP(VLoadVectorGatherMasked) {
sel(SubRegSize, Dst.Z(), CMPPredicate, TempDst.Z(), IncomingDst.Z());
} else {
LOGMAN_THROW_A_FMT(!Is256Bit, "Can't emulate this gather load in the backend! Programming error!");
Emulate128BitGather(IROp->Size, IROp->ElementSize, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg,
VectorIndexSize, DataElementOffsetStart, IndexElementOffsetStart, OffsetScale);
}
}
// FEX needs to use a temporary destination vector register in a couple of instances.
// When Dst overlaps MaskReg, VectorIndexLow, or VectorIndexHigh
// Due to x86 gather instruction limitations, it is highly likely that a destination temporary isn't required.
const bool NeedsDestTmp = Dst == MaskReg || Dst == VectorIndexLow || (VectorIndexHigh.has_value() && Dst == *VectorIndexHigh);
DEF_OP(VLoadVectorGatherMaskedQPS) {
const auto Op = IROp->C<IR::IROp_VLoadVectorGatherMaskedQPS>();
// If the incoming destination isn't the destination then we need to move.
const bool NeedsIncomingDestMove = Dst != IncomingDst || NeedsDestTmp;
/// This instruction behaves similarly to the non-QPS version except for some STRICT limitations
/// - Only supports 32-bit element data size!
/// - Only supports 64-bit element address size!
/// - Only masks elements based on 32-bit element data size! (NOT ADDR SIZE!)
/// - Optimally uses SVE's `ld1w {zt.D}` variant instruction!
/// - Only outputs a single 128-bit result, while consuming 128-bit or 256-bit of address indexes!
/// - Matches VGATHERQPS/VPGATHERQD behaviour!
const auto OffsetScale = Op->OffsetScale;
const auto Dst = GetVReg(Node);
const auto IncomingDst = GetVReg(Op->Incoming.ID());
///< Adventurers beware, emulated ASIMD style gather masked load operation.
// Number of elements to load is calculated by the number of index elements available.
size_t NumAddrElements = (VectorIndexHigh.has_value() ? 32 : 16) / VectorIndexSize;
// The number of elements is clamped by the resulting register size.
size_t NumDataElements = std::min<size_t>(IROp->Size / IROp->ElementSize, NumAddrElements);
const auto MaskReg = GetVReg(Op->MaskReg.ID());
std::optional<ARMEmitter::Register> BaseAddr = !Op->AddrBase.IsInvalid() ? std::make_optional(GetReg(Op->AddrBase.ID())) : std::nullopt;
const auto VectorIndexLow = GetVReg(Op->VectorIndexLow.ID());
std::optional<ARMEmitter::VRegister> VectorIndexHigh =
!Op->VectorIndexHigh.IsInvalid() ? std::make_optional(GetVReg(Op->VectorIndexHigh.ID())) : std::nullopt;
size_t IndexElementsSizeBytes = NumAddrElements * VectorIndexSize;
if (IndexElementsSizeBytes > 16) {
// We must have a high register in this case.
LOGMAN_THROW_A_FMT(VectorIndexHigh.has_value(), "Need High vector index register!");
///< If the host supports SVE and the offset scale matches SVE limitations then it can do an SVE style load.
if (HostSupportsSVE128 && (OffsetScale == 1 || OffsetScale == 4)) {
ARMEmitter::SVEModType ModType = ARMEmitter::SVEModType::MOD_NONE;
if (OffsetScale != 1) {
ModType = ARMEmitter::SVEModType::MOD_LSL;
}
auto ResultReg = Dst;
if (NeedsDestTmp) {
// Use VTMP1 as the temporary destination
ResultReg = VTMP1;
}
auto WorkingReg = TMP1;
auto TempMemReg = TMP2;
const uint64_t ElementSizeInBits = IROp->ElementSize * 8;
const auto CMPPredicate = ARMEmitter::PReg::p0;
const auto CMPPredicate2 = ARMEmitter::PReg::p1;
if (NeedsIncomingDestMove) {
mov(ResultReg.Q(), IncomingDst.Q());
}
const auto GoverningPredicate = PRED_TMP_16B;
for (size_t i = DataElementOffsetStart, IndexElement = IndexElementOffsetStart; i < NumDataElements; ++i, ++IndexElement) {
ARMEmitter::SingleUseForwardLabel Skip {};
// Extract mask element
PerformMove(IROp->ElementSize, WorkingReg, MaskReg, i);
// Check if the sign bit is set for the given element size.
// This will set the predicate bits for elements [0, 1, 2, 3]
// We then use punpklo to extend the low results to be for 64-bit elements.
cmplt(ARMEmitter::SubRegSize::i32Bit, CMPPredicate, GoverningPredicate.Zeroing(), MaskReg.Z(), 0);
punpklo(CMPPredicate2, CMPPredicate);
auto TempDst = VTMP1;
// Skip if the mask's sign bit isn't set
tbz(WorkingReg, ElementSizeInBits - 1, &Skip);
// Extract Index Element
if ((IndexElement * VectorIndexSize) >= 16) {
// Fetch from the high index register.
PerformSMove(VectorIndexSize, WorkingReg, *VectorIndexHigh, IndexElement - (16 / VectorIndexSize));
} else {
// Fetch from the low index register.
PerformSMove(VectorIndexSize, WorkingReg, VectorIndexLow, IndexElement);
}
// Calculate memory position for this gather load
if (BaseAddr.has_value()) {
if (VectorIndexSize == 4) {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ExtendedType::SXTW, FEXCore::ilog2(OffsetScale));
auto GatherExtend = [this](ARMEmitter::VRegister Dst, std::optional<ARMEmitter::Register> BaseAddr, ARMEmitter::VRegister VectorIndex,
ARMEmitter::PRegister CMPPredicate, ARMEmitter::SVEModType ModType, uint8_t OffsetScale) {
// No need to load a temporary register in the case that we weren't provided a base address and there is no scaling.
uint8_t SVEScale = FEXCore::ilog2(OffsetScale);
ARMEmitter::SVEMemOperand MemDst {ARMEmitter::SVEMemOperand(VectorIndex.Z(), 0)};
if (BaseAddr.has_value() || OffsetScale != 1) {
ARMEmitter::Register AddrReg = TMP1;
if (BaseAddr.has_value()) {
AddrReg = *BaseAddr;
} else {
add(ARMEmitter::Size::i64Bit, TempMemReg, *BaseAddr, WorkingReg, ARMEmitter::ShiftType::LSL, FEXCore::ilog2(OffsetScale));
}
} else {
///< In this case we have no base address, All addresses come from the vector register itself
if (VectorIndexSize == 4) {
// Sign extend and shift in to the 64-bit register
sbfiz(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale), 32);
} else {
lsl(ARMEmitter::Size::i64Bit, TempMemReg, WorkingReg, FEXCore::ilog2(OffsetScale));
///< OpcodeDispatcher didn't provide a Base address while SVE requires one.
LoadConstant(ARMEmitter::Size::i64Bit, AddrReg, 0);
}
MemDst = ARMEmitter::SVEMemOperand(AddrReg.X(), VectorIndex.Z(), ModType, SVEScale);
}
// Now that the address is calculated. Do the load.
switch (IROp->ElementSize) {
case 1: ld1<ARMEmitter::SubRegSize::i8Bit>(ResultReg.Q(), i, TempMemReg); break;
case 2: ld1<ARMEmitter::SubRegSize::i16Bit>(ResultReg.Q(), i, TempMemReg); break;
case 4: ld1<ARMEmitter::SubRegSize::i32Bit>(ResultReg.Q(), i, TempMemReg); break;
case 8: ld1<ARMEmitter::SubRegSize::i64Bit>(ResultReg.Q(), i, TempMemReg); break;
case 16: ldr(ResultReg.Q(), TempMemReg, 0); break;
default: LOGMAN_MSG_A_FMT("Unhandled {} size: {}", __func__, IROp->ElementSize); FEX_UNREACHABLE;
}
ld1w<ARMEmitter::SubRegSize::i64Bit>(Dst.Z(), CMPPredicate.Zeroing(), MemDst);
};
Bind(&Skip);
}
GatherExtend(TempDst, BaseAddr, VectorIndexLow, CMPPredicate2, ModType, OffsetScale);
if (NeedsDestTmp) {
// Move result.
mov(Dst.Q(), ResultReg.Q());
if (VectorIndexHigh.has_value()) {
punpkhi(CMPPredicate2, CMPPredicate);
GatherExtend(VTMP2, BaseAddr, *VectorIndexHigh, CMPPredicate2, ModType, OffsetScale);
// Move elements to the lower half.
uzp1(ARMEmitter::SubRegSize::i32Bit, TempDst.Q(), TempDst.Q(), VTMP2.Q());
///< Merge elements based on predicate.
sel(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), CMPPredicate, TempDst.Z(), IncomingDst.Z());
} else {
// Move elements to the lower half.
xtn(ARMEmitter::SubRegSize::i32Bit, TempDst.Q(), TempDst.Q());
///< Merge elements based on predicate.
sel(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), CMPPredicate, TempDst.Z(), IncomingDst.Z());
}
} else {
Emulate128BitGather(16, 4, Dst, IncomingDst, BaseAddr, VectorIndexLow, VectorIndexHigh, MaskReg, 8, 0, 0, OffsetScale);
}
}
@@ -2222,7 +2287,7 @@ DEF_OP(VStoreNonTemporalPair) {
const auto Op = IROp->C<IR::IROp_VStoreNonTemporalPair>();
const auto OpSize = IROp->Size;
const auto Is128Bit = OpSize == Core::CPUState::XMM_SSE_REG_SIZE;
[[maybe_unused]] const auto Is128Bit = OpSize == Core::CPUState::XMM_SSE_REG_SIZE;
LOGMAN_THROW_A_FMT(Is128Bit, "This IR operation only operates at 128-bit wide");
const auto ValueLow = GetVReg(Op->ValueLow.ID());
@@ -2234,5 +2299,31 @@ DEF_OP(VStoreNonTemporalPair) {
stnp(ValueLow.Q(), ValueHigh.Q(), MemReg, Offset);
}
DEF_OP(VLoadNonTemporal) {
const auto Op = IROp->C<IR::IROp_VLoadNonTemporal>();
const auto OpSize = IROp->Size;
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Is128Bit = OpSize == Core::CPUState::XMM_SSE_REG_SIZE;
const auto Dst = GetVReg(Node);
const auto MemReg = GetReg(Op->Addr.ID());
const auto Offset = Op->Offset;
if (Is256Bit) {
LOGMAN_THROW_A_FMT(HostSupportsSVE256, "Need SVE256 support in order to use VStoreNonTemporal with 256-bit operation");
const auto GoverningPredicate = PRED_TMP_32B.Zeroing();
const auto OffsetScaled = Offset / 32;
ldnt1b(Dst.Z(), GoverningPredicate, MemReg, OffsetScaled);
} else if (Is128Bit && HostSupportsSVE128) {
const auto GoverningPredicate = PRED_TMP_16B.Zeroing();
const auto OffsetScaled = Offset / 16;
ldnt1b(Dst.Z(), GoverningPredicate, MemReg, OffsetScaled);
} else {
// Treat the non-temporal store as a regular vector store in this case for compatibility
ldr(Dst.Q(), MemReg, Offset);
}
}
#undef DEF_OP
} // namespace FEXCore::CPU
@@ -98,6 +98,7 @@ DEF_OP(GetRoundingMode) {
DEF_OP(SetRoundingMode) {
auto Op = IROp->C<IR::IROp_SetRoundingMode>();
auto Src = GetReg(Op->RoundMode.ID());
auto MXCSR = GetReg(Op->MXCSR.ID());
// As above, setup the rounding flags in [31:30]
rbit(ARMEmitter::Size::i32Bit, TMP2, Src);
@@ -116,6 +117,11 @@ DEF_OP(SetRoundingMode) {
lsr(ARMEmitter::Size::i64Bit, TMP2, Src, 2);
bfi(ARMEmitter::Size::i64Bit, TMP1, TMP2, 24, 1);
if (Op->SetDAZ && HostSupportsAFP) {
// Extract DAZ from MXCSR and insert to in FPCR.FIZ
bfxil(ARMEmitter::Size::i64Bit, TMP1, MXCSR, 6, 1);
}
// Now save the new FPCR
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
}
@@ -227,7 +233,7 @@ DEF_OP(ProcessorID) {
// Now that we are done in the syscall we need to carefully peel back the state
// First unspill the registers from before
FillStaticRegs(false, SpillMask);
FillStaticRegs(false, SpillMask, ~0U, ARMEmitter::Reg::r8, ARMEmitter::Reg::r2);
// Now the registers we've spilled are back in their original host registers
// We can safely claim we are no longer in a syscall
@@ -53,7 +53,7 @@ DEF_OP(Copy) {
DEF_OP(Swap1) {
auto Op = IROp->C<IR::IROp_Swap1>();
auto A = GetReg(Op->A.ID()), B = GetReg(Op->B.ID());
LOGMAN_THROW_AA_FMT(B == GetReg(Node), "Invariant");
LOGMAN_THROW_A_FMT(B == GetReg(Node), "Invariant");
mov(ARMEmitter::Size::i64Bit, TMP1, A);
mov(ARMEmitter::Size::i64Bit, A, B);
@@ -188,6 +188,30 @@ namespace FEXCore::CPU {
VFScalarOperation(IROp->Size, ElementSize, Op->ZeroUpperBits, ScalarEmit, Dst, Vector1, Vector2); \
}
#define DEF_FMAOP_SCALAR_INSERT(FEXOp, ARMOp) \
DEF_OP(FEXOp) { \
const auto Op = IROp->C<IR::IROp_##FEXOp>(); \
const auto ElementSize = Op->Header.ElementSize; \
\
auto ScalarEmit = \
[this, ElementSize](ARMEmitter::VRegister Dst, ARMEmitter::VRegister Src1, ARMEmitter::VRegister Src2, ARMEmitter::VRegister Src3) { \
if (ElementSize == 2) { \
ARMOp(Dst.H(), Src1.H(), Src2.H(), Src3.H()); \
} else if (ElementSize == 4) { \
ARMOp(Dst.S(), Src1.S(), Src2.S(), Src3.S()); \
} else if (ElementSize == 8) { \
ARMOp(Dst.D(), Src1.D(), Src2.D(), Src3.D()); \
} \
}; \
\
const auto Dst = GetVReg(Node); \
const auto Vector1 = GetVReg(Op->Vector1.ID()); \
const auto Vector2 = GetVReg(Op->Vector2.ID()); \
const auto Addend = GetVReg(Op->Addend.ID()); \
\
VFScalarFMAOperation(IROp->Size, ElementSize, ScalarEmit, Dst, Vector1, Vector2, Addend); \
}
DEF_UNOP(VAbs, abs, true)
DEF_UNOP(VPopcount, cnt, true)
DEF_UNOP(VNeg, neg, false)
@@ -224,6 +248,35 @@ DEF_FBINOP_SCALAR_INSERT(VFSubScalarInsert, fsub)
DEF_FBINOP_SCALAR_INSERT(VFMulScalarInsert, fmul)
DEF_FBINOP_SCALAR_INSERT(VFDivScalarInsert, fdiv)
DEF_FMAOP_SCALAR_INSERT(VFMLAScalarInsert, fmadd)
DEF_FMAOP_SCALAR_INSERT(VFMLSScalarInsert, fnmsub)
DEF_FMAOP_SCALAR_INSERT(VFNMLAScalarInsert, fmsub)
DEF_FMAOP_SCALAR_INSERT(VFNMLSScalarInsert, fnmadd)
void Arm64JITCore::VFScalarFMAOperation(uint8_t OpSize, uint8_t ElementSize, ScalarFMAOpCaller ScalarEmit, ARMEmitter::VRegister Dst,
ARMEmitter::VRegister Vector1, ARMEmitter::VRegister Vector2, ARMEmitter::VRegister Addend) {
LOGMAN_THROW_A_FMT(OpSize == Core::CPUState::XMM_SSE_REG_SIZE, "256-bit unsupported", __func__);
LOGMAN_THROW_AA_FMT(ElementSize == 2 || ElementSize == 4 || ElementSize == 8, "Invalid size");
const auto SubRegSize = ARMEmitter::ToVectorSizePair(ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
ARMEmitter::SubRegSize::i64Bit);
if (Dst != Vector1 && Dst != Vector2 && Dst != Addend && HostSupportsAFP) {
// If destination doesnt overlap any incoming register then move the adder to the destination first.
mov(Dst.Q(), Addend.Q());
Dst = Addend;
}
if (HostSupportsAFP && Dst == Addend) {
///< Exactly matches ARM scalar FMA semantics
// If the host CPU supports AFP then scalar does an insert without modifying upper bits.
ScalarEmit(Dst, Vector1, Vector2, Addend);
} else {
// No overlap between addr and destination or host doesn't support AFP, need to emit in to a temporary then insert.
ScalarEmit(VTMP1, Vector1, Vector2, Addend);
ins(SubRegSize.Vector, Dst.Q(), 0, VTMP1.Q(), 0);
}
}
// VFScalarOperation performs the operation described through ScalarEmit between Vector1 and Vector2,
// storing it into Dst. This is a scalar operation, so the only lowest element of each vector is used for the operation.
@@ -3165,6 +3218,48 @@ DEF_OP(VSXTL2) {
}
}
DEF_OP(VSSHLL) {
const auto Op = IROp->C<IR::IROp_VSSHLL>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto BitShift = Op->BitShift;
LOGMAN_THROW_A_FMT(BitShift < ((IROp->ElementSize >> 1) * 8), "Bitshift size too large for source element size: {} < {}", BitShift,
(IROp->ElementSize >> 1) * 8);
if (Is256Bit) {
sunpklo(SubRegSize, Dst.Z(), Vector.Z());
lsl(SubRegSize, Dst.Z(), Dst.Z(), BitShift);
} else {
sshll(SubRegSize, Dst.D(), Vector.D(), BitShift);
}
}
DEF_OP(VSSHLL2) {
const auto Op = IROp->C<IR::IROp_VSSHLL2>();
const auto OpSize = IROp->Size;
const auto SubRegSize = ConvertSubRegSize248(IROp);
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
const auto Dst = GetVReg(Node);
const auto Vector = GetVReg(Op->Vector.ID());
const auto BitShift = Op->BitShift;
LOGMAN_THROW_A_FMT(BitShift < ((IROp->ElementSize >> 1) * 8), "Bitshift size too large for source element size: {} < {}", BitShift,
(IROp->ElementSize >> 1) * 8);
if (Is256Bit) {
sunpkhi(SubRegSize, Dst.Z(), Vector.Z());
lsl(SubRegSize, Dst.Z(), Dst.Z(), BitShift);
} else {
sshll2(SubRegSize, Dst.Q(), Vector.Q(), BitShift);
}
}
DEF_OP(VUXTL) {
const auto Op = IROp->C<IR::IROp_VUXTL>();
const auto OpSize = IROp->Size;
@@ -4005,12 +4100,16 @@ DEF_OP(VFMLA) {
const auto Mask = PRED_TMP_32B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fmla(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else {
@@ -4026,7 +4125,11 @@ DEF_OP(VFMLA) {
}
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Q(), VectorAddend.Q());
}
if (OpSize == 16) {
@@ -4035,7 +4138,7 @@ DEF_OP(VFMLA) {
fmla(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Q(), DestTmp.Q());
}
}
@@ -4063,24 +4166,32 @@ DEF_OP(VFMLS) {
const auto Mask = PRED_TMP_32B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fnmls(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else if (HostSupportsSVE128 && Is128Bit) {
const auto Mask = PRED_TMP_16B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fnmls(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else {
@@ -4096,15 +4207,29 @@ DEF_OP(VFMLS) {
}
// Addend needs to get negated to match correct behaviour here.
ARMEmitter::VRegister DestTmp = VTMP1;
ARMEmitter::VRegister DestTmp = Dst;
if (Dst == Vector1 || Dst == Vector2) {
DestTmp = VTMP1;
}
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
fmla(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), DestTmp.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmla(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fmla(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
mov(Dst.D(), DestTmp.D());
}
if (DestTmp != Dst) {
if (Is128Bit) {
mov(Dst.Q(), DestTmp.Q());
} else {
mov(Dst.D(), DestTmp.D());
}
}
}
}
@@ -4130,12 +4255,16 @@ DEF_OP(VFNMLA) {
const auto Mask = PRED_TMP_32B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fmls(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else {
@@ -4152,7 +4281,11 @@ DEF_OP(VFNMLA) {
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Q(), VectorAddend.Q());
}
if (OpSize == 16) {
@@ -4161,7 +4294,7 @@ DEF_OP(VFNMLA) {
fmls(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
}
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Q(), DestTmp.Q());
}
}
@@ -4190,24 +4323,32 @@ DEF_OP(VFNMLS) {
const auto Mask = PRED_TMP_32B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fnmla(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else if (HostSupportsSVE128 && Is128Bit) {
const auto Mask = PRED_TMP_16B.Merging();
ARMEmitter::VRegister DestTmp = Dst;
if (Dst != VectorAddend) {
DestTmp = VTMP1;
if (Dst != Vector1 && Dst != Vector2) {
DestTmp = Dst;
} else {
DestTmp = VTMP1;
}
mov(DestTmp.Z(), VectorAddend.Z());
}
fnmla(SubRegSize, DestTmp.Z(), Mask, Vector1.Z(), Vector2.Z());
if (Dst != VectorAddend) {
if (Dst != DestTmp) {
mov(Dst.Z(), DestTmp.Z());
}
} else {
@@ -4223,15 +4364,29 @@ DEF_OP(VFNMLS) {
}
// Addend needs to get negated to match correct behaviour here.
ARMEmitter::VRegister DestTmp = VTMP1;
ARMEmitter::VRegister DestTmp = Dst;
if (Dst == Vector1 || Dst == Vector2) {
DestTmp = VTMP1;
}
if (Is128Bit) {
fneg(SubRegSize, DestTmp.Q(), VectorAddend.Q());
fmls(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
mov(Dst.Q(), DestTmp.Q());
} else {
fneg(SubRegSize, DestTmp.D(), VectorAddend.D());
}
if (Is128Bit) {
fmls(SubRegSize, DestTmp.Q(), Vector1.Q(), Vector2.Q());
} else {
fmls(SubRegSize, DestTmp.D(), Vector1.D(), Vector2.D());
mov(Dst.D(), DestTmp.D());
}
if (DestTmp != Dst) {
if (Is128Bit) {
mov(Dst.Q(), DestTmp.Q());
} else {
mov(Dst.D(), DestTmp.D());
}
}
}
}
@@ -17,6 +17,5 @@ class CPUBackend;
[[nodiscard]]
fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread);
CPUBackendFeatures GetArm64JITBackendFeatures();
} // namespace FEXCore::CPU
@@ -13,9 +13,6 @@
#include <stddef.h>
#include <utility>
#include <mutex>
#ifdef _M_ARM_64EC
#include <winnt.h>
#endif
namespace FEXCore {
@@ -70,24 +67,6 @@ public:
return 0;
}
#ifdef _M_ARM_64EC
bool CheckPageEC(uint64_t Address) {
if (!RtlIsEcCode(Address)) {
return false;
}
std::lock_guard<std::recursive_mutex> lk(WriteLock);
// Mark L2 entry for this page as EC by setting the LSB, this can then be
// checked by the dispatcher to see if it needs to perform a call/return to
// EC code.
const auto PageIndex = (Address & (VirtualMemSize - 1)) >> 12;
const auto Pointers = reinterpret_cast<uintptr_t*>(PagePointer);
Pointers[PageIndex] |= 1;
return true;
}
#endif
fextl::map<uint64_t, fextl::vector<uint64_t>> CodePages;
// Appends Block {Address} to CodePages [Start, Start + Length)
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -258,68 +258,7 @@ void OpDispatchBuilder::CalculateAF(Ref Src1, Ref Src2) {
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
}
void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
if (CurrentDeferredFlags.Type == FlagsGenerationType::TYPE_NONE) {
// Nothing to do
if (NZCVDirty && CachedNZCV) {
_StoreNZCV(CachedNZCV);
}
CachedNZCV = nullptr;
NZCVDirty = false;
return;
}
switch (CurrentDeferredFlags.Type) {
case FlagsGenerationType::TYPE_SUB:
CalculateFlags_SUB(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Sources.TwoSrcImmediate.Src1,
CurrentDeferredFlags.Sources.TwoSrcImmediate.Src2, CurrentDeferredFlags.Sources.TwoSrcImmediate.UpdateCF);
break;
case FlagsGenerationType::TYPE_MUL:
CalculateFlags_MUL(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res, CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_UMUL: CalculateFlags_UMUL(CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_LOGICAL:
CalculateFlags_Logical(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res, CurrentDeferredFlags.Sources.TwoSource.Src1,
CurrentDeferredFlags.Sources.TwoSource.Src2);
break;
case FlagsGenerationType::TYPE_LSHLI:
CalculateFlags_ShiftLeftImmediate(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1, CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_LSHRI:
CalculateFlags_ShiftRightImmediate(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1, CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_LSHRDI:
CalculateFlags_ShiftRightDoubleImmediate(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1, CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_ASHRI:
CalculateFlags_SignShiftRightImmediate(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res,
CurrentDeferredFlags.Sources.OneSrcImmediate.Src1, CurrentDeferredFlags.Sources.OneSrcImmediate.Imm);
break;
case FlagsGenerationType::TYPE_BEXTR: CalculateFlags_BEXTR(CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_BLSI: CalculateFlags_BLSI(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_BLSMSK:
CalculateFlags_BLSMSK(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res, CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_BLSR:
CalculateFlags_BLSR(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res, CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_POPCOUNT: CalculateFlags_POPCOUNT(CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_BZHI:
CalculateFlags_BZHI(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res, CurrentDeferredFlags.Sources.OneSource.Src1);
break;
case FlagsGenerationType::TYPE_ZCNT: CalculateFlags_ZCNT(CurrentDeferredFlags.SrcSize, CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_RDRAND: CalculateFlags_RDRAND(CurrentDeferredFlags.Res); break;
case FlagsGenerationType::TYPE_NONE:
default: ERROR_AND_DIE_FMT("Unhandled flags type {}", CurrentDeferredFlags.Type);
}
// Done calculating
CurrentDeferredFlags.Type = FlagsGenerationType::TYPE_NONE;
void OpDispatchBuilder::CalculateDeferredFlags() {
if (NZCVDirty && CachedNZCV) {
_StoreNZCV(CachedNZCV);
}
@@ -383,15 +322,14 @@ Ref OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, Ref Src1, Ref Src2) {
} else {
// Zero extend for correct comparison behaviour with Src1 = 0xffff.
Src1 = _Bfe(OpSize, SrcSize * 8, 0, Src1);
Src2 = _Bfe(OpSize, SrcSize * 8, 0, Src2);
auto CF = GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
auto Src1MinusCF = _Sub(OpSize, Src1, CF);
auto Src2PlusCF = _Adc(OpSize, _Constant(0), Src2);
Res = _Sub(OpSize, Src1MinusCF, Src2);
Res = _Sub(OpSize, Src1, Src2PlusCF);
Res = _Bfe(OpSize, SrcSize * 8, 0, Res);
// Need to zero-extend for correct comparisons below
auto SelectCF = _Select(FEXCore::IR::COND_ULT, Src1MinusCF, Res, One, Zero);
auto SelectCF = _Select(FEXCore::IR::COND_ULT, Src1, Src2PlusCF, One, Zero);
SetNZ_ZeroCV(SrcSize, Res);
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(SelectCF);
@@ -509,12 +447,13 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Ref U
// CF
{
// Extract the last bit shifted in to CF
// Extract the last bit shifted in to CF. Shift is already masked, but for
// 8/16-bit it might be >= SrcSizeBits, in which case CF is cleared. There's
// nothing to do in that case since we already cleared CF above.
auto SrcSizeBits = SrcSize * 8;
if (SrcSizeBits < Shift) {
Shift &= (SrcSizeBits - 1);
if (Shift < SrcSizeBits) {
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Src1, SrcSizeBits - Shift, true);
}
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Src1, SrcSizeBits - Shift, true);
}
CalculatePF(UnmaskedRes);
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -373,7 +373,7 @@ struct ThunkHandler_impl final : public ThunkHandler {
Thread = _Thread;
}
void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) override {
void AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) override {
for (auto& Definition : Definitions) {
Thunks.emplace(Definition.Sum, Definition.ThunkFunction);
}
+1 -1
View File
@@ -35,6 +35,6 @@ public:
static fextl::unique_ptr<ThunkHandler> Create();
virtual void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) = 0;
virtual void AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) = 0;
};
}; // namespace FEXCore
+1 -1
View File
@@ -83,7 +83,7 @@ struct AOTIRCacheEntry {
AOTIRInlineIndex* Array;
void* FilePtr;
size_t Size;
std::unique_ptr<FEXCore::HLE::SourcecodeMap> SourcecodeMap;
fextl::unique_ptr<FEXCore::HLE::SourcecodeMap> SourcecodeMap;
fextl::string FileId;
fextl::string Filename;
bool ContainsCode;
+418 -30
View File
@@ -168,7 +168,7 @@
"SwitchGen": false,
"JITDispatchOverride": "NoOp"
},
"IRHeader SSA:$Blocks, u64:$OriginalRIP, u32:$BlockCount, u32:$NumHostInstructions": {
"IRHeader SSA:$Blocks, u64:$OriginalRIP, u32:$BlockCount, u32:$NumHostInstructions, i1:$HasX87{false}": {
"SwitchGen": false,
"JITDispatchOverride": "NoOp"
},
@@ -226,7 +226,7 @@
"DestSize": "4"
},
"SetRoundingMode GPR:$RoundMode": {
"SetRoundingMode GPR:$RoundMode, i1:$SetDAZ, GPR:$MXCSR": {
"Desc": ["Sets the current rounding mode options for the thread"
],
"HasSideEffects": true
@@ -286,7 +286,7 @@
"Desc": ["Exits the current JIT function with a target RIP"
],
"HasSideEffects": true,
"DestSize": "GetOpSize(_NewRIP)"
"DestSize": "GetOpSize(NewRIP)"
},
"Break BreakDefinition:$Reason": {
"HasSideEffects": true
@@ -489,24 +489,6 @@
"DestSize": "8"
},
"GPR = LoadFlag u32:$Flag": {
"Desc": ["Loads an x86-64 flag from the context object",
"Specialized to allow flexible implementation of flag handling"
],
"DestSize": "1"
},
"StoreFlag GPR:$Value, u32:$Flag": {
"HasSideEffects": true,
"Desc": ["Stores 1-bit of the flag in to the specified x86-64 flag",
"Specialized to allow flexible implementation of flag handling"
],
"DestSize": "1"
},
"GPR = GetHostFlag GPR:$Value, u8:$Flag": {
},
"SSA = LoadMem RegisterClass:$Class, u8:#Size, GPR:$Addr, GPR:$Offset, u8:$Align, MemOffsetType:$OffsetType, u8:$OffsetScale": {
"DestSize": "Size"
},
@@ -570,6 +552,22 @@
"$VectorIndexElementSize == OpSize::i32Bit || $VectorIndexElementSize == OpSize::i64Bit"
]
},
"FPR = VLoadVectorGatherMaskedQPS u8:#RegisterSize, u8:#ElementSize, FPR:$Incoming, FPR:$MaskReg, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, u8:$OffsetScale": {
"Desc": [
"Does a masked load similar to VPGATHERQPS where the upper bit of each element",
"determines whether or not that element will be loaded from memory.",
"Most of VSIB encoding is passed directly through to the IR operation.",
"Only supports the case of 32-bit data element sizes from 64-bit addresses"
],
"TiedSource": 0,
"ImplicitFlagClobber": true,
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"EmitValidation": [
"ElementSize == OpSize::i32Bit",
"RegisterSize != FEXCore::IR::OpSize::i256Bit && \"What does 256-bit mean in this context?\""
]
},
"FPR = VLoadVectorElement u8:#RegisterSize, u8:#ElementSize, FPR:$DstSrc, u8:$Index, GPR:$Addr": {
"Desc": ["Does a memory load to a single element of a vector.",
"Leaves the rest of the vector's data intact.",
@@ -650,7 +648,7 @@
"Desc": ["Does a cacheline prefetch operation"
],
"EmitValidation": [
"_CacheLevel > 0 && _CacheLevel < 4"
"CacheLevel > 0 && CacheLevel < 4"
],
"HasSideEffects": true,
"DestSize": "8"
@@ -663,7 +661,7 @@
"HasSideEffects": true,
"DestSize": "RegisterSize",
"EmitValidation": [
"_Offset % RegisterSize == 0",
"Offset % RegisterSize == 0",
"RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i256Bit"
]
},
@@ -675,9 +673,21 @@
"HasSideEffects": true,
"DestSize": "RegisterSize",
"EmitValidation": [
"_Offset % RegisterSize == 0",
"Offset % RegisterSize == 0",
"RegisterSize == FEXCore::IR::OpSize::i128Bit"
]
},
"FPR = VLoadNonTemporal u8:#RegisterSize, GPR:$Addr, i8:$Offset": {
"Desc": ["Does a non-temporal memory load of a vector.",
"Matches arm64 SVE ldnt1b semantics.",
"Specifically weak-memory model ordered to match x86 non-temporal stores."
],
"HasSideEffects": true,
"DestSize": "RegisterSize",
"EmitValidation": [
"Offset % RegisterSize == 0",
"RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i256Bit"
]
}
},
"Atomic": {
@@ -1058,7 +1068,7 @@
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit",
"_Shift != ShiftType::ROR"
"Shift != ShiftType::ROR"
]
},
"GPR = AddWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
@@ -1158,7 +1168,7 @@
"DestSize": "Size",
"EmitValidation": [
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit",
"_Shift != ShiftType::ROR"
"Shift != ShiftType::ROR"
]
},
"GPR = SubWithFlags OpSize:#Size, GPR:$Src1, GPR:$Src2": {
@@ -1443,7 +1453,7 @@
"DestSize": "ResultSize",
"ImplicitFlagClobber": true,
"EmitValidation": [
"_CompareSize == FEXCore::IR::OpSize::i32Bit || _CompareSize == FEXCore::IR::OpSize::i64Bit || _CompareSize == FEXCore::IR::OpSize::i128Bit",
"CompareSize == FEXCore::IR::OpSize::i32Bit || CompareSize == FEXCore::IR::OpSize::i64Bit || CompareSize == FEXCore::IR::OpSize::i128Bit",
"ResultSize == FEXCore::IR::OpSize::i32Bit || ResultSize == FEXCore::IR::OpSize::i64Bit",
"WalkFindRegClass($Cmp1) == WalkFindRegClass($Cmp2)"
]
@@ -1706,6 +1716,42 @@
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = VFMLAScalarInsert u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
"Desc": [
"Dest = (Vector1 * Vector2) + Addend",
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"TiedSource": 2
},
"FPR = VFMLSScalarInsert u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
"Desc": [
"Dest = (Vector1 * Vector2) - Addend",
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"TiedSource": 2
},
"FPR = VFNMLAScalarInsert u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
"Desc": [
"Dest = (-Vector1 * Vector2) + Addend",
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"TiedSource": 2
},
"FPR = VFNMLSScalarInsert u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
"Desc": [
"Dest = (-Vector1 * Vector2) - Addend",
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize",
"TiedSource": 2
}
},
"Vector": {
@@ -1878,6 +1924,18 @@
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VSSHLL u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift{0}": {
"Desc": "Sign extends elements from the source element size to the next size up",
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VSSHLL2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift{0}": {
"Desc": ["Sign extends elements from the source element size to the next size up",
"Source elements come from the upper half of the register"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / (ElementSize << 1)"
},
"FPR = VUXTL u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
"Desc": "Zero extends elements from the source element size to the next size up",
"DestSize": "RegisterSize",
@@ -2431,6 +2489,13 @@
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / ElementSize"
},
"FPR = Vector_F64ToI32 u8:#RegisterSize, FPR:$Vector, RoundType:$Round, i1:$EnsureZeroUpperHalf": {
"Desc": ["Vector op: Rounds 64-bit float to 32-bit integral with round mode",
"Matches CVTPD2DQ/CVTTPD2DQ behaviour"
],
"DestSize": "RegisterSize",
"NumElements": "RegisterSize / FEXCore::IR::OpSize::i32Bit"
}
},
"Crypto": {
@@ -2521,34 +2586,283 @@
}
},
"F80": {
"GPR = SyncStackToSlow": {
"Desc": [
"Synchronizes the virtual stack environment to the physical registers.",
"Returns the current stack top."
],
"X87": true,
"HasSideEffects": true,
"DestSize": 8
},
"StackForceSlow": {
"Desc": [
"Forces the slow path."
],
"X87": true,
"HasSideEffects": true
},
"InitStack": {
"Desc": [
"Initializes the stack by marking all tags as invalid and setting top to zero."
],
"X87": true,
"HasSideEffects": true
},
"IncStackTop": {
"Desc": [
"Increase stack top-pointer."
],
"X87": true,
"HasSideEffects": true
},
"DecStackTop": {
"Desc": [
"Decrease stack top-pointer."
],
"X87": true,
"HasSideEffects": true
},
"InvalidateStack u8:$StackLocation": {
"Desc": [
"Marks the value in TOP+$StackLocation as empty / invalid 0b11.",
"If the StackLocation is 0xff, we invalidate all locations."
],
"X87": true,
"HasSideEffects": true
},
"PushStack FPR:$X80Src, SSA:$OriginalValue, u8:$LoadSize, i1:$Float": {
"Desc": [
"Pushes the provided X80Src source on to the x87 stack.",
"Tracks OriginalValue as the original value of X80Src.",
"Opsize is 128bit for F80 values, 64-bit for low precision.",
"LoadSize the original load size, i.e. of size of OriginalValue.",
"Float: 80-bit, 64-bit, 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"EmitValidation": [
"WalkFindRegClass($OriginalValue) == FPRClass || WalkFindRegClass($OriginalValue) == GPRClass"
],
"HasSideEffects": true,
"X87": true
},
"CopyPushStack u8:$StackLocation": {
"Desc": [
"Pushes an element already on the stack onto the top."
],
"HasSideEffects": true,
"X87": true
},
"StoreStackMemory GPR:$Addr, OpSize:$SourceSize, i1:$Float, u8:$StoreSize": {
"Desc": [
"Takes the top value off the x87 stack and stores it to memory.",
"SourceSize is 128bit for F80 values, 64-bit for low precision.",
"StoreSize is the store size for conversion:",
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"HasSideEffects": true,
"X87": true
},
"StoreStackToStack u8:$StackLocation": {
"Desc": [
"Takes the top value off the x87 stack and stores it to stack location TOP+StackLocation",
"Float: 80-bit, 64-bit, or 32-bit",
"Int: 64-bit, 32-bit, 16-bit"
],
"HasSideEffects": true,
"X87": true
},
"PopStackDestroy": {
"Desc": [
"Pops the top value off the stack but doesn't save it anywhere."
],
"HasSideEffects": true,
"X87": true
},
"FPR = ReadStackValue u8:$StackLocation": {
"Desc": [
"Reads a value off the stack at the offset"
],
"DestSize": "16",
"X87": true
},
"GPR = StackValidTag u8:$StackLocation": {
"Desc": [
"Returns 1 if the value in location TOP+$StackLocation is valid, 0 otherwise."
],
"DestSize": 4,
"X87": true
},
"F80AddStack u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Adds two stack locations together, storing the result in to the first stack location"
],
"HasSideEffects": true,
"X87": true
},
"F80AddValue u8:$SrcStack, FPR:$X80Src": {
"Desc": [
"Adds a operand value to a stack location. The result stored in to the stack location provided."
],
"HasSideEffects": true,
"X87": true
},
"FPR = F80Add FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80SubStack u8:$DstStack, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Subtracts the value in stack location TOP+$SrcStack2 from the value in stack location TOP+$SrcStack1.",
"The result is stored in stack location TOP+$DstStack."
],
"HasSideEffects": true,
"X87": true
},
"F80SubValue u8:$SrcStack, FPR:$X80Src": {
"Desc": [
"Subtracts the value $X80Src from the value in stack location TOP+$SrcStack.",
"The result is stored in stack location TOP."
],
"HasSideEffects": true,
"X87": true
},
"F80SubRValue FPR:$X80Src, u8:$SrcStack": {
"Desc": [
"Subtracts the value in stack location TOP+$SrcStack from the value $X80Src.",
"The result is stored in stack location TOP."
],
"HasSideEffects": true,
"X87": true
},
"FPR = F80Sub FPR:$X80Src1, FPR:$X80Src2": {
"Desc": [
"Subtracts the value in $X80Src1 from the value in $X80Src2.",
"The result is returned.",
"`FPR = X80Src2 - X80Src1`"
],
"DestSize": "16",
"JITDispatch": false
},
"F80MulStack u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Multiplies two stack locations together, storing the result in to the first stack location"
],
"HasSideEffects": true,
"X87": true
},
"F80MulValue u8:$SrcStack, FPR:$X80Src": {
"Desc": [
"Multiplies a operand value to a stack location. The result stored in to the stack location provided."
],
"HasSideEffects": true,
"X87": true
},
"FPR = F80Mul FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80DivStack u8:$DstStack, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Divides the value in stack location TOP+$SrcStack1 by the value in stack location TOP+$SrcStack2.",
"The result is stored in stack location TOP+$DstStack.",
"`FPR|Stack[TOP+DstStack] = Stack[TOP+SrcStack1] / Stack[TOP+SrcStack2]`"
],
"HasSideEffects": true,
"X87": true
},
"F80DivValue u8:$SrcStack, FPR:$X80Src": {
"Desc": [
"Divides the value in stack location TOP+$SrcStack by the value $X80Src.",
"The result is stored in stack location TOP and returned.",
"`FPR|Stack[TOP] = Stack[TOP+SrcStack] / X80Src`"
],
"HasSideEffects": true,
"X87": true
},
"F80DivRValue FPR:$X80Src, u8:$SrcStack": {
"Desc": [
"Divides the value X80Src by the value in stack location TOP+$SrcStack.",
"The result is stored in stack location TOP.",
"`FPR|Stack[TOP] = X80Src / Stack[TOP+SrcStack]`"
],
"HasSideEffects": true,
"X87": true
},
"FPR = F80Div FPR:$X80Src1, FPR:$X80Src2": {
"Desc": [
"Divides the value in $X80Src1 by the value in $X80Src2.",
"The result is returned.",
"`FPR = X80Src1 / X80Src2`"
],
"DestSize": "16",
"JITDispatch": false
},
"F80StackXchange u8:$SrcStack": {
"Desc": [
"Exchanges the value at the top of the stack with the value at TOP+$SrcStack."
],
"X87": true,
"HasSideEffects": true
},
"FPR = F80StackChangeSign": {
"Desc": [
"Complements the sign bit of the value at the top of the stack.",
"Returns the new value at the top of the stack."
],
"HasSideEffects": true,
"DestSize": "16",
"X87": true
},
"FPR = F80StackAbs": {
"Desc": [
"Clears the sign bit of the value at the top of the stack.",
"Returns the new value at the top of the stack."
],
"HasSideEffects": true,
"DestSize": "16",
"X87": true
},
"F80PTANStack": {
"Desc": [
"Computes the approximate tangent of the source operand in register ST(0), stores the result in ST(0), and pushes a 1.0 onto the FPU register stack."
],
"X87": true,
"HasSideEffects": true
},
"FPR = F80ATANStack": {
"Desc": [
"Computes arctan(st1/st0) and stores it in st0. Then pops the stack."
],
"DestSize": "16",
"X87": true,
"HasSideEffects": true
},
"FPR = F80ATAN FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80FPREMStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80FPREM FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80FPREM1Stack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80FPREM1 FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80SCALEStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80SCALE FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
@@ -2569,10 +2883,21 @@
"DestSize": "16",
"JITDispatch": false
},
"F80RoundStack": {
"Desc": [
"Replaces the value at the top of the stack with its nearest integral value."
],
"X87": true,
"HasSideEffects": true
},
"FPR = F80Round FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
},
"F80F2XM1Stack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80F2XM1 FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
@@ -2581,18 +2906,38 @@
"DestSize": "16",
"JITDispatch": false
},
"F80SINStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80SIN FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
},
"F80COSStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80COS FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
},
"F80SINCOSStack": {
"X87": true,
"HasSideEffects": true
},
"F80SQRTStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80SQRT FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
},
"F80XTRACTStack": {
"X87": true,
"HasSideEffects": true
},
"FPR = F80XTRACT_EXP FPR:$X80Src": {
"DestSize": "16",
"JITDispatch": false
@@ -2601,8 +2946,32 @@
"DestSize": "16",
"JITDispatch": false
},
"GPR = F80Cmp FPR:$X80Src1, FPR:$X80Src2, u32:$Flags": {
"Desc": ["Does a scalar unordered compare and stores the asked for flags in to a GPR",
"GPR = F80StackTest u8:$SrcStack": {
"Desc": [
"Does comparison between value in stack at TOP + SrcStack"
],
"DestSize": "4",
"X87": true
},
"GPR = F80CmpStack u8:$SrcStack": {
"Desc": [
"Does a scalar unordered compare between the value at the top of the stack and the value in stack position TOP+$SrcStack and stores the flags in to a GPR",
"Ordering flag result is true if either float input is NaN"
],
"DestSize": "4",
"X87": true
},
"GPR = F80CmpValue FPR:$X80Src": {
"Desc": [
"Does a scalar unordered compare between the value at the top of the stack and $X80Src and stores the asked for flags in to a GPR",
"Ordering flag result is true if either float input is NaN"
],
"DestSize": "4",
"HasSideEffects": true,
"X87": true
},
"GPR = F80Cmp FPR:$X80Src1, FPR:$X80Src2": {
"Desc": ["Does a scalar unordered compare and stores the flags in to a GPR",
"Ordering flag result is true if either float input is NaN"
],
"DestSize": "4",
@@ -2616,10 +2985,29 @@
"DestSize": "16",
"JITDispatch": false
},
"FPR = F80FYL2XStack": {
"Desc": [
"Computes ST1 * log2(ST0)",
"Stores the result in ST1, and pops the top of the stack.",
"Returns the new value at the top of the stack, i.e. the result of the operation."
],
"HasSideEffects": true,
"DestSize": "16",
"X87": true
},
"FPR = F80FYL2X FPR:$X80Src1, FPR:$X80Src2": {
"DestSize": "16",
"JITDispatch": false
},
"F80VBSLStack u8:#RegisterSize, FPR:$VectorMask, u8:$SrcStack1, u8:$SrcStack2": {
"Desc": [
"Does a vector bitwise select.",
"If the bit in the field is 1 then the corresponding bit is pulled from VectorTrue",
"If the bit in the field is 0 then the corresponding bit is pulled from VectorFalse",
"Writes the result to the top of the stack."
],
"X87": true,
"HasSideEffects": true
}
},
"Backend": {
+3 -3
View File
@@ -343,9 +343,9 @@ protected:
return Ptr;
}
virtual void SaveNZCV(IROps Op) {
// Overriden by dispatcher, stubbed for IR tests
}
// Overriden by dispatcher, stubbed for IR tests
virtual void RecordX87Use() {}
virtual void SaveNZCV(IROps Op) {}
Ref CurrentWriteCursor = nullptr;
+1 -1
View File
@@ -70,7 +70,7 @@ void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx) {
FEX_CONFIG_OPT(DisablePasses, O0);
if (!DisablePasses()) {
InsertPass(CreateContextLoadStoreElimination(ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256));
InsertPass(CreateX87StackOptimizationPass());
InsertPass(CreateDeadStoreElimination());
InsertPass(CreateConstProp(ctx->HostFeatures.SupportsTSOImm9, &ctx->CPUID));
InsertPass(CreateDeadFlagCalculationEliminination());
+1 -1
View File
@@ -17,10 +17,10 @@ class RegisterAllocationPass;
class RegisterAllocationData;
fextl::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool SupportsTSOImm9, const FEXCore::CPUIDEmu* CPUID);
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsSVE256);
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination();
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass();
fextl::unique_ptr<FEXCore::IR::Pass> CreateX87StackOptimizationPass();
namespace Validation {
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
@@ -5,13 +5,7 @@ tags: ir|opts
desc: ConstProp, ZExt elim, const pooling, fcmp reduction, const inlining
$end_info$
*/
// aarch64 heuristics
#include "aarch64/assembler-aarch64.h"
#include "aarch64/cpu-aarch64.h"
#include "aarch64/disasm-aarch64.h"
#include "aarch64/assembler-aarch64.h"
#include <CodeEmitter/Emitter.h>
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/PassManager.h"
@@ -56,10 +50,7 @@ static bool IsImmLogical(uint64_t imm, unsigned width) {
if (width < 32) {
width = 32;
}
return vixl::aarch64::Assembler::IsImmLogical(imm, width);
}
static bool IsImmAddSub(uint64_t imm) {
return vixl::aarch64::Assembler::IsImmAddSub(imm);
return ARMEmitter::Emitter::IsImmLogical(imm, width);
}
static bool IsBfeAlreadyDone(IREmitter* IREmit, OrderedNodeWrapper src, uint64_t Width) {
@@ -166,7 +157,7 @@ void ConstProp::ConstantPropagation(IREmitter* IREmit, const IRListView& Current
} else if (IsConstant1 && IsConstant2 && IROp->Op == OP_SUB) {
uint64_t NewConstant = (Constant1 - Constant2) & getMask(IROp);
IREmit->ReplaceWithConstant(CodeNode, NewConstant);
} else if (IsConstant2 && !IsImmAddSub(Constant2) && IsImmAddSub(-Constant2)) {
} else if (IsConstant2 && !ARMEmitter::IsImmAddSub(Constant2) && ARMEmitter::IsImmAddSub(-Constant2)) {
// If the second argument is constant, the immediate is not ImmAddSub, but when negated is.
// So, negate the operation to negate (and inline) the constant.
if (IROp->Op == OP_ADD) {
@@ -611,7 +602,7 @@ void ConstProp::ConstantInlining(IREmitter* IREmit, const IRListView& CurrentIR)
if (IREmit->IsValueConstant(IROp->Args[1], &Constant2)) {
// We don't allow 8/16-bit operations to have constants, since no
// constant would be in bounds after the JIT's 24/16 shift.
if (IsImmAddSub(Constant2) && IROp->Size >= 4) {
if (ARMEmitter::IsImmAddSub(Constant2) && IROp->Size >= 4) {
IREmit->SetWriteCursor(CurrentIR.GetNode(IROp->Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, CreateInlineConstant(IREmit, Constant2));
}
@@ -629,7 +620,8 @@ void ConstProp::ConstantInlining(IREmitter* IREmit, const IRListView& CurrentIR)
break;
}
case OP_ADC:
case OP_ADCWITHFLAGS: {
case OP_ADCWITHFLAGS:
case OP_STORECONTEXT: {
uint64_t Constant1 {};
if (IREmit->IsValueConstant(IROp->Args[0], &Constant1)) {
if (Constant1 == 0) {
@@ -655,7 +647,7 @@ void ConstProp::ConstantInlining(IREmitter* IREmit, const IRListView& CurrentIR)
case OP_CONDSUBNZCV: {
uint64_t Constant2 {};
if (IREmit->IsValueConstant(IROp->Args[1], &Constant2)) {
if (IsImmAddSub(Constant2)) {
if (ARMEmitter::IsImmAddSub(Constant2)) {
IREmit->SetWriteCursor(CurrentIR.GetNode(IROp->Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, CreateInlineConstant(IREmit, Constant2));
}
@@ -683,7 +675,7 @@ void ConstProp::ConstantInlining(IREmitter* IREmit, const IRListView& CurrentIR)
case OP_SELECT: {
uint64_t Constant1 {};
if (IREmit->IsValueConstant(IROp->Args[1], &Constant1)) {
if (IsImmAddSub(Constant1)) {
if (ARMEmitter::IsImmAddSub(Constant1)) {
IREmit->SetWriteCursor(CurrentIR.GetNode(IROp->Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, CreateInlineConstant(IREmit, Constant1));
}
@@ -725,7 +717,7 @@ void ConstProp::ConstantInlining(IREmitter* IREmit, const IRListView& CurrentIR)
case OP_CONDJUMP: {
uint64_t Constant2 {};
if (IREmit->IsValueConstant(IROp->Args[1], &Constant2)) {
if (IsImmAddSub(Constant2)) {
if (ARMEmitter::IsImmAddSub(Constant2)) {
IREmit->SetWriteCursor(CurrentIR.GetNode(IROp->Args[1]));
IREmit->ReplaceNodeArgument(CodeNode, 1, CreateInlineConstant(IREmit, Constant2));
}
@@ -1,772 +0,0 @@
// SPDX-License-Identifier: MIT
/*
$info$
tags: ir|opts
desc: Transforms ContextLoad/Store to temporaries, similar to mem2reg
$end_info$
*/
#include "Interface/IR/IR.h"
#include "Interface/IR/IREmitter.h"
#include "Interface/IR/Passes.h"
#include "Interface/IR/PassManager.h"
#include <FEXCore/Core/CoreState.h>
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/Utils/EnumOperators.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/Utils/Profiler.h>
#include <FEXCore/fextl/unordered_map.h>
#include <FEXCore/fextl/vector.h>
#include <array>
#include <memory>
#include <stddef.h>
#include <stdint.h>
#include <unordered_map>
#include <utility>
namespace {
struct ContextMemberClassification {
size_t Offset;
uint16_t Size;
};
enum class LastAccessType {
NONE = (0b000 << 0), ///< Was never previously accessed
WRITE = (0b001 << 0), ///< Was fully overwritten
READ = (0b010 << 0), ///< Was fully read
INVALID = (0b011 << 0), ///< Accessing this is invalid
MASK = (0b011 << 0),
PARTIAL = (0b100 << 0),
PARTIAL_WRITE = (PARTIAL | WRITE), ///< Was partially written
PARTIAL_READ = (PARTIAL | READ), ///< Was partially read
};
FEX_DEF_NUM_OPS(LastAccessType);
static bool IsWriteAccess(LastAccessType Type) {
return (Type & LastAccessType::MASK) == LastAccessType::WRITE;
}
static bool IsReadAccess(LastAccessType Type) {
return (Type & LastAccessType::MASK) == LastAccessType::READ;
}
[[maybe_unused]]
static bool IsInvalidAccess(LastAccessType Type) {
return (Type & LastAccessType::MASK) == LastAccessType::INVALID;
}
[[maybe_unused]]
static bool IsPartialAccess(LastAccessType Type) {
return (Type & LastAccessType::PARTIAL) == LastAccessType::PARTIAL;
}
[[maybe_unused]]
static bool IsFullAccess(LastAccessType Type) {
return (Type & LastAccessType::PARTIAL) == LastAccessType::NONE;
}
struct ContextMemberInfo {
ContextMemberClassification Class;
LastAccessType Accessed;
FEXCore::IR::RegisterClassType AccessRegClass;
uint32_t AccessOffset;
uint8_t AccessSize;
///< The last value that was loaded or stored.
FEXCore::IR::Ref ValueNode;
///< With a store access, the store node that is doing the operation.
FEXCore::IR::Ref StoreNode;
};
struct ContextInfo {
fextl::vector<ContextMemberInfo*> Lookup;
fextl::vector<ContextMemberInfo> ClassificationInfo;
};
static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool SupportsAVX256) {
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader),
sizeof(FEXCore::Core::CPUState::InlineJITBlockHeader),
},
LastAccessType::INVALID,
FEXCore::IR::InvalidClass,
});
// DeferredSignalRefCount
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount),
sizeof(FEXCore::Core::CPUState::DeferredSignalRefCount),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, avx_high[0][0]) + FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * i,
FEXCore::Core::CPUState::XMM_SSE_REG_SIZE,
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, rip),
sizeof(FEXCore::Core::CPUState::rip),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GPRS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gregs[0]) + sizeof(FEXCore::Core::CPUState::gregs[0]) * i,
FEXCore::Core::CPUState::GPR_REG_SIZE,
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, _pad),
sizeof(FEXCore::Core::CPUState::_pad),
},
LastAccessType::INVALID,
FEXCore::IR::InvalidClass,
});
static_assert(offsetof(FEXCore::Core::CPUState, xmm.avx.data[0][0]) == 416, "What");
static_assert(FEXCore::Core::CPUState::XMM_AVX_REG_SIZE == 32, "What");
if (SupportsAVX256) {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, xmm.avx.data[0][0]) + FEXCore::Core::CPUState::XMM_AVX_REG_SIZE * i,
FEXCore::Core::CPUState::XMM_AVX_REG_SIZE,
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
} else {
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, xmm.sse.data[0][0]) + FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * i,
FEXCore::Core::CPUState::XMM_SSE_REG_SIZE,
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, xmm.sse.pad[0][0]),
static_cast<uint16_t>(FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * FEXCore::Core::CPUState::NUM_XMMS),
},
LastAccessType::INVALID,
FEXCore::IR::InvalidClass,
});
}
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, es_idx),
sizeof(FEXCore::Core::CPUState::es_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, cs_idx),
sizeof(FEXCore::Core::CPUState::cs_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, ss_idx),
sizeof(FEXCore::Core::CPUState::ss_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, ds_idx),
sizeof(FEXCore::Core::CPUState::ds_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gs_idx),
sizeof(FEXCore::Core::CPUState::gs_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, fs_idx),
sizeof(FEXCore::Core::CPUState::fs_idx),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, _pad2),
sizeof(FEXCore::Core::CPUState::_pad2),
},
LastAccessType::INVALID,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, es_cached),
sizeof(FEXCore::Core::CPUState::es_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, cs_cached),
sizeof(FEXCore::Core::CPUState::cs_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, ss_cached),
sizeof(FEXCore::Core::CPUState::ss_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, ds_cached),
sizeof(FEXCore::Core::CPUState::ds_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gs_cached),
sizeof(FEXCore::Core::CPUState::gs_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, fs_cached),
sizeof(FEXCore::Core::CPUState::fs_cached),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, flags[0]) + sizeof(FEXCore::Core::CPUState::flags[0]) * i,
FEXCore::Core::CPUState::FLAG_SIZE,
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, pf_raw),
sizeof(FEXCore::Core::CPUState::pf_raw),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, af_raw),
sizeof(FEXCore::Core::CPUState::af_raw),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_MMS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {offsetof(FEXCore::Core::CPUState, mm[0][0]) + sizeof(FEXCore::Core::CPUState::mm[0]) * i,
FEXCore::Core::CPUState::MM_REG_SIZE},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
// GDTs
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GDTS; ++i) {
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, gdt[0]) + sizeof(FEXCore::Core::CPUState::gdt[0]) * i,
sizeof(FEXCore::Core::CPUState::gdt[0]),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
}
// FCW
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, FCW),
sizeof(FEXCore::Core::CPUState::FCW),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
// AbridgedFTW
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, AbridgedFTW),
sizeof(FEXCore::Core::CPUState::AbridgedFTW),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
// _pad3
ContextClassification->emplace_back(ContextMemberInfo {
ContextMemberClassification {
offsetof(FEXCore::Core::CPUState, _pad3),
sizeof(FEXCore::Core::CPUState::_pad3),
},
LastAccessType::NONE,
FEXCore::IR::InvalidClass,
});
[[maybe_unused]] size_t ClassifiedStructSize {};
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
for (auto& it : *ContextClassification) {
LOGMAN_THROW_A_FMT(it.Class.Offset == ContextClassificationInfo->Lookup.size(), "Offset mismatch (offset={})", it.Class.Offset);
for (int i = 0; i < it.Class.Size; i++) {
ContextClassificationInfo->Lookup.push_back(&it);
}
ClassifiedStructSize += it.Class.Size;
}
LOGMAN_THROW_AA_FMT(ClassifiedStructSize == sizeof(FEXCore::Core::CPUState),
"Classified CPUStruct size doesn't match real CPUState struct size! {} (classified) != {} (real)",
ClassifiedStructSize, sizeof(FEXCore::Core::CPUState));
LOGMAN_THROW_A_FMT(ContextClassificationInfo->Lookup.size() == sizeof(FEXCore::Core::CPUState),
"Classified lookup size doesn't match real CPUState struct size! {} (classified) != {} (real)",
ContextClassificationInfo->Lookup.size(), sizeof(FEXCore::Core::CPUState));
}
static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo, bool SupportsAVX256) {
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
auto SetAccess = [&](size_t Offset, LastAccessType Access) {
ContextClassification->at(Offset).Accessed = Access;
ContextClassification->at(Offset).AccessRegClass = FEXCore::IR::InvalidClass;
ContextClassification->at(Offset).AccessOffset = 0;
ContextClassification->at(Offset).StoreNode = nullptr;
};
size_t Offset = 0;
///< InlineJITBlockHeader
SetAccess(Offset++, LastAccessType::INVALID);
// DeferredSignalRefCount
SetAccess(Offset++, LastAccessType::INVALID);
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
///< avx_high
SetAccess(Offset++, LastAccessType::NONE);
}
// rip
SetAccess(Offset++, LastAccessType::NONE);
///< gregs
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GPRS; ++i) {
SetAccess(Offset++, LastAccessType::NONE);
}
// pad
SetAccess(Offset++, LastAccessType::NONE);
// xmm
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
SetAccess(Offset++, LastAccessType::NONE);
}
// xmm_pad
if (!SupportsAVX256) {
SetAccess(Offset++, LastAccessType::NONE);
}
// Segment indexes
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
// Pad2
SetAccess(Offset++, LastAccessType::INVALID);
// Segments
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
SetAccess(Offset++, LastAccessType::NONE);
///< flags
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
SetAccess(Offset++, LastAccessType::NONE);
}
///< pf_raw
SetAccess(Offset++, LastAccessType::NONE);
///< af_raw
SetAccess(Offset++, LastAccessType::NONE);
///< mm
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_MMS; ++i) {
SetAccess(Offset++, LastAccessType::NONE);
}
///< gdt
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GDTS; ++i) {
SetAccess(Offset++, LastAccessType::NONE);
}
///< FCW
SetAccess(Offset++, LastAccessType::NONE);
///< AbridgedFTW
SetAccess(Offset++, LastAccessType::NONE);
// pad3
SetAccess(Offset++, LastAccessType::INVALID);
}
struct BlockInfo {
fextl::vector<FEXCore::IR::Ref> Predecessors;
fextl::vector<FEXCore::IR::Ref> Successors;
ContextInfo IncomingClassifiedStruct;
ContextInfo OutgoingClassifiedStruct;
};
class RCLSE final : public FEXCore::IR::Pass {
public:
explicit RCLSE(bool SupportsAVX256)
: SupportsAVX256 {SupportsAVX256} {
ClassifyContextStruct(&ClassifiedStruct, SupportsAVX256);
}
void Run(FEXCore::IR::IREmitter* IREmit) override;
private:
ContextInfo ClassifiedStruct;
fextl::unordered_map<FEXCore::IR::NodeID, BlockInfo> OffsetToBlockMap;
bool SupportsAVX256;
ContextMemberInfo* FindMemberInfo(ContextInfo* ClassifiedInfo, uint32_t Offset, uint8_t Size);
ContextMemberInfo* RecordAccess(ContextMemberInfo* Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
LastAccessType AccessType, FEXCore::IR::Ref Node, FEXCore::IR::Ref StoreNode = nullptr);
ContextMemberInfo* RecordAccess(ContextInfo* ClassifiedInfo, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
LastAccessType AccessType, FEXCore::IR::Ref Node, FEXCore::IR::Ref StoreNode = nullptr);
void HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::Ref CodeNode, unsigned Flag);
// Classify context loads and stores.
void ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::NodeIterator BlockEnd);
void ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::Ref ValueNode);
// Block local Passes
void RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit);
unsigned OffsetForReg(FEXCore::IR::RegisterClassType Class, unsigned Reg, unsigned Size) {
if (Class == FEXCore::IR::FPRClass) {
return Size == 32 ? offsetof(FEXCore::Core::CPUState, xmm.avx.data[Reg][0]) : offsetof(FEXCore::Core::CPUState, xmm.sse.data[Reg][0]);
} else if (Reg == FEXCore::Core::CPUState::PF_AS_GREG) {
return offsetof(FEXCore::Core::CPUState, pf_raw);
} else if (Reg == FEXCore::Core::CPUState::AF_AS_GREG) {
return offsetof(FEXCore::Core::CPUState, af_raw);
} else {
return offsetof(FEXCore::Core::CPUState, gregs[Reg]);
}
}
};
ContextMemberInfo* RCLSE::FindMemberInfo(ContextInfo* ContextClassificationInfo, uint32_t Offset, uint8_t Size) {
return ContextClassificationInfo->Lookup.at(Offset);
}
ContextMemberInfo* RCLSE::RecordAccess(ContextMemberInfo* Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
LastAccessType AccessType, FEXCore::IR::Ref ValueNode, FEXCore::IR::Ref StoreNode) {
LOGMAN_THROW_AA_FMT((Offset + Size) <= (Info->Class.Offset + Info->Class.Size), "Access to context item went over member size");
LOGMAN_THROW_AA_FMT(Info->Accessed != LastAccessType::INVALID, "Tried to access invalid member");
// If we aren't fully overwriting the member then it is a partial write that we need to track
if (Size < Info->Class.Size) {
AccessType = AccessType == LastAccessType::WRITE ? LastAccessType::PARTIAL_WRITE : LastAccessType::PARTIAL_READ;
}
if (Size > Info->Class.Size) {
LOGMAN_MSG_A_FMT("Can't handle this");
}
Info->Accessed = AccessType;
Info->AccessRegClass = RegClass;
Info->AccessOffset = Offset;
Info->AccessSize = Size;
Info->ValueNode = ValueNode;
if (StoreNode != nullptr) {
Info->StoreNode = StoreNode;
}
return Info;
}
ContextMemberInfo* RCLSE::RecordAccess(ContextInfo* ClassifiedInfo, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
LastAccessType AccessType, FEXCore::IR::Ref ValueNode, FEXCore::IR::Ref StoreNode) {
ContextMemberInfo* Info = FindMemberInfo(ClassifiedInfo, Offset, Size);
return RecordAccess(Info, RegClass, Offset, Size, AccessType, ValueNode, StoreNode);
}
void RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
uint32_t Offset, uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::NodeIterator BlockEnd) {
auto Info = FindMemberInfo(LocalInfo, Offset, Size);
ContextMemberInfo PreviousMemberInfoCopy = *Info;
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, CodeNode);
if (PreviousMemberInfoCopy.AccessRegClass == Info->AccessRegClass && PreviousMemberInfoCopy.AccessOffset == Info->AccessOffset &&
PreviousMemberInfoCopy.AccessSize == Size) {
// This optimizes two cases:
// - Previous access was a load, and we have a redundant load of the same value.
// - Previous access was a store, and we are redundantly loading immediately after the store. Eliminating the store.
IREmit->ReplaceAllUsesWithRange(CodeNode, PreviousMemberInfoCopy.ValueNode, IREmit->GetIterator(IREmit->WrapNode(CodeNode)), BlockEnd);
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, PreviousMemberInfoCopy.ValueNode);
}
// TODO: Optimize the case of partial loads.
}
void RCLSE::ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
uint32_t Offset, uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::Ref ValueNode) {
auto Info = FindMemberInfo(LocalInfo, Offset, Size);
ContextMemberInfo PreviousMemberInfoCopy = *Info;
RecordAccess(Info, Class, Offset, Size, LastAccessType::WRITE, ValueNode, CodeNode);
if (PreviousMemberInfoCopy.AccessRegClass == Info->AccessRegClass && PreviousMemberInfoCopy.AccessOffset == Info->AccessOffset &&
PreviousMemberInfoCopy.AccessSize == Size && PreviousMemberInfoCopy.Accessed == LastAccessType::WRITE) {
// This optimizes redundant stores with no intervening load
// TODO: this is causing RA to fall over in some titles, disabling for now.
// Revisit when the new RA lands.
#if 0
IREmit->Remove(PreviousMemberInfoCopy.StoreNode);
#endif
}
// TODO: Optimize the case of partial stores.
}
void RCLSE::HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::Ref CodeNode, unsigned Flag) {
const auto FlagOffset = offsetof(FEXCore::Core::CPUState, flags[Flag]);
auto Info = FindMemberInfo(LocalInfo, FlagOffset, 1);
LastAccessType LastAccess = Info->Accessed;
auto LastValueNode = Info->ValueNode;
if (IsWriteAccess(LastAccess)) { // 1 byte so always a full write
// If the last store matches this load value then we can replace the loaded value with the previous valid one
IREmit->SetWriteCursor(CodeNode);
IREmit->ReplaceAllUsesWith(CodeNode, LastValueNode);
RecordAccess(Info, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::READ, LastValueNode);
} else if (IsReadAccess(LastAccess)) {
IREmit->ReplaceAllUsesWith(CodeNode, LastValueNode);
RecordAccess(Info, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::READ, LastValueNode);
}
}
/**
* @brief This pass removes redundant pairs of storecontext and loadcontext ops
*
* eg.
* %26 i128 = LoadMem %25 i64, 0x10
* (%%27) StoreContext %26 i128, 0x10, 0xb0
* %28 i128 = LoadContext 0x10, 0x90
* %29 i128 = LoadContext 0x10, 0xb0
* Converts to
* %26 i128 = LoadMem %25 i64, 0x10
* (%%27) StoreContext %26 i128, 0x10, 0xb0
* %28 i128 = LoadContext 0x10, 0x90
*
* eg.
* %6 i128 = LoadContext 0x10, 0x90
* %7 i128 = LoadContext 0x10, 0x90
* %8 i128 = VXor %7 i128, %6 i128
* Converts to
* %6 i128 = LoadContext 0x10, 0x90
* %7 i128 = VXor %6 i128, %6 i128
*
* eg.
* (%%189) StoreContext %188 i128, 0x10, 0xa0
* %190 i128 = LoadContext 0x10, 0x90
* %192 i128 = VAdd %188 i128, %190 i128, 0x10, 0x4
* (%%193) StoreContext %192 i128, 0x10, 0xa0
* Converts to
* %173 i128 = LoadContext 0x10, 0x90
* %175 i128 = VAdd %172 i128, %173 i128, 0x10, 0x4
* (%%176) StoreContext %175 i128, 0x10, 0xa0
*/
void RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
using namespace FEXCore;
using namespace FEXCore::IR;
auto CurrentIR = IREmit->ViewIR();
auto OriginalWriteCursor = IREmit->GetWriteCursor();
// XXX: Walk the list and calculate the control flow
ContextInfo& LocalInfo = ClassifiedStruct;
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
auto BlockOp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
auto BlockEnd = IREmit->GetIterator(BlockOp->Last);
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STORECONTEXT) {
auto Op = IROp->CW<IR::IROp_StoreContext>();
ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
} else if (IROp->Op == OP_STOREREGISTER) {
auto Op = IROp->CW<IR::IROp_StoreRegister>();
auto Offset = OffsetForReg(Op->Class, Op->Reg, IROp->Size);
ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
} else if (IROp->Op == OP_LOADREGISTER) {
auto Op = IROp->CW<IR::IROp_LoadRegister>();
auto Offset = OffsetForReg(Op->Class, Op->Reg, IROp->Size);
ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Offset, IROp->Size, CodeNode, BlockEnd);
} else if (IROp->Op == OP_LOADCONTEXT) {
auto Op = IROp->CW<IR::IROp_LoadContext>();
ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, BlockEnd);
} else if (IROp->Op == OP_STOREFLAG) {
const auto Op = IROp->CW<IR::IROp_StoreFlag>();
const auto FlagOffset = offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag;
auto Info = FindMemberInfo(&LocalInfo, FlagOffset, 1);
auto LastStoreNode = Info->StoreNode;
RecordAccess(&LocalInfo, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::WRITE, CurrentIR.GetNode(Op->Header.Args[0]), CodeNode);
// Flags don't alias, so we can take the simple route here. Kill any flags that have been overwritten
if (LastStoreNode != nullptr) {
IREmit->Remove(LastStoreNode);
}
} else if (IROp->Op == OP_INVALIDATEFLAGS) {
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
// Loop through non-reserved flag stores and eliminate unused ones.
for (size_t F = 0; F < Core::CPUState::NUM_EFLAG_BITS; F++) {
if (!(Op->Flags & (1ULL << F))) {
continue;
}
const auto FlagOffset = offsetof(FEXCore::Core::CPUState, flags[0]) + F;
auto Info = FindMemberInfo(&LocalInfo, FlagOffset, 1);
auto LastStoreNode = Info->StoreNode;
// Flags don't alias, so we can take the simple route here. Kill any flags that have been invalidated without a read.
if (LastStoreNode != nullptr) {
IREmit->SetWriteCursor(CodeNode);
RecordAccess(&LocalInfo, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::WRITE, IREmit->_Constant(0), CodeNode);
IREmit->Remove(LastStoreNode);
}
}
} else if (IROp->Op == OP_LOADFLAG) {
const auto Op = IROp->CW<IR::IROp_LoadFlag>();
HandleLoadFlag(IREmit, &LocalInfo, CodeNode, Op->Flag);
} else if (IROp->Op == OP_LOADDF) {
HandleLoadFlag(IREmit, &LocalInfo, CodeNode, X86State::RFLAG_DF_RAW_LOC);
} else if (IROp->Op == OP_SYSCALL || IROp->Op == OP_INLINESYSCALL) {
FEXCore::IR::SyscallFlags Flags {};
if (IROp->Op == OP_SYSCALL) {
auto Op = IROp->C<IR::IROp_Syscall>();
Flags = Op->Flags;
} else {
auto Op = IROp->C<IR::IROp_InlineSyscall>();
Flags = Op->Flags;
}
if ((Flags & FEXCore::IR::SyscallFlags::OPTIMIZETHROUGH) != FEXCore::IR::SyscallFlags::OPTIMIZETHROUGH) {
// We can't track through these
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
}
} else if (IROp->Op == OP_STORECONTEXTINDEXED || IROp->Op == OP_LOADCONTEXTINDEXED || IROp->Op == OP_BREAK) {
// We can't track through these
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
}
}
}
IREmit->SetWriteCursor(OriginalWriteCursor);
}
void RCLSE::Run(FEXCore::IR::IREmitter* IREmit) {
FEXCORE_PROFILE_SCOPED("PassManager::RCLSE");
RedundantStoreLoadElimination(IREmit);
}
} // namespace
namespace FEXCore::IR {
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsAVX256) {
return fextl::make_unique<RCLSE>(SupportsAVX256);
}
} // namespace FEXCore::IR
@@ -42,17 +42,16 @@ struct ReadWriteKill {
};
struct Info {
ReadWriteKill flag;
ReadWriteKill reg;
};
/**
* @brief This is a temporary pass to detect simple multiblock dead flag/reg stores
* @brief This is a temporary pass to detect simple multiblock dead reg stores
*
* First pass computes which flags/regs are read and written per block
* First pass computes which regs are read and written per block
*
* Second pass computes which flags/regs are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead flags/regs across multiple blocks.
* Second pass computes which regs are stored, but overwritten by the next block(s).
* It also propagates this information a few times to catch dead regs across multiple blocks.
*
* Third pass removes the dead stores.
*
@@ -64,28 +63,14 @@ void DeadStoreElimination::Run(IREmitter* IREmit) {
fextl::vector<Info> InfoMap(CurrentIR.GetSSACount());
// Pass 1
// Compute flags/regs read/writes per block
// Compute regs read/writes per block
// This is conservative and doesn't try to be smart about loads after writes
{
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
BlockInfo.flag.writes |= 1UL << Op->Flag;
} else if (IROp->Op == OP_INVALIDATEFLAGS) {
auto Op = IROp->C<IR::IROp_InvalidateFlags>();
BlockInfo.flag.writes |= Op->Flags;
} else if (IROp->Op == OP_LOADFLAG) {
auto Op = IROp->C<IR::IROp_LoadFlag>();
BlockInfo.flag.reads |= 1UL << Op->Flag;
} else if (IROp->Op == OP_LOADDF) {
BlockInfo.flag.reads |= 1UL << X86State::RFLAG_DF_RAW_LOC;
} else if (IROp->Op == OP_STOREREGISTER) {
if (IROp->Op == OP_STOREREGISTER) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
BlockInfo.reg.writes |= RegBit(Op->Class, Op->Reg);
} else if (IROp->Op == OP_LOADREGISTER) {
@@ -111,11 +96,9 @@ void DeadStoreElimination::Run(IREmitter* IREmit) {
auto& TargetInfo = InfoMap[Op->Header.Args[0].ID().Value];
// stores to remove are written by the next block but not read
BlockInfo.flag.kill = TargetInfo.flag.writes & ~(TargetInfo.flag.reads) & ~BlockInfo.flag.reads;
BlockInfo.reg.kill = TargetInfo.reg.writes & ~(TargetInfo.reg.reads) & ~BlockInfo.reg.reads;
// Flags that are written by the next block can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
// If written by the next block can be considered as written by this block, if not read
BlockInfo.reg.writes |= BlockInfo.reg.kill & ~BlockInfo.reg.reads;
} else if (IROp->Op == OP_CONDJUMP) {
auto Op = IROp->C<IR::IROp_CondJump>();
@@ -125,14 +108,10 @@ void DeadStoreElimination::Run(IREmitter* IREmit) {
auto& FalseTargetInfo = InfoMap[Op->FalseBlock.ID().Value];
// stores to remove are written by the next blocks but not read
BlockInfo.flag.kill = TrueTargetInfo.flag.writes & ~(TrueTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
BlockInfo.reg.kill = TrueTargetInfo.reg.writes & ~(TrueTargetInfo.reg.reads) & ~BlockInfo.reg.reads;
BlockInfo.flag.kill &= FalseTargetInfo.flag.writes & ~(FalseTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
BlockInfo.reg.kill &= FalseTargetInfo.reg.writes & ~(FalseTargetInfo.reg.reads) & ~BlockInfo.reg.reads;
// Flags that are written by the next blocks can be considered as written by this block, if not read
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
// if written by the next blocks can be considered as written by this block, if not read
BlockInfo.reg.writes |= BlockInfo.reg.kill & ~BlockInfo.reg.reads;
}
}
@@ -145,14 +124,7 @@ void DeadStoreElimination::Run(IREmitter* IREmit) {
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
if (IROp->Op == OP_STOREFLAG) {
auto Op = IROp->C<IR::IROp_StoreFlag>();
// If this StoreFlag is never read, remove it
if (BlockInfo.flag.kill & (1UL << Op->Flag)) {
IREmit->Remove(CodeNode);
}
} else if (IROp->Op == OP_STOREREGISTER) {
if (IROp->Op == OP_STOREREGISTER) {
auto Op = IROp->C<IR::IROp_StoreRegister>();
// If this OP_STOREREGISTER is never read, remove it
@@ -461,7 +461,12 @@ private:
// If that scalar is free because it is killed by this instruction, it
// needs to be shuffled too, since the copy would clobber it.
for (auto s = 0; s < IR::GetRAArgs(Pivot->Op); ++s) {
const PhysicalRegister ClobberReg = SSAToReg[Pivot->Args[s].ID().Value];
// It is possible that the argument is to be remapped, but the actual
// remapping in the IR only happens later in the pass so we need to
// Map() explicitly. This can be hit with SRA shuffles.
Ref New = Map(IR->GetNode(Pivot->Args[s]));
const PhysicalRegister ClobberReg = SSAToReg[IR->GetID(New).Value];
if (ClobberReg.Class == GPRClass && ClobberReg.Reg == NewReg) {
Clobber = IR->GetNode(Pivot->Args[s]);
break;
File diff suppressed because it is too large. Load diff
+14 -21
View File
@@ -8,27 +8,21 @@ $end_info$
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/LogManager.h>
#include <FEXCore/fextl/fmt.h>
#include <FEXCore/fextl/vector.h>
#include <cstdarg>
#include <cstdio>
#include <malloc.h>
namespace LogMan {
namespace Throw {
fextl::vector<ThrowHandler> Handlers;
void InstallHandler(ThrowHandler Handler) {
Handlers.emplace_back(Handler);
ThrowHandler Handler {};
void InstallHandler(ThrowHandler _Handler) {
Handler = _Handler;
}
void UnInstallHandlers() {
Handlers.clear();
void UnInstallHandler() {
Handler = nullptr;
}
void MFmt(const char* fmt, const fmt::format_args& args) {
auto msg = fextl::fmt::vformat(fmt, args);
for (auto& Handler : Handlers) {
if (Handler) {
auto msg = fextl::fmt::vformat(fmt, args);
Handler(msg.c_str());
}
@@ -37,18 +31,17 @@ namespace Throw {
} // namespace Throw
namespace Msg {
fextl::vector<MsgHandler> Handlers;
void InstallHandler(MsgHandler Handler) {
Handlers.emplace_back(Handler);
MsgHandler Handler {};
void InstallHandler(MsgHandler _Handler) {
Handler = _Handler;
}
void UnInstallHandlers() {
Handlers.clear();
void UnInstallHandler() {
Handler = nullptr;
}
void MFmtImpl(DebugLevels level, const char* fmt, const fmt::format_args& args) {
const auto msg = fextl::fmt::vformat(fmt, args);
for (auto& Handler : Handlers) {
if (Handler) {
const auto msg = fextl::fmt::vformat(fmt, args);
Handler(level, msg.c_str());
}
}
+3 -3
View File
@@ -24,12 +24,12 @@ namespace FEXCore::Utils::SpinWaitLock {
#define LOADEXCLUSIVE(LoadExclusiveOp, RegSize) \
/* Prime the exclusive monitor with the passed in address. */ \
#LoadExclusiveOp " %" #RegSize "[Result], [%[Futex]];"
#LoadExclusiveOp " %" #RegSize "[Result], [%[Futex]];\n"
#define SPINLOOP_BODY(LoadAtomicOp, RegSize) \
/* WFE will wait for either the memory to change or spurious wake-up. */ \
"wfe;" /* Load with acquire to get the result of memory. */ \
#LoadAtomicOp " %" #RegSize "[Result], [%[Futex]]; "
"wfe;\n" /* Load with acquire to get the result of memory. */ \
#LoadAtomicOp " %" #RegSize "[Result], [%[Futex]];\n"
#define SPINLOOP_WFE_LDX_8BIT LOADEXCLUSIVE(ldaxrb, w)
#define SPINLOOP_WFE_LDX_16BIT LOADEXCLUSIVE(ldaxrh, w)
+1 -5
View File
@@ -16,11 +16,10 @@
namespace FEXCore::Telemetry {
#ifndef FEX_DISABLE_TELEMETRY
static std::array<Value, FEXCore::Telemetry::TelemetryType::TYPE_LAST> TelemetryValues = {{}};
std::array<Value, FEXCore::Telemetry::TelemetryType::TYPE_LAST> TelemetryValues = {{}};
const std::array<std::string_view, FEXCore::Telemetry::TelemetryType::TYPE_LAST> TelemetryNames {
"64byte Split Locks",
"16byte Split atomics",
"VEX instructions (AVX)",
"EVEX instructions (AVX512)",
"16bit CAS Tear",
"32bit CAS Tear",
@@ -80,8 +79,5 @@ void Shutdown(const fextl::string& ApplicationName) {
}
}
Value& GetTelemetryValue(TelemetryType Type) {
return TelemetryValues.at(Type);
}
#endif
} // namespace FEXCore::Telemetry
+1
View File
@@ -187,6 +187,7 @@ public:
}
void Set(ConfigOption Option, const char* Data) {
LOGMAN_THROW_AA_FMT(Data != nullptr, "Data can't be null");
OptionMap[Option].emplace_back(fextl::string(Data));
}
+5 -12
View File
@@ -16,10 +16,11 @@
#include <ostream>
#include <mutex>
#include <shared_mutex>
#include <span>
namespace FEXCore {
class CodeLoader;
class HostFeatures;
struct HostFeatures;
class ForkableSharedMutex;
} // namespace FEXCore
@@ -101,7 +102,7 @@ using CodeRangeInvalidationFn = std::function<void(uint64_t start, uint64_t Leng
using CustomCPUFactoryType = std::function<fextl::unique_ptr<CPU::CPUBackend>(Context*, Core::InternalThreadState* Thread)>;
using CustomIREntrypointHandler = std::function<void(uintptr_t Entrypoint, IR::IREmitter*)>;
using ExitHandler = std::function<void(uint64_t ThreadId, ExitReason)>;
using ExitHandler = std::function<void(Core::InternalThreadState* Thread, ExitReason)>;
using AOTIRCodeFileWriterFn = std::function<void(const fextl::string& fileid, const fextl::string& filename)>;
using AOTIRLoaderCBFn = std::function<int(const fextl::string&)>;
@@ -118,7 +119,7 @@ public:
*
* @return a new context object
*/
FEX_DEFAULT_VISIBILITY static fextl::unique_ptr<FEXCore::Context::Context> CreateNewContext();
FEX_DEFAULT_VISIBILITY static fextl::unique_ptr<FEXCore::Context::Context> CreateNewContext(const FEXCore::HostFeatures& Features);
/**
* @brief Allows setting up in memory code and other things prior to launchign code execution
@@ -163,14 +164,6 @@ public:
*/
FEX_DEFAULT_VISIBILITY virtual void SetCustomCPUBackendFactory(CustomCPUFactoryType Factory) = 0;
/**
* @brief Retrieves a feature struct indicating certain supported aspects from
* the hose.
*
* @param CTX A valid non-null context instance.
*/
FEX_DEFAULT_VISIBILITY virtual HostFeatures GetHostFeatures() const = 0;
FEX_DEFAULT_VISIBILITY virtual void HandleCallback(FEXCore::Core::InternalThreadState* Thread, uint64_t RIP) = 0;
///< State reconstruction helpers
@@ -264,7 +257,7 @@ public:
* @param CTX A valid non-null context instance.
* @param Definitions A vector of thunk definitions that the frontend controls
*/
FEX_DEFAULT_VISIBILITY virtual void AppendThunkDefinitions(const fextl::vector<FEXCore::IR::ThunkDefinition>& Definitions) = 0;
FEX_DEFAULT_VISIBILITY virtual void AppendThunkDefinitions(std::span<const FEXCore::IR::ThunkDefinition> Definitions) = 0;
FEX_DEFAULT_VISIBILITY virtual void GetVDSOSigReturn(VDSOSigReturn* VDSOPointers) = 0;
+6 -10
View File
@@ -3,7 +3,6 @@
#include <FEXCore/Core/X86Enums.h>
#include <FEXCore/IR/IR.h>
#include <FEXCore/HLE/Linux/ThreadManagement.h>
#include <FEXCore/Utils/CompilerDefs.h>
#include <FEXCore/Utils/Telemetry.h>
@@ -104,7 +103,7 @@ struct CPUState {
// Raw segment register indexes
uint16_t es_idx {}, cs_idx {}, ss_idx {}, ds_idx {};
uint16_t gs_idx {}, fs_idx {};
uint16_t _pad2[2];
uint32_t mxcsr {};
// Segment registers holding base addresses
uint32_t es_cached {}, cs_cached {}, ss_cached {}, ds_cached {};
@@ -162,6 +161,10 @@ struct CPUState {
// we encode DF as 1/-1 within the JIT, so we have to write 0x1 here to
// zero DF.
flags[X86State::RFLAG_DF_RAW_LOC] = 0x1;
// Default mxcsr value
// All exception masks enabled.
mxcsr = 0x1F80;
}
};
static_assert(std::is_trivially_copyable_v<CPUState>, "Needs to be trivial");
@@ -185,14 +188,7 @@ enum FallbackHandlerIndex {
OPINDEX_F80CVTINT_TRUNC2,
OPINDEX_F80CVTINT_TRUNC4,
OPINDEX_F80CVTINT_TRUNC8,
OPINDEX_F80CMP_0,
OPINDEX_F80CMP_1,
OPINDEX_F80CMP_2,
OPINDEX_F80CMP_3,
OPINDEX_F80CMP_4,
OPINDEX_F80CMP_5,
OPINDEX_F80CMP_6,
OPINDEX_F80CMP_7,
OPINDEX_F80CMP,
OPINDEX_F80CVTTOINT_2,
OPINDEX_F80CVTTOINT_4,
Loaded 100 of 316 files, more files were not shown because too many files have changed in this diff. Show more