Compare commits

..
437 Commits
Author SHA1 Message Date
Ryan Houdek 3ba84ad06a Docs: Update for release FEX-2507 2025-07-07 23:49:56 -07:00
Ryan Houdek c6aae9e05a Merge pull request #4634 from Sonicadvance1/fix_horizon
EmulatedFiles: Emulate `current_clocksource`
2025-07-07 21:16:54 -07:00
Ryan Houdek 95b4618833 Merge pull request #4644 from ChanthMiao/fix/sigframe_mistake
Fix: wrong magic value in fpstate.
2025-07-06 17:07:31 -07:00
Changwei Miao 1fe17d55d9 Fix: wrong magic value in fpstate.
FEX should only set fpx_sw_bytes.magic1 with FP_XSTATE_MAGIC when
avx is enabled. Otherwise it may cause segfault in ntdll::save_context,
which requires access to extended xstate info if magic1 equals FP_XSTATE_MAGIC.

Signed-off-by: Changwei Miao <chanthmiao@outlook.com>
2025-07-06 18:44:37 +08:00
Ryan Houdek ead73371d9 EmulatedFiles: Emulate current_clocksource
WINE uses this to determine TSC frequency and because it doesn't say
`tsc` on ARM devices, it was ignoring TSC and instead using CPU maximum
frequency.

This was causing Horizon to think the TSC ran at whatever the max
frequency of a core was  (1.8Ghz to 2.6Ghz depending?) This was causing
all of Horizon Zero Dawn's physics to run at slower than real time
speeds because our 1Ghz (on Orion) TSC is significantly lower than the
max clock speeds of the cores.

This is still a bug in Wine that it is using the maximum CPU clock speed
in the case of current_clocksource not being TSC, but that's a battle
for a different time.
2025-07-05 21:10:53 -07:00
Ryan Houdek 6f089a4323 Merge pull request #4643 from tstellar/llvm-21
Fix build with LLVM >= 21
2025-07-05 16:51:38 -07:00
Ryan Houdek 640f024551 Merge pull request #4641 from Sonicadvance1/optimize_sincos
JIT: Optimize x87 FSINCOS
2025-07-05 16:26:37 -07:00
Tom Stellard 99920f89dd Fix build with LLVM >= 21 2025-07-05 16:50:34 +00:00
Ryan Houdek 8ea276267f InstcountCI: Update 2025-07-03 18:05:52 -07:00
Ryan Houdek 6ebbd91245 JIT: Optimize x87 FSINCOS
Turns out Bayonetta hammers SINCOS, our splitting the operation is
actually harming the performance of games that heavily use FSINCOS. We
instead can actually combine the operation which improves performance.
Not enough to get the game running full speed consistently on my Radxa,
but good numbers in my microbenchmark.

```
Test, Total Cycles, Total Runs, Cycles Average, Internal Loops, Average cycles per internal, per/second

64-bit:
Before:
FSIN, 2691031290, 50000, 53820.6, 1000, 53.8206, 18580.237319
FCOS, 2719397120, 50000, 54387.9, 1000, 54.3879, 18386.428239
FSINCOS, 5586917530, 50000, 111738, 1000, 111.738, 8949.478801

After:
FSIN, 2669959250, 50000, 53399.2, 1000, 53.3992, 18726.877573
FCOS, 2740942260, 50000, 54818.8, 1000, 54.8188, 18241.901965
FSINCOS, 3189472870, 50000, 63789.5, 1000, 63.7895, 15676.571659

80-bit:
Before:
FSIN, 24702939380, 50000, 494059, 1000, 494.059, 2024.050629
FCOS, 19127131020, 50000, 382543, 1000, 382.543, 2614.087808
FSINCOS, 40386785260, 50000, 807736, 1000, 807.736, 1238.028719

After:
FSIN, 24869980710, 50000, 497400, 1000, 497.4, 2010.455922
FCOS, 19131849590, 50000, 382637, 1000, 382.637, 2613.443084
FSINCOS, 38329985570, 50000, 766600, 1000, 766.6, 1304.461749

Improvement 64-bit: 1.75x
Improvement 80-bit: 1.05x
```

Only a minor improvement at 80-bit precision since cephes doesn't provide a combined sincos operation, but the f64 implementation is significantly improved, allowing 75% more operations per second.

Disabled in the simulator because we can't easily support pairs of
vector registers being returned.
2025-07-03 18:05:51 -07:00
Ryan Houdek fa0a54deb9 Merge pull request #4640 from alyssarosenzweig/bug/fix-hades
Fix Hades
2025-07-03 14:53:01 -07:00
Alyssa Rosenzweig 046043090f unittests: add move merging test
this hits a nasty case with post-RA merging. fails on main.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:56 -04:00
Alyssa Rosenzweig 61150a18cc RegisterAllocationPass: fix bookkeeping with merging
this fixes Hades.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-03 17:22:53 -04:00
Ryan Houdek 95ca20cfee Merge pull request #4636 from alyssarosenzweig/bug/ra-invariant
RegisterAllocationPass: assert an invariant in post-RA prop
2025-07-02 18:17:08 -07:00
Alyssa Rosenzweig 5a536d47fd RegisterAllocationPass: assert an invariant in post-RA prop
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:53:47 -04:00
Ryan Houdek afbc7da027 Merge pull request #4629 from alyssarosenzweig/opt/cpuid-basic
Optimize some constant cpuid/xgetbv cases
2025-07-02 10:25:01 -07:00
Alyssa Rosenzweig c093c08c40 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 360d8c629e RegisterAllocationPass: optimize cpuid
for constant function where we don't have a leaf. this isn't fully general but
we can't do better without a more general post-RA optimizer. i'm not inclined to
do that unless/until we get hot blocks demonstrating its value (that we can
compare against the JIT time hit of the heavier-duty optimizer.)

however this special case we can (and should) optimize for now.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig abb41d39e4 RegisterAllocationPass: optimize xgetbv
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig d4eb4ef594 IR: plumb CPUID into RA pass
for cpuid folding.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Alyssa Rosenzweig 02f45854e8 IR: include a fence in CPUID
easier for post-RA to chew thru.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-07-02 13:04:49 -04:00
Ryan Houdek 62de1004df Merge pull request #4635 from neobrain/fix_libfwd_wl_more
LibraryForwarding/wayland: Add new interface objects
2025-07-02 09:01:11 -07:00
Tony Wasserka 7f216ca02f LibraryForwarding/wayland: Add new interface objects 2025-07-02 16:51:03 +02:00
LC 492b0fdda8 Merge pull request #4632 from neobrain/feature_nix
Build: Add nix-based helpers to facilitate cross-compilation
2025-07-01 16:23:58 -04:00
LC bb072c0112 Merge pull request #4633 from Sonicadvance1/noexec_test
unittests: Adds unittest for no-exec testing
2025-07-01 16:20:44 -04:00
Ryan Houdek afabe7cb47 Merge pull request #4627 from alyssarosenzweig/opt/long-div-peephole-ready
Optimize long division
2025-06-30 13:59:48 -07:00
Ryan Houdek 38e0fc2434 unittests: Adds unittest for no-exec testing
In preparation for #4474
2025-06-30 13:19:33 -07:00
Alyssa Rosenzweig 16a70eafc6 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:59:27 -04:00
Alyssa Rosenzweig de4becc26e OpcodeDispatcher: mask certain divisors
needed for fusing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Alyssa Rosenzweig af23f4325f OpcodeDispatcher: reorder xor-with-self sequence
this lets us peephole fuse things even when there are flags calculated in the
way

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-30 15:58:33 -04:00
Ryan Houdek 94af96df8f Merge pull request #4631 from neobrain/fix_libfwd_wl_cutter
LibraryForwarding: Fix various Wayland issues
2025-06-30 11:38:13 -07:00
Ryan Houdek e685ab818e Merge pull request #4619 from neobrain/refactor_drop_config_h_in
Remove code generation build step for install prefix
2025-06-30 11:35:26 -07:00
Ryan Houdek 5f2a72b65b Merge pull request #4624 from Sonicadvance1/static_analysis_fixes
Some static code analysis fixes
2025-06-30 11:28:06 -07:00
Tony Wasserka c1842a6167 Build: Add nix-based helpers to manage toolchains for ARM64EC/WOW64
The cmake_configure_woa*.sh scripts will automatically install any required
cross-toolchains required to enable either ARM64EC or WOW64 builds of FEX,
and it will initialize the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell WineOnArm/shell.nix`,
which will make the toolchain available via environment variables. This also
generates a meson crossfile for building VKD3D or vkd3d-proton.
2025-06-30 16:50:13 +02:00
Tony Wasserka 3f3907b5d1 Build: Add nix-based helpers to manage toolchains for FEXLinuxTests
The cmake_enable_flt.sh script will automatically install the required
cross-toolchains required to build FEXLinuxTests, and it will reconfigure
the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell FEXLinuxTests/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 1899465390 Build: Add nix-based helpers to manage toolchains for library forwarding
The cmake_enable_libfwd.sh script will automatically install any required
cross- toolchains and development headers required to enable library
forwarding, and it will reconfigure the current build folder appropriately.

For advanced uses, a shell can be opened by running `nix-shell LibraryForwarding/shell.nix`,
which will make the toolchain available via environment variables.
2025-06-30 16:50:13 +02:00
Tony Wasserka 18360d4ccb LibraryForwarding/wayland: Add more method signatures 2025-06-27 10:57:10 +02:00
Tony Wasserka a691c3cd99 LibraryForwarding/wayland: Fix mprotect call when crossing page boundaries 2025-06-27 10:57:10 +02:00
Tony Wasserka 70bc561bbf LibraryForwarding/wayland: Fix wl_proxy_marshal_array_constructor 2025-06-27 10:57:10 +02:00
Tony Wasserka b9222d8431 LibraryForwarding/unittests: Fix build with clang 20 2025-06-27 10:57:10 +02:00
Tony Wasserka a5fad89e57 Merge pull request #4621 from neobrain/feature_update_readme
Update Readme.md
2025-06-26 21:54:53 +02:00
Tony Wasserka a14360b89d Update Readme.md 2025-06-26 21:41:23 +02:00
Tony Wasserka c9aaedd217 CPack: Update package description 2025-06-26 21:41:23 +02:00
Alyssa Rosenzweig cda15ce9ea RegisterAllocationPass: optimize long divsion
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 3f3b6ad337 IR: merge regular/long divisions
this makes it a lot easier to turn a long division into a non-long division,
just by nulling out a source.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:29 -04:00
Alyssa Rosenzweig 62410c4381 InstCountCI: add another udiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-24 16:10:22 -04:00
Ryan Houdek cf82b56dd8 CPUBackend: Remove unused variable
CID 482006
2025-06-19 16:51:30 -07:00
Ryan Houdek 4a74bea7ab InstcountCI: Don't use a global static initializer for CodeSizeValidation
Relies on fmt facet initialization order which isn't guaranteed to have
correct initialization order.

CID 482003
2025-06-19 16:51:26 -07:00
Ryan Houdek c6d8e60ef8 OpcodeDispatcher: Make sure to initialize ArithRef
CID 482002
2025-06-19 16:46:15 -07:00
Ryan Houdek 1212cd526a VDSOEmulation: Sanitize sysconf result
Unlikely to fail but make sure.

CID 482021
2025-06-19 16:43:53 -07:00
Ryan Houdek 79a8ed53b6 CPUBackend: Make sure to zero initialize variable
CID 482022
2025-06-19 16:41:24 -07:00
Ryan Houdek a6ce115d9c FEXLoader: Handle bad LogFile path
Just go silent but print a log about it.

CID 482031
CID 482023
2025-06-19 16:40:03 -07:00
Ryan Houdek 646a5a7f9e CodeEmitter/ASIMD: Removes redundant ternary selection
Redundant and unnecessary.

CID 482017
2025-06-19 16:31:08 -07:00
Ryan Houdek 7f71b6f1b2 Passes: Use move instead of copy semantics
To initialize in-place.

CID 482035
2025-06-19 16:22:39 -07:00
Ryan Houdek 16d5ca447f FEXCore: DebugData is never null now
This is a required data structure to exist.

CID 482036
2025-06-19 16:21:19 -07:00
Ryan Houdek 3d0c20a263 Merge pull request #4615 from neobrain/feature_logging_qol
Improve log message formatting
2025-06-19 12:50:44 -07:00
Ryan Houdek 1c38b8b046 Merge pull request #4614 from alyssarosenzweig/opt/easy-mov-elim
Merge moves that are immediately consumed
2025-06-19 12:26:07 -07:00
Alyssa Rosenzweig e1124480be InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 5243f50ed1 RegisterAllocationPass: merge 32-bit mov + 64-bit and
mov wA, wB
  and xA, xA, ...

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig cde805147f RegisterAllocationPass: merge 32-bit moves
mov wA, wB
  op wA, wA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 7e39eb3df2 RegisterAllocationPass: merge full size moves
mov xA, xB
  op xA, xA, ..

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig af366d4480 RegisterAllocationPass: skip inlineconstant in RA
similar reasoning as guestopcode.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 33ef98aae7 RegisterAllocationPass: refactor push/pop merge
to make way for move merging.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig c16db2db4a RegisterAllocationPass: simplify an expression
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 059d980c33 IR: give StoreRegister a precoloured destination
this will eliminate an annoying special case in post-RA opts.

No difference proven at 95.0% confidence

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Alyssa Rosenzweig 581381fd86 IR: make 0 the invalid physical register
so zero init works as expected

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-19 13:44:23 -04:00
Ryan Houdek 0072b289bb Merge pull request #4622 from neobrain/feature_fexconfig_logo
FEXConfig: Add icon
2025-06-18 01:46:30 -07:00
Tony Wasserka e2bd79087e FEXConfig: Add icon 2025-06-18 10:33:07 +02:00
Ryan Houdek 6cd78fc90d Merge pull request #4618 from neobrain/feature_infer_32bit_libfwd_paths
LibraryForwarding: Infer folders for 32-bit wrappers automatically
2025-06-17 17:55:38 -07:00
Ryan Houdek c346aca241 Merge pull request #4620 from neobrain/refactor_data_files
Move data files to Data/
2025-06-17 17:54:28 -07:00
Ryan Houdek bf83569f0b Merge pull request #4623 from neobrain/feature_tracy_0_12
External: Update Tracy submodule to version 0.12.1
2025-06-17 17:53:47 -07:00
Tony Wasserka 1502f04a8a External: Update Tracy submodule to version 0.12.1
Notably, this update adds flame graph functionality for aggregated data.
2025-06-17 17:51:52 +02:00
Tony Wasserka 7bd9d0ae23 Data: Move Dockerfile 2025-06-17 16:40:42 +02:00
Tony Wasserka cf57afdf26 Data: Move CI folder to Data/CI 2025-06-17 16:40:42 +02:00
Tony Wasserka 43e6aebc7a Data: Move CMake support scripts to Data/CMake 2025-06-17 16:40:42 +02:00
Tony Wasserka 22780993e1 Data: Move CPack files to Data/CMake/ 2025-06-17 16:40:42 +02:00
Tony Wasserka 578dcee9af Data: Move toolchain files to a central location 2025-06-17 16:40:42 +02:00
Tony Wasserka 61d77e3f9b LibraryForwarding: Infer folders for 32-bit wrappers automatically
There's no need to bother the user to select these paths manually.
Instead, just use the same folder names with _32 appended.

Fixes #4588.
2025-06-17 11:57:35 +02:00
Tony Wasserka 4ce0acba80 Remove now unused Config.h.in 2025-06-17 11:40:32 +02:00
Tony Wasserka d137212222 FEXGetConfig: Infer install prefix from executable path 2025-06-17 11:40:32 +02:00
Tony Wasserka 9ad4e3a6a0 FEXBash: Clean up and fix FEXInterpreter lookup
Previously, the first attempt to look up a FEXInterpreter would always
fail due to a missing path separator.

Additionally, fallback lookup now uses /proc/self/exe to find a path relative
to the FEXBash executable. The previous use of FindContainerPrefix does not
seem to be required anymore in current Steam versions.
2025-06-17 11:40:32 +02:00
Tony Wasserka 57627d4fcf LogManager: Drop unused STDOUT/STDERR log levels 2025-06-16 13:54:03 +02:00
Tony Wasserka 23b69271eb Use consistent log message formatting for all modules 2025-06-16 13:54:03 +02:00
LC 9d2f557666 Merge pull request #4616 from alyssarosenzweig/ici/32bit-div
InstructionCountCI: add 32-bit division cases
2025-06-13 13:18:55 -04:00
Alyssa Rosenzweig 4f9e352ff0 InstructionCountCI: add 32-bit division cases
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-13 13:05:35 -04:00
Tony Wasserka 3e85e60a30 LogManager: Use colors for logging when possible 2025-06-13 15:14:47 +02:00
Tony Wasserka febce21b21 FEXServer: Separate PID and TID with a pipe instead of a period
Since most software considers the former a word boundary but not the latter,
this allows the individual values to be copy-pasted more conveniently.
2025-06-13 14:54:22 +02:00
Tony Wasserka 755364e2df FEXServer: Clean up time display for logging
This is now relative to the time of the first message. Furthermore, display
precision is limited to milliseconds (which are actually zero-padded now!).
2025-06-13 14:54:13 +02:00
Tony Wasserka df63979773 LogManager: Avoid ^C being printed when quitting foreground FEXServer 2025-06-13 14:31:25 +02:00
Tony Wasserka 992d86bbc1 LogManager: Shorten debug level strings to a single letter 2025-06-13 14:31:25 +02:00
Ryan Houdek 534b338161 Merge pull request #4612 from alyssarosenzweig/ir/merge-divrem-2
IR: merge integer division & remainder
2025-06-12 14:29:02 -07:00
Alyssa Rosenzweig ef250f936c InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig e57130e364 JIT: drop UDiv extensions
we already extend in the dispatcher.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig eedcb35270 IR: merge ldiv/lrem handlers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-12 16:45:49 -04:00
Alyssa Rosenzweig 1eb470083c IR: merge div/rem opcodes
it's simpler & faster to calculate both together, matching the x86 semantic.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-10 17:08:20 -04:00
LC 1f15a4e35b Merge pull request #4611 from Sonicadvance1/ubisoft_ptrace
Linux: Implement enough of ptrace to allow Ubisoft launcher
2025-06-10 11:35:49 -04:00
Ryan Houdek 7ed9bea16b Linux: Implement enough of ptrace to allow Ubisoft launcher
Wine does some minimal attaching, peeking, poking, and detaching to have
a different process inspect another's TEB region. Ubisoft's launcher
does this to check if a debugger is attached and rejects it in the case
that it is.

While this implementation isn't all-encompassing, it is good enough for
this family of games.
2025-06-09 16:32:42 -07:00
Ryan Houdek 8b1383d235 Docs: Update for release FEX-2506 2025-06-04 10:48:24 -07:00
LC a73fab3bb5 Merge pull request #4609 from Sonicadvance1/add_plucky_install
Scripts/InstallFEX: Add plucky
2025-06-03 16:30:45 -04:00
Ryan Houdek 7bd64d9c53 Merge pull request #4608 from alyssarosenzweig/ici/case
InstructionCountCI: fix sdiv case
2025-06-03 12:50:36 -07:00
Ryan Houdek 1786c2f157 Scripts/InstallFEX: Add plucky
This was added to the PPA last month.
2025-06-03 12:31:40 -07:00
Alyssa Rosenzweig 958b671736 InstructionCountCI: fix sdiv case
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:30:42 -04:00
Ryan Houdek 6549b66cf6 Merge pull request #4606 from alyssarosenzweig/opt/cdq
Optimize CDQ
2025-06-03 12:24:15 -07:00
Alyssa Rosenzweig 0c855a5ce3 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:39 -04:00
Alyssa Rosenzweig 6864d48dcf OpcodeDispatcher: optimize cdq
prereq to optimizing sign-ext+ldiv in a reasonable way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-03 15:00:06 -04:00
Ryan Houdek 99816a23a8 Merge pull request #4583 from neobrain/refactor_syscalls_unify
LinuxSyscalls: Reduce code duplication between 32-bit and 64-bit paths
2025-06-03 09:54:09 -07:00
Ryan Houdek 2ff9546523 Merge pull request #4599 from pmatos/FEXServerFind
Fixes FEXServer path search
2025-06-03 09:52:39 -07:00
Ryan Houdek 06541f21d6 Merge pull request #4605 from neobrain/fix_cmake_full_libdir2
Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values
2025-06-03 09:52:26 -07:00
Tony Wasserka 45a37edd4a Config: Use CMAKE_INSTALL_FULL_LIBDIR when templating default values 2025-06-03 17:45:29 +02:00
Ryan Houdek 2a713a1f51 Merge pull request #4604 from neobrain/fix_cmake_install_prefix
CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
2025-06-03 08:43:24 -07:00
Tony Wasserka 3e104de377 CMake: Generate DATA_DIRECTORY dynamically unless explicitly set
This allows changes to CMAKE_INSTALL_PREFIX to be automatically picked up and
propagated properly. Previously, you had to update 4 variables in 3 files by
hand to do so.
2025-06-03 16:58:06 +02:00
Paulo Matos b22b316e70 Removes handling of FEXServer from the FEXBash wrapper
This was not necessary since FEXInterpreter is already doing it,
and doing it properly. The previous implementation if FEXBash was
incomplete.
2025-06-03 12:24:03 +02:00
Tony Wasserka df461546c5 LinuxSyscalls: Make error return values consistent 2025-06-03 11:10:08 +02:00
Tony Wasserka 1c1c43cd86 LinuxSyscalls: Fix formatting 2025-06-03 11:10:08 +02:00
Tony Wasserka 3eac9f937e LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:08 +02:00
Tony Wasserka f13a0d8e84 LinuxSyscalls: Unify shmdt implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ca12dc9213 LinuxSyscalls: Unify shmat implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 365ed2cd70 LinuxSyscalls: Unify mprotect implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka 7efbfed0bf LinuxSyscalls: Unify mmap and munmap implementations 2025-06-03 11:10:07 +02:00
Tony Wasserka ed502738c6 LinuxSyscalls: Make Get32BitAllocator interface virtual 2025-06-03 11:10:07 +02:00
Ryan Houdek ba162bb058 Merge pull request #4603 from neobrain/fix_cmake_full_libdir
LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
2025-06-02 15:49:22 -07:00
Tony Wasserka f194d35913 LibraryForwarding: Use CMAKE_INSTALL_FULL_LIBDIR instead of constructing the library install paths manually
This fixes issues in the nix build, where CMAKE_INSTALL_LIBDIR is an
absolute path instead of the typical "lib(64)".
2025-06-03 00:37:23 +02:00
Ryan Houdek e81c84e00a Merge pull request #4602 from Sonicadvance1/add_instcountci_tests
InstcountCI: Adds tests for instructions discovered by #4597
2025-06-02 12:45:51 -07:00
Ryan Houdek 68dc9030bc Merge pull request #4601 from Sonicadvance1/fix_futimesat_flake
unittests/futimesat: Fixes flake
2025-06-02 12:45:26 -07:00
Ryan Houdek fedad275e7 Merge pull request #4597 from alyssarosenzweig/opt/drop-pile-of-constprop
Constant fold on the fly
2025-06-02 12:34:53 -07:00
Ryan Houdek f6b4c76d76 InstcountCI: Adds tests for instructions discovered by #4597
Apparently I completely missed that cpuid, xgetbv, syscall,
l{u,}{div,rem} were failing to hit their optimized cases for inlining
and avoiding 128-bit software divide.

The divisions are a clear performance regression for 64-bit applications
since that is the only real way to do a 64-bit division on x86, I added
those specifically because it sped up games.

CPUID depends heavily on the game, since some games use that as a
serialization instruction fairly heavily.

XGETBV is trivial since it matches behaviour of CPUID (and is basically
an extension of it).

Syscall inlining can save a decent amount of time, again heavily depends
on game.

Adds multi-inst tests for all of these situations so that once it gets
fixed (Apparently broken once RCLSE got stripped out), we can see that
they keep working. Obviously tests couldn't have existed in instcountCI
before since we didn't support multi-instruction tests.
2025-06-02 12:10:39 -07:00
Alyssa Rosenzweig 27854aa091 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig d966ae145e ConstProp: merge inline + pooling
now that the algebraic/folding opts are gone, we can do this in one pass for a
2.5% speedup:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44474704    0.46750433    0.45455258    0.45446569  0.0044727894
+  50    0.43149892    0.45173984    0.44252267    0.44295575  0.0045621814
Difference at 95.0% confidence
	-0.0115099 +/- 0.00179263
	-2.53263% +/- 0.394447%
	(Student's t, pooled s = 0.00451771)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:27:02 -04:00
Alyssa Rosenzweig 4cb37e6a1b ConstProp: drop constant folding and algebraic opts
No longer needed.

The total difference from the beginning of this series (all the prep work to
make this change possible) plus this commit is a modest 0.4% win.

    N           Min           Max        Median           Avg        Stddev
x 100     0.4472467    0.46646308    0.45708424    0.45713057  0.0040838243
+ 100    0.44707586    0.46581227    0.45479448    0.45509309  0.0037548573
Difference at 95.0% confidence
	-0.00203748 +/- 0.00108734
	-0.445711% +/- 0.237862%
	(Student's t, pooled s = 0.00392279)

...in addition to a net deletion of 144 lines of code.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:24:58 -04:00
Alyssa Rosenzweig d9da81e99b ConstProp: drop dead cross-instr opts
these can't work due to architectural limitations. they could be ported to
post-RA passes, I think, but having them here now is not helping anything and
they're in the way.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 7d734740be Addressing: avoid a Bfe
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 06299cca4b OpcodeDispatcher: optimize out Bfi for storereg
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig af5aaab38b OpcodeDispatcher: optimize a few ALU-with-constant ops
instead of relying on ConstProp for this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 77bb01d384 OpcodeDispatcher: don't generate pointless Xor for AF
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig a9eb1bd4e6 OpcodeDispatcher: don't generate pointless Bfe for moves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 83ef2da95f OpcodeDispatcher: avoid zero shift in SHLD
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ba13ccadb8 OpcodeDispatcher: use ArithRef for 8/16-bit imul
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 4a2dee873d OpcodeDispatcher: use ArithRef for small rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 589906e6e6 OpcodeDispatcher: generalize AF on constants optimization
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig f0fcf6d9e6 OpcodeDispatcher: optimize SVE vmovmskpd
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 1591ced5a7 OpcodeDispatcher: use ArithRef for ADC/SBB flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 0d235d63d0 OpcodeDispatcher: use ArithRef for bit tests
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 820d5c1447 OpcodeDispatcher: use ArithRef for rotates
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 8093e8f5f1 OpcodeDispatcher: optimize LoadDir
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 802857e411 OpcodeDispatcher: add ArithRef helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig bf3a1839a2 OpcodeDispatcher: transform DF ourselves
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig ad7844d7da OpcodeDispatcher: do not generate useless VMov
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Alyssa Rosenzweig 3ff9128a8a unittests: add asm test for mov ah, 0
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 14:15:51 -04:00
Ryan Houdek c865eb98ea Merge pull request #4598 from alyssarosenzweig/idc
Stop using hashmap in DCE
2025-06-02 10:26:08 -07:00
Ryan Houdek 217a9228f8 unittests/futimesat: Fixes flake
Due to interactions between file times and the lack of granularity in
futimesat, if we don't remove the nanoseconds then this test can flake.

Easy enough.
2025-06-02 10:21:42 -07:00
Alyssa Rosenzweig 105d0a36ad RedundantFlagCalculationElimination: dont use hashmap
combined results on node from this and the previous commit:

    N           Min           Max        Median           Avg        Stddev
x  50    0.44433582    0.46988457    0.45344824    0.45320041  0.0043288623
+  50     0.4385365    0.46615359    0.45026179    0.44996216  0.0045742233
Difference at 95.0% confidence
	-0.00323825 +/- 0.00176704
	-0.71453% +/- 0.389903%
	(Student's t, pooled s = 0.00445323)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Alyssa Rosenzweig 263279d5dd IR: index blocks
this will let us avoid a costly hashmap in DCE.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-06-02 13:02:44 -04:00
Ryan Houdek 861ecbec0f Merge pull request #4600 from neobrain/fix_libfwd_cross_target
LibraryForwarding: Change target env to gnu
2025-06-02 09:21:46 -07:00
Ryan Houdek 5ff9bb3669 Merge pull request #4596 from pmatos/JSONOpts
Error when JSON config includes unknown options
2025-06-02 08:50:47 -07:00
Ryan Houdek e794584bb5 Merge pull request #4479 from neobrain/feature_codebuffer_sharing
Reduce JIT time by 25% by sharing code buffers between threads
2025-06-02 08:50:17 -07:00
Tony Wasserka a792dd0703 CPUBackend: Clean up CodeBuffer size limits 2025-06-01 22:45:50 +02:00
Tony Wasserka a9a6a645bf Arm64Emitter: Disable PC-relative constant encoding
This no longer works since the JIT output is now relocated before execution.
2025-06-01 22:45:50 +02:00
Tony Wasserka 7c93becd5f JIT: Increase estimate for CodeBuffer space use
The previous bound was exceeded during Steam startup before.
2025-06-01 22:45:50 +02:00
Tony Wasserka 4bbaef58e9 JIT: Re-enable parallel compilation by compiling to a temporary buffer 2025-06-01 22:45:50 +02:00
Tony Wasserka 95791a985a Core: Minimize the time CodeBufferWriteMutex is held 2025-06-01 22:45:50 +02:00
Tony Wasserka 0dfefe9730 Core: Re-check LookupCache before running compiler backend
This further reduces lock contention by skipping the backend phase in case
another thread raced the active one for the same block.
2025-06-01 22:45:49 +02:00
Tony Wasserka 8481c797df FEXCore/JIT: Extend LookupCache lock to all of ExitFunctionLink 2025-06-01 22:44:49 +02:00
Tony Wasserka a503b5e20b Rename CodeBufferManager reference 2025-06-01 22:44:49 +02:00
Tony Wasserka 4078840ef1 Core: Reduce JIT time by sharing CodeBuffers between threads
This is changes the interface of CodeBuffer to that of a partially persistent
data structure based on reference counting:
- Exactly one CodeBuffer is now designated as "active", which means data can
  be *appended* to it
- Lossy modifications to the active CodeBuffer will not invalidate any data
  in use by other threads, which enables save sharing across threads
- Instead, such lossy modifications trigger a new "version" of the data in
  the modifying thread. Old versions of the CodeBuffer persist as read-only
  data for use by the other threads.
- The other threads can update their version of the CodeBuffer. This will
  decrease the reference count and eventually trigger deallocation of the
  old version
2025-06-01 22:44:49 +02:00
Tony Wasserka ab51958b26 Move AllocateNewCodeBuffer from CPUBackend to a new CodeBufferManager interface 2025-06-01 22:42:55 +02:00
Tony Wasserka 3376587b6a CPUBackend: Manage CodeBuffers using shared_ptr
This is required for sharing CodeBuffers between threads anyway, but it also
allows use of the constructor/destructor to manage memory automatically.
2025-06-01 22:42:55 +02:00
Tony Wasserka 6681d7dcf9 LookupCache: Split L3 cache into a dedicated interface
This data isn't really a cache, since the JIT is directly responsible of
writing its contents. Instead it be considered the source to populate the
L1/L2 caches from.

Furthermore, splitting off this data allows it to be shared across threads
in the future without affecting L1/L2 caches.
2025-06-01 22:42:55 +02:00
Tony Wasserka 34224481c6 LookupCache: Prefer empty() over a size check 2025-06-01 22:42:55 +02:00
Tony Wasserka 1837aaabe4 fextl: Add shared_ptr and make_shared 2025-06-01 22:42:55 +02:00
Tony Wasserka 5811914a78 LibraryForwarding: Change target env to gnu
This is required for clang to find architecture-specific libstdc++ headers as
distributed by NixOS.
2025-06-01 21:40:33 +02:00
Paulo Matos a109a4efa1 Error when JSON config includes unknown options 2025-06-01 11:59:51 +02:00
Alyssa Rosenzweig a08a6ce5de Merge pull request #4576 from Sonicadvance1/fix_vma_race
Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
2025-05-30 08:14:38 -04:00
Ryan Houdek 3dc8a3ddc1 Merge pull request #4595 from neobrain/refactor_config_templates
Config: Clean up use of templates
2025-05-29 12:26:34 -07:00
Ryan Houdek 5e103365f7 Merge pull request #4594 from neobrain/fix_self_move
Async: Don't destruct on self-moves
2025-05-29 12:26:24 -07:00
Ryan Houdek cdea8d7f74 Merge pull request #4593 from alyssarosenzweig/ici/zeroing-sub-regs
InstructionCountCI: add more cases for mov 0/~0
2025-05-29 12:26:15 -07:00
Ryan Houdek ef6dc3d802 Merge pull request #4592 from alyssarosenzweig/opt/x87-tag
OpcodeDispatcher: optimize X87FTWTag
2025-05-29 12:26:04 -07:00
Ryan Houdek e6edb349ba Linux: Fixes some vestigial mmap handling
Missed this in previous commits, had to double check that I got them
all.
2025-05-29 12:15:26 -07:00
Ryan Houdek ad132267ec Linux/SMCTracking: Fixes nasty race condition causing invalid memory tracking
This has been a bug that we have technically lived with ever since SMC
tracking was introduced. The problem boils down to the fact that memory
management syscalls from multiple threads can race our SMC tracking.

This was only uncovered due to recent changes in the Steam client where
downloading games has more aggressively started reallocating memory.
This causes Steam to oversubscribe the CPU by a small margin, causing
threads to context switch more heavily during memory management.

The strace that finally managed to capture this:
```
41574 munmap(0xba84e000, 724992 <unfinished ...>
<...>
41227 mmap(NULL, 540672, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -3, 0 <unfinished ...>
<...>
41574 <... munmap resumed>)             = 0
<...>
41227 <... mmap resumed>)               = 0xba87b000
```

While FEX's tracking linearly was:
```
mmap, 0xba87b000, 0x84000, 0x3, 0x22, 0xfffffffd, 0x0
munmap, 0xba84e000, 0xb1000
```

The way munmap and mmap perfectly interleave while getting context switched meant that the kernel's view of munmap then mmap didn't match our view of mmap completing first then munmap happening afterwards.
The kernel/strace is obviously the correct view in this instance.

This all comes down to how these threads are racing the VMA tracking
mutex after the syscall happens and not guaranteeing sequential
consistency that matches the kernel's view.

The only way to correct this sanely is to extend the locking period to
also encompass the syscalls getting executed. This is a bit tricky since
the VMA tracking needs to ensure that the lock is no longer held once
ThreadManager invalidation occurs so a callback to do the syscall
operation is about the only sane approach here. Luckily we now have
fextl::move_only_function.

Fixes consistent crashes with Steam game downloads (and maybe some
chromium crashes?)
2025-05-29 12:15:26 -07:00
Ryan Houdek 6f837281ef Merge pull request #4591 from Sonicadvance1/fix_warning2
FEXCore/CPUID: Remove warning
2025-05-29 11:21:49 -07:00
Tony Wasserka 2a1d29d2df Config: Clean up use of templates 2025-05-29 18:38:35 +02:00
Tony Wasserka bdceb4ca89 Async: Don't destruct on self-moves
FEXServer's logger performs such self-moves for the log pipe posix_descriptor
of short-lived clients. This resulted in a double-close previously, which
could interfere with file operations on other threads (typically crashing
FEXServer in effect).
2025-05-29 18:28:56 +02:00
Alyssa Rosenzweig 656bb928cf InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig 0fbe69ebcf OpcodeDispatcher: optimize xor-with-self flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:53:38 -04:00
Alyssa Rosenzweig ece817c691 OpcodeDispatcher: optimize logical flags
seems to be strictly better.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:34:50 -04:00
Alyssa Rosenzweig 3fbc8204b7 OpcodeDispatcher: clean up logical flags
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:28:31 -04:00
Alyssa Rosenzweig 44a5481254 InstructionCountCI: add more cases for mov 0/~0
some obvious opportunities to improve here! although mostly I want this to
regression test my constprop rework.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:26:12 -04:00
Alyssa Rosenzweig cf2ff90f87 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:14:06 -04:00
Alyssa Rosenzweig b8dd5d95b0 OpcodeDispatcher: optimize X87FTWTag
using bit twiddling tricks :-)

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-28 13:13:16 -04:00
Ryan Houdek 7cd52febc2 FEXCore/CPUID: Remove warning
This is currently only used on win32 builds because Linux doesn't
understand TPIDRRO.
2025-05-28 09:28:41 -07:00
Ryan Houdek 8e079c1965 Merge pull request #4590 from Sonicadvance1/remove_stack_set_arm
TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
2025-05-27 09:21:14 -07:00
Ryan Houdek 4928af5a64 Merge pull request #4585 from alyssarosenzweig/opt/pair-push-pop
Pair push/pop
2025-05-27 09:17:47 -07:00
Alyssa Rosenzweig 289df740cd InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:13 -04:00
Alyssa Rosenzweig 6ad7392cd7 RegisterAllocationPass: pair push/pop
as a simple post-RA peephole. much much easier to do post-RA than pre-RA.

This isn't a post-RA /pass/ in the traditional sense... it's done while
assigning registers to coalesce the passes over the IR, since we pay per-pass
and we can merge the walks over the IR.

Closes: #4480
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:33:12 -04:00
Alyssa Rosenzweig b15d5f299c RegisterAllocationPass: drop dead IP increment
written but never read.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-27 11:32:33 -04:00
Ryan Houdek 5c7c959dd9 TestHarnessRunner: Stop setting the guest RSP for ARM test runner.
This matches behaviour with the x86 host runner that RSP isn't
guaranteed to be set to a valid memory location. Make sure when running
tests on ARM that it gets the same behaviour.
2025-05-26 13:02:07 -07:00
Alyssa Rosenzweig 6927c7577a RegisterAllocationPass: delete trivial instructions
lots of instructions only exist for RA, so RA can garbage collect them before
post-RA passes (including the JIT) deals with them. this simplifies our life
now, and makes post-RA passes a LOT simpler for little cost.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 88682a457a IR: add RemovePostRA helper
Remove blows up because of use tracking, but we can do a much simpler version
for post-RA and elide lots of checks from trying to make Remove more general.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 61ae53cc03 IR: add paired PushTwo/PopTwo helpers
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 0e67f30103 OpcodeDispatcher: make Push do the right thing and use it
this both optimizes and bug-fixes pusha while deleting a snotton of code.

Closes: https://github.com/FEX-Emu/FEX/issues/4589
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig babd6e9a7b RegisterAllocationPass: allow Copy on the input IR
useful for pusha.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Alyssa Rosenzweig 76caa2c6e3 unittests: add pop-to-same-reg test
an earlier version of this PR failed this.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-26 15:52:23 -04:00
Ryan Houdek dc9f8aa855 Merge pull request #4580 from alyssarosenzweig/ir/inline-ra
IR: Inline registers into the IR
2025-05-26 09:44:51 -07:00
Ryan Houdek ded8b3284a Merge pull request #4584 from neobrain/feature_assert_source_location
LogManager: Print source location when failing assertions
2025-05-24 01:15:11 -07:00
Tony Wasserka 2e24ee7a5f LogManager: Print source location when failing assertions 2025-05-24 09:35:11 +02:00
Tony Wasserka 048ae597d2 CodeEmitter: Fix incorrect macro parameter passing
Macros aren't aware of C++ templates, so the comma is considered a macro
argument separator unless the argument is wrapped in parentheses.
2025-05-24 09:35:11 +02:00
Ryan Houdek 0147f7aa19 Merge pull request #4587 from stanfordzhang/main
Fix callee saved floating-point arguments order issue
2025-05-23 20:13:58 -07:00
StanfordZhang 23cda2c961 Update Arm64Emitter.cpp
fix callee saved floating-point arguments issue
2025-05-23 22:14:37 +08:00
Alyssa Rosenzweig feb67658e1 RegisterAllocationPass: exploit new IR
Now that we can just set registers directly, we can simplify RA a lot. All the
Map/Unmap nonsense - it all goes away. We just assign registers as we go and
everything clicks into place naturally.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 65ee1fafa8 IR: drop RegisterAllocationData
This sideband is now unused, registers are encoded directly in the IR. So we can
garbage collect all this code for quite some savings.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 49f8332c5b JIT: use registers directly from the IR
This is the flag day change from the series, using all the new shiny
infrastructre we added.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig afce108ed7 JIT: make almost all the DEF_OPs common
this deduplicates a bunch of #defines, letting us change the signature easier.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 30f3b545af RegisterAllocationPass: use Header spill slots instead
Removes even more RAData dependence.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig bf1597920c IR: use post-RA flag
rather than implicitly depending on the RA data.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig f66bf3811e RegisterAllocationPass: set PostRA flag
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig a01e29ac99 IR: extend the IR header with RA info
Beyond the actual registers allocated, there are two pieces of sideband data we
store in the RAData object:

* # of spill slots (explicitly)
* whether RA has run (implicitly by the existence of RAData)

We want to get rid of RAData, so we'll move these to the header.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 3136a5e2f8 IR: add helpers to extract physical registers from IR
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig c4f00a05df IR: add space for registers right in OrderedNode *
This again follows the same idea of eliminating the RAData sideband.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig e9a8f8a9ff IR: generate builders that take OrderedNodeWrapper
when we need to materialize instructions inside RA without having a
corresponding OrderedNode* source, we want to just pass a OrderedNodeWrapper
with an encoded register. generate appropriate builders so this is possible.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 261ae7a195 IR: add immediates into OrderedNodeWrapper
this will let us encode registers directly inside OrderedNodeWrapper, rather
than pointers to OrderedNode *. that will let us speed up RA & onwards.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-19 13:42:59 -04:00
Alyssa Rosenzweig 7eaf5ae9e0 IR: drop FillRegister original source
this is now unused, and it's problematic with future work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig 10a02449b1 IR: drop RA validation
There's no reasonable way to keep this around without adding significant
complexity to RA. This series prefers to drop complexity from RA, lessening the
need for validation in the first place.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-16 15:18:06 -04:00
Alyssa Rosenzweig debc57e8c7 Merge pull request #4575 from alyssarosenzweig/cleanup/no-implicit-size
IR: drop DestSize inference
2025-05-16 15:17:53 -04:00
Tony Wasserka 7a4fff8e5b Merge pull request #4578 from cjacek/libgcc
Use -print-libgcc-file-name to get the compiler-rt file name
2025-05-16 13:40:25 +02:00
Jacek Caban ec0a2a8671 Use -print-libgcc-file-name to get the compiler-rt file name
Avoid hardcoding the file name. Upstream Clang currently expects the aarch64 variant of
compiler-rt. Changing this to use arm64ec in the file name could be problematic in the future,
if the MinGW toolchain gains support for ARM64X. Querying Clang for the correct file name is
the most forward-compatible approach.
2025-05-16 11:24:57 +02:00
LC b8b516f7b6 Merge pull request #4563 from Sonicadvance1/minor_vma_cleanup2
SyscallsVMATracking: Minor cleanup
2025-05-15 19:15:20 -04:00
Alyssa Rosenzweig 3b1e91b1fc IR: drop DestSize inference
no more users, and it's problematic for upcoming work.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 15:00:03 -04:00
Alyssa Rosenzweig 8532593d91 IR: specify DestSize for RMWHandle
seems to just have been an oversight.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:58:05 -04:00
Alyssa Rosenzweig 55bd16e2c8 IR: make FillRegister sizes explicit
instead of hacking around it.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:57:51 -04:00
Alyssa Rosenzweig c7797d56c9 IR: don't use GetOpSize in ExitFunction
nothing else does this, and it complicates upcoming refactor to move away from
IR builder helpers doing IR dereferencing.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-15 14:51:11 -04:00
Ryan Houdek 9b573effd1 Merge pull request #4573 from alyssarosenzweig/cleanup/jit-id-2
JIT: use .ID() even less
2025-05-14 13:36:33 -07:00
Alyssa Rosenzweig dfd1aedae5 JIT: use .ID() even less
oops, missed a "/g" with th sed.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 16:23:22 -04:00
Ryan Houdek 221ae2d7b4 Merge pull request #4571 from alyssarosenzweig/opt/constprop-xor-1
ConstProp: optimize XOR with all-1
2025-05-14 12:59:12 -07:00
Ryan Houdek 9127d206b5 Merge pull request #4572 from alyssarosenzweig/cleanup/jit-id
JIT: stop using .ID() pattern
2025-05-14 12:59:01 -07:00
Alyssa Rosenzweig ecc6fea54e JIT: stop using .ID() pattern
sed -ie 's/.ID()//' *

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:46:32 -04:00
Alyssa Rosenzweig 99446da7c1 JIT: add Reg helpers taking OrderedNodeWrappers
more ergonomic and will give us freedom to migrate things easier soon.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:43:20 -04:00
Alyssa Rosenzweig ad75563a26 IR: drop irrelevant reference to x86
we don't run the JIT on x86 anymore.

Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 15:32:07 -04:00
Alyssa Rosenzweig 197af972d9 InstCountCI: Update
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Alyssa Rosenzweig fbd706c191 ConstProp: optimize XOR with all-1
Signed-off-by: Alyssa Rosenzweig <alyssa@rosenzweig.io>
2025-05-14 14:48:29 -04:00
Ryan Houdek 6b85fa5611 SignalScopeGuards: Review 2025-05-14 10:56:44 -07:00
Ryan Houdek fae66a921f SyscallsVMATracking: Moves MappedResources implementation details to private 2025-05-14 10:56:44 -07:00
Ryan Houdek e6fc462e9d SyscallsVMATracking: Rename VMA tracking functions
Makes it more easily reflect what this functions are actually doing and
add a couple lines of documentation to help with the inherent opaqueness
of these.

NFC, just helps my brain wrap this more easily.
2025-05-14 10:56:44 -07:00
Ryan Houdek d11a265a53 SyscallsVMATracking: Adds validation that thread has ownership of VMA mutex
This way we can capture any programming bugs. Can only check the
functions that actually require unique locks rather than shared locks.
2025-05-14 10:56:44 -07:00
Ryan Houdek 395d870814 SignalScopeGuards: Add checks for locks being held by the calling thread
pthreads allows us to check if mutex/rwlock is currently locked by the
calling thread. This can give us some safety in code expecting locks to
be in place, allowing us to find programming bugs.
2025-05-14 10:56:44 -07:00
Ryan Houdek fec1ffaa6b Merge pull request #4566 from neobrain/feature_pool_alloc_size
ThreadPoolAllocator: Add support for updating the size of managed data
2025-05-14 10:53:34 -07:00
Tony Wasserka d4fdb28e72 PoolBufferWithTimedRetirement: Add support for updating the buffer size
This should be done at low frequency since it may unclaim the buffer.
2025-05-14 13:37:46 +02:00
Tony Wasserka 39a5c2021e ThreadPoolAllocator: Rename FixedSizePoolAllocation to PoolBufferWithTimedRetirement
This more accurately reflects that the core feature of the helper is the
timer-based unclaiming of buffers instead of the allocation size.
2025-05-14 13:29:44 +02:00
Tony Wasserka f41501444d ThreadPoolAllocator: Add a dedicated interface to try reowning a buffer without fallback 2025-05-14 13:29:44 +02:00
Tony Wasserka 86b26b80ce FixedSizePooledAllocation: Clean up documentation 2025-05-14 13:29:44 +02:00
Ryan Houdek ae9a5b1125 Merge pull request #4487 from pmatos/GCCTargetTestsN1
Run gcc target tests with block size 1
2025-05-13 04:45:52 -07:00
LC 38579807c2 Merge pull request #4567 from neobrain/fix_glibcxx_debug
ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
2025-05-09 15:17:48 -04:00
Tony Wasserka c19119bcd6 ThreadPoolAllocator: Fix assertion in _GLIBCXX_DEBUG builds
Default-constructed iterators can't be copied.
2025-05-09 15:02:49 +02:00
LC 47f1ad693d Merge pull request #4561 from Sonicadvance1/minor_vma_cleanup
LinuxEmulation: Minor cleanup by separating VMA definitions
2025-05-08 15:21:44 -04:00
Ryan Houdek 2164d7bb96 Merge pull request #4565 from neobrain/fix_async_timeout
FEXServer: Don't time out while clients are still connected
2025-05-08 12:14:56 -07:00
Tony Wasserka c326e2d669 FEXServer: Don't time out while clients are still connected 2025-05-08 11:56:19 +02:00
Tony Wasserka 8eaf45414c Async: Add run_one interface to enable more fine-grained event loop control 2025-05-08 11:56:19 +02:00
Ryan Houdek b3297d106e Merge pull request #4564 from bylaws/arm64ec
CMake: Allow disabling explicit -mcpu usage
2025-05-07 15:30:04 -07:00
Ryan Houdek 2d0e19e6a7 Merge pull request #4562 from WhatAmISupposedToPutHere/main
Windows: Fix building with llvm-libcxx
2025-05-07 15:06:22 -07:00
Billy Laws fd2ee4dc46 CMake: Allow disabling explicit -mcpu usage 2025-05-07 22:48:05 +01:00
Sasha Finkelstein b1bbc37c59 Windows: Fix building with llvm-libcxx
Libcxx uses GetSystemTimePreciseAsFileTime if _WIN32_WINNT specifies
a version new enough to have it.
2025-05-07 23:47:24 +02:00
Ryan Houdek 56409d4f2b LinuxEmulation: Minor cleanup by separating VMA definitions
NFC, just moving this to its own header. It's already a huge PITA to
read. I just want to try and preserve some sanity while attempting to
fix #4557
2025-05-07 13:55:44 -07:00
Paulo Matos ec0683b729 Run gcc target tests with block size 1
info files for the test runner like Disabled_Tests, Known_Failures,
etc, receive not just the filename but the test name (which is the test name
and potentially its running config).

Increase the timeout of gcc tests to 30secs.
2025-05-07 16:14:53 +02:00
LC 89e5041e70 Merge pull request #4559 from Sonicadvance1/we_require_more_machicolations!
LinuxEmulation: Implement custom longjump that is fortification safe
2025-05-07 09:07:12 -04:00
LC 8b3e7312e0 Merge pull request #4560 from Sonicadvance1/fix_bad_define_check
LinuxEmulation: Fix bad compile time definition check
2025-05-06 22:28:41 -04:00
Ryan Houdek 3c76a9176d LinuxEmulation: Fix bad compile time definition check
We don't want this to be compiled out if the definition doesn't exist.
Actually define it in that case.
2025-05-06 18:39:58 -07:00
Ryan Houdek a37def2c22 LinuxEmulation: Implement custom longjump that is fortification safe
With fortifications enabled, glibc long jump has some additional checks
in place that break because we do a stack pivot. The only way around
this is to do our own long jumps. Luckily this is trivial.

Fixes #4558
2025-05-06 15:29:42 -07:00
Ryan Houdek ee4ef0fe87 Docs: Update for release FEX-2505 2025-05-05 11:37:17 -07:00
LC 1dff7073de Merge pull request #4554 from Sonicadvance1/remove_unused_argument
NFC: FEXCore: Removes unused argument on CreateThread
2025-05-05 14:30:24 -04:00
Ryan Houdek 92a82c3134 FEXCore: Removes unused argument on CreateThread
ParentTID is purely a Linux construct and has been moved entirely to the
frontend at this point. Remove this argument which is now unused.
2025-05-05 11:17:02 -07:00
Ryan Houdek a6a203c483 Merge pull request #4553 from neobrain/fix_align16b 2025-05-05 10:35:55 -07:00
Tony Wasserka cdaa65f6fd Arm64Emitter: Fix overalignment in Align16B
Previously, 16 additional bytes were emitted if the buffer was already
aligned.
2025-05-05 16:02:42 +02:00
Ryan Houdek 31e5f706b8 Merge pull request #4551 from OFFTKP/fadvise64
Fix 32-bit fadvise64
2025-05-02 18:37:55 -07:00
offtkp 00e05558ce Formatting 2025-05-03 04:28:23 +03:00
offtkp c02b88baff Fix 32-bit fadvise64 2025-05-03 03:56:16 +03:00
Ryan Houdek 794c80edad Merge pull request #4549 from OFFTKP/patch-2
Marshal freeram in sysinfo
2025-05-02 17:34:52 -07:00
Paris Oplopoios fb939600e2 Marshal freeram in sysinfo 2025-05-03 03:23:18 +03:00
Ryan Houdek 9dbbd44d09 Merge pull request #4544 from Sonicadvance1/fhu_ring_buffer
FHU: Add a non-block thread local ringbuffer
2025-05-02 11:22:39 -07:00
Ryan Houdek afcd93fe5f Merge pull request #4547 from Sonicadvance1/dont_go_chasing_sigsegv_waterfalls
Linux/SMCTracking: Stop calling mprotect on a memory region times the number of threads
2025-05-02 11:22:30 -07:00
Ryan Houdek 118faa5380 FHU: Add a non-block thread local ringbuffer
I keep rewriting this thing when I want to see some history on
something. Throw it in a FHU utility so I can stop wasting my time.
2025-05-02 01:19:00 -07:00
Ryan Houdek 7ed17f68c5 FEXCore: Remove unused InvalidateGuestCodeRange with callback 2025-05-02 01:15:50 -07:00
Ryan Houdek 96363143de Linux/SMCTracking: Stop calling mprotect on a memory region times the number of threads
I noticed this cascade of mprotects when poking at Crypt of the
Necrodancer, since it consistently is invalidating code. I saw us
calling mprotect on the same page 32 times in a tight loop and thought
surely this isn't FEX doing this.

Turns out we were calling the callback after invalidating each thread.
It should instead be done once at the end of invalidating the thread's
caches while still holding the locks.

Fixes this weird cascade of mprotects that equal the number of FEX
threads.
2025-05-01 21:14:05 -07:00
Ryan Houdek cf5fcb2371 Merge pull request #4546 from OFFTKP/shmdt
Make shmdt reset the unmapped pages in the 32-bit allocator
2025-05-01 20:17:50 -07:00
offtkp 7d3699050e Make shmdt reset the unmapped pages in the 32-bit allocator 2025-05-02 02:49:39 +03:00
LC 40e0a71f49 Merge pull request #4545 from Sonicadvance1/fix_fexserver_with_images
FEXServer: Fixes squashfs/erofs with newer fuse releases
2025-05-01 17:55:57 -04:00
Ryan Houdek cdbcdfe397 FEXServer: Fixes squashfs/erofs with newer fuse releases
Newer fuse releases changed how they are waiting on child processes to
exit. Setting the signal action to SIG_IGN would cause
erofsfuse/squashfuse to inherit the ignored action and cause their
internal `wait4` syscalls to fail with ECHLD.

Set the action back to default inside the FEXServer because our original
reasoning for setting the ignoring is no longer valid. FEXInterpreter
still ignores SIGCHLD while launching FEXServer.

Maybe fixes the muvm thing people have been complaining about.

Also fixes accidental comma delimiter usage.
2025-05-01 13:47:46 -07:00
LC 4f2d2e646e Merge pull request #4542 from Sonicadvance1/fexcore_reconstructions_getting_saved_today
FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
2025-05-01 13:28:16 -04:00
LC f97cd24647 Merge pull request #4543 from Sonicadvance1/fix_my_reducing_of_x87_today
FEXCore: Fixes x87 reduced precision
2025-05-01 13:25:30 -04:00
Ryan Houdek e5e75ad1ef InstcountCI: Update 2025-04-30 17:54:55 -07:00
Ryan Houdek fc052efb91 FEXCore: Fixes x87 reduced precision
With the change from #4538 I had accidentally broken x87 reduced
precision.

This is due to the fact that we accidentally lost ABI information about
interpreter fallbacks supporting `preserve_all` or not. So now instead
of having some ABI callbacks supporting it and some not, just force
usage of `preserve_all` if it is supported by the compiler entirely.

Fixes Steam when x87 reduced precision is enabled.
2025-04-30 17:37:10 -07:00
Ryan Houdek c5754145c5 FEXCore/JIT: Switch over to new vl64pair for RIP reconstruction
As said in the implementation of this struct commit message. This new
pair struct optimizes specific cases of small forward only increments
that can fit in to 8-bit space, and small forward or backward jump cases
that fit in to 16-bit space.

Some stats of this change:
- Steam: 5.88MB down to 4.34MB. 73.8% the space consumed
- Steamwebhelper: 15.8MB down to 13.24MB. 83.8% space consumed
- Sonic Mania: 3.6MB down to 2.58MB. 71.6% space consumed

As for absolute stats when compared to all code buffer size:
- Steam: 86MB of code buffer to 5.88MB -> 4.34MB of RIP reconstruction.
  - 6.8% -> 5% code buffer space used for RIP reconstruction
- Steamwebhelper: 285MB of code buffer to 17MB -> 14.26MB of RIP reconstruction.
  - 5.9% -> 4.9% code buffer space used for RIP reconstruction
- Sonic Mania: 48.53MB of code buffer to 3.55MB -> 2.53MB of RIP reconstruction.
  - 7.3% -> 5.2% code buffer space used for RIP reconstruction
2025-04-30 15:41:43 -07:00
Ryan Houdek 9986e69622 FEXCore/APITests: Extend tests for vl64pair 2025-04-30 15:41:20 -07:00
Ryan Houdek 00acf4d327 FEXCore/Utils: Implement a new VL64Pair struct type
This new variable length pair of integer implementation is taking direct
advantage of the most common aspects of FEX's JIT in that most x86
instructions are <= 8-bytes in length, and the ARM implementations of
those are /usually/ 16 instructions in length or less. Also only
unsigned offsets in this implementation since the common case is forward
incrementing.

This converts a majority of 16-bit vl encodings in to an 8-bit encoding
instead, shaving space off the RIP reconstruction data.

An additional optimization is for the 16-bit pair of integers, we
continue this optimization through but with more bits and changing over
to signed. This captures the second most common cases of /slightly/
larger increments and small loops.

Pairs of 32-bit and 64-bit integers are unoptimized since they are
uncommon.
2025-04-30 15:36:49 -07:00
Ryan Houdek 00ff549044 Merge pull request #4538 from Sonicadvance1/fex_interpreters_are_dancing
JIT: Move interpreter ABI handlers in to the dispatcher
2025-04-30 11:51:29 -07:00
Ryan Houdek 75a39c8939 InstcountCI: Update 2025-04-29 22:35:39 -07:00
Ryan Houdek 5b9354cc74 JIT: Move interpreter ABI handlers in to the dispatcher
Fixes #4535

A handful of improvements on this.
* Reduces codegen around interpreter fallbacks
* Keeps ABI handling code in common Dispatcher code
* Improves I$ hitrate by most of the heavy code staying in Dispatcher

This cuts the amount of codegen inside the JIT for most interpreter
fallbacks by 1/2 or 1/3, by only doing the minimal amount of work in the
code blocks and doing most things in the dispatcher. The cost of which
is an additional branch per operation.

This should bring marginal performance improvements, but it should also
basically fall within noise. The bigger thing to care about here is a
smaller amount of code being generated for x87 blocks.
2025-04-29 22:35:39 -07:00
Ryan Houdek f7049a6478 Arm64Emitter: On spill return stack used and stop clobbering TMP4
TMP4 was used before we passed in a tmp register. Now use that temp
register.

Also return the amount of stack used on the push function. This will be
used in a bit.
2025-04-29 22:31:35 -07:00
Ryan Houdek e1d032b5a6 Merge pull request #4540 from pmatos/MProtectLastPage
mprotect last page of CodeBuffer
2025-04-27 11:07:14 -07:00
LC f99691b0eb Merge pull request #4541 from Sonicadvance1/f80_cephes_softfloat_prep
80-bit cephes prep work
2025-04-26 09:18:57 -04:00
Ryan Houdek 10d18f2f17 Softfloat-3e: Adds missing extF80_le file 2025-04-25 17:30:18 -07:00
Ryan Houdek a236700c55 cephes_128bit: Rename a bunch of variables
These are going to conflict with the 80-bit implementation otherwise.
2025-04-25 17:30:18 -07:00
Ryan Houdek 34d62fcea6 cephes: Split 128-bit to its own folder
We are soon going to learn how to operate in f80.
2025-04-25 17:30:18 -07:00
Paulo Matos 791502afef Protect last page of CodeBuffer
protects last page of codebuffer. This should cause a
SIGSEGV if we try to access it. Until now it was possible to go over
and access out of bounds.

In addition, there a couple of clang-tidy fixes which should be NFC.
2025-04-25 20:41:18 +02:00
Ryan Houdek 99114e1fc2 Merge pull request #4528 from neobrain/feature_unlink_threadsafe
JIT: Make code patching during (un-/)linking thread-safe
2025-04-25 08:22:38 -07:00
Ryan Houdek 3c8bb53d1a Merge pull request #4539 from cjacek/mincore
Link to mincore instead of kernelbase on ARM64EC
2025-04-25 08:22:19 -07:00
Jacek Caban bd0682b518 Link to mincore instead of kernelbase on ARM64EC
The kernelbase import library is not available in upstream mingw-w64. Instead, similar to MSVC,
we can use mincore, which allows importing the relevant functions via apisets.
2025-04-25 11:38:16 +02:00
Tony Wasserka e89a913b69 JIT: Make memory write visible to other threads reading the same location 2025-04-25 09:47:13 +02:00
Tony Wasserka 293d77d412 JIT: Make code patching during (un-/)linking thread-safe 2025-04-25 09:21:09 +02:00
Ryan Houdek 6beb4b0f8b Merge pull request #4533 from Sonicadvance1/align_tail_in_the_pale_moonlight
JIT: Align JITCodeTail to native alignment
2025-04-24 16:18:45 -07:00
LC 23a8462688 Merge pull request #4536 from Sonicadvance1/eternal_sunshine_of_a_vectorless_mind
JIT: Moves VPCMPESTRX handler to use vectors
2025-04-24 13:35:33 -04:00
LC 02f90fbc66 Merge pull request #4534 from Sonicadvance1/spaaaaaaace
CodeEmitter: Fix clang-format
2025-04-24 13:34:19 -04:00
Ryan Houdek 8cc23d0cf8 Merge pull request #4537 from sdpoueme/main
updated Dockerfile to reflect latest FEX-Emu releases
2025-04-24 10:28:24 -07:00
Serge Poueme 60b5b1414d updated Dockerfile to reflect latest FEX-Emu releases
Add multi-stage Dockerfile for FEX emulator build

- Stage 1 (Builder): Sets up build environment with Ubuntu 22.04
  - Installs development tools and dependencies
  - Builds FEX using clang-13 with optimized settings
  - Configures cmake with LTO enabled and tests disabled

- Stage 2 (Runner): Creates minimal runtime image
  - Includes only necessary runtime libraries
  - Copies built binaries from builder stage
2025-04-24 10:13:04 -07:00
Ryan Houdek cbab344822 InstcountCI: Update 2025-04-24 10:11:25 -07:00
Ryan Houdek 32764ddf81 JIT: Moves VPCMPESTRX handler to use vectors
I pushed this off from the previous changes that were converting things
to vector as less important. It has now become more important to keep
these in vector registers until beyond the ABI boundary.

This will reduce burden on our JIT backend and just changes where the
movement in to GPRs occurs. Necessary for #4535
2025-04-24 10:09:26 -07:00
Ryan Houdek e31127e39e CodeEmitter: Fix clang-format 2025-04-24 08:41:23 -07:00
Ryan Houdek 735f537846 JIT: Align JITCodeTail to native alignment
Removes UB
2025-04-24 08:40:43 -07:00
Ryan Houdek b7790e10e9 Merge pull request #4532 from pmatos/X87StateBlockReset
X87 state block reset
2025-04-24 08:37:32 -07:00
Ryan Houdek 3f7ad04054 Merge pull request #4531 from Sonicadvance1/instcountci_change_size
InstcountCI: Changes how code size is calculated
2025-04-24 08:35:33 -07:00
Paulo Matos 670ddf9cd8 instcountci: Reset MMXState to X87 at the start of each block 2025-04-24 15:02:46 +02:00
Paulo Matos d377e26106 Reset MMXState to X87 at the start of each block
Ensures that blocks always start with the same state independently of predecessors
which allows independent compilation of blocks.
Starting in the X87 state is better than starting in MMX state because
MMX state is more work to initialize.
2025-04-24 15:02:41 +02:00
Ryan Houdek 674f939690 Merge pull request #4523 from bylaws/maptrack
Improve tracking of executable mappings
2025-04-23 13:08:12 -07:00
Ryan Houdek 64abfb1afe InstcountCI: Update 2025-04-23 12:56:46 -07:00
Ryan Houdek adcefe62cf InstcountCI: Changes how code size is calculated
Due to how jit block tail padding is working, there's no real good way
to determine the true "implementation size" of an instruction without
the backend being aware of wanting to investigate it.

Trying to inject another instruction, or another IR operation actually
subtly changes codegen in a way that gives invalid results. The only
real way to get around this is to inject a known token in to the
instruction stream as we `ExitFunction`.

So inject a `udf #0x420f`, and change the scanning behaviour to find the
first one and cut everything else off afterwards.

This already scoops out some code in some game blocks that were
accidentally landing ExitFunction code in the json.

This also has been tested to work with #4528 with its InstCountCI
specific changes reverted.

This means we don't need to play subtle padding tricks in the JIT to get
the information we want in InstcountCI.
2025-04-23 12:56:46 -07:00
Ryan Houdek bf77fa4275 unittests/Emitter: Adds udf test 2025-04-23 12:53:36 -07:00
Ryan Houdek e7fce2de40 CodeEmitter: Adds udf support 2025-04-23 12:14:44 -07:00
Ryan Houdek 8864637602 Merge pull request #4465 from Sonicadvance1/ubsan_fixes
Various: UBSAN fixes around unaligned accesses
2025-04-23 00:54:49 -07:00
Ryan Houdek 1a0d97d5f7 Various: UBSAN fixes around unaligned accesses 2025-04-23 00:31:09 -07:00
Ryan Houdek 9ba46e395d Merge pull request #4530 from neobrain/fix_instcountci_script
Scripts/InstCountCI: Restrict git-add to relevant JSON files only
2025-04-22 06:10:43 -07:00
Tony Wasserka be99fa0b05 Scripts/InstCountCI: Restrict git-add to relevant JSON files only 2025-04-22 13:34:28 +02:00
Ryan Houdek c42b1858c5 Merge pull request #4529 from lioncash/tests
ASIMD_Tests: Enable PMULL/PMULL2 tests
2025-04-21 19:03:01 -07:00
Lioncache cc11de16bf ASIMD_Tests: Enable PMULL/PMULL2 tests
vixl now supports these. We can also get rid of the invalid data sizes,
since only halfword and 128-bit variants are defined.
2025-04-21 20:03:37 -04:00
Ryan Houdek 3ea7c2b014 Merge pull request #4522 from bylaws/persona
LinuxEmulation: Copy the host persona on thread startup
2025-04-21 00:59:01 -07:00
Ryan Houdek e0e9f5aa16 Merge pull request #4527 from lioncash/concept
CodeEmitter/ASIMDOps: Constrain Q and D register requirements with concept
2025-04-20 18:14:15 -07:00
Lioncache 63a2d1e1db CodeEmitter/ASIMDOps: Move three emitter helpers to private section
These don't need to be public.
2025-04-20 15:55:53 -04:00
Lioncache 644263a764 CodeEmitter/ASIMDOps: Constrain Q and D register requirements with concept
Pulls these out into a single concept instead of having the same lengthy
requirements clause.
2025-04-20 15:55:49 -04:00
Ryan Houdek f912295690 Merge pull request #4524 from bylaws/faulto
X86Tables: Set FLAGS_BLOCK_END for more faulting ops
2025-04-19 19:02:09 -07:00
Ryan Houdek 79e57c320c Merge pull request #4525 from lioncash/pac
CodeEmitter/LoadstoreOps: Add Load/store register (PAC) group
2025-04-19 18:29:18 -07:00
Lioncache 4e6f183f77 CodeEmitter/LoadstoreOps: Add Load/store register (PAC) group
Eh, what the heck. Gets rid of the last TODO marker in the base load-stores.
2025-04-19 13:00:00 -04:00
Billy Laws 23017af874 TestHarnessRunner: Call the frontend mapping callback for created mappings 2025-04-19 15:29:23 +01:00
LC 4d02126b77 Merge pull request #4517 from Sonicadvance1/ruining_another_set_of_memory_leaks
LinuxEmulation: Fixes remaining memory leak on pthread teardown
2025-04-19 09:44:26 -04:00
Billy Laws 30fefbd57d X86Tables: Set FLAGS_BLOCK_END for more faulting ops 2025-04-19 14:19:19 +01:00
Billy Laws 38304fa63a LinuxEmulation: Copy the host persona on thread startup 2025-04-19 14:10:55 +01:00
Billy Laws 6525b8dd14 VDSO_Emulation: Map the VDSO thunk as executable 2025-04-19 14:09:48 +01:00
Billy Laws 0155719b1a VDSO_Emulation: Ensure the 32 bit sigreturn mapping is tracked as executable 2025-04-19 14:09:48 +01:00
Billy Laws 966dce6e95 SyscallsSMCTracking: Handle shmat SHM_EXEC flag 2025-04-19 14:09:48 +01:00
Billy Laws e3088686db InvalidationTracker: Track executable mappings 2025-04-19 14:09:48 +01:00
Billy Laws 771162cf14 InvalidationTracker: Track the protection of regions mapped at startup
Code can be injected into a process by e.g. the chromium sandbox before
FEX is loaded.
2025-04-19 14:09:48 +01:00
Billy Laws 1a2398d2ce Windows: Call the protection callback for executable FEX mappings 2025-04-19 14:09:48 +01:00
Billy Laws 0f0e969d8f Windows: Call the image map callback for ntdll 2025-04-19 14:09:48 +01:00
Ryan Houdek c8371087e6 Merge pull request #4519 from neobrain/fix_fexconfig_layout
FEXConfig: Fix layout issues on Qt 6.9
2025-04-18 09:25:08 -07:00
Ryan Houdek 3da6fc3972 Merge pull request #4518 from neobrain/refactor_warn_fixes
JIT: Fix warning about unused variable
2025-04-18 09:24:49 -07:00
Ryan Houdek 0cbbd91e72 Merge pull request #4521 from lioncash/mem
CodeEmitter/LoadStoreOps: Add Memory Copy and Memory Set category
2025-04-18 09:24:24 -07:00
Lioncache d71aca9e9b CodeEmitter/LoadStoreOps: Add Memory Copy and Memory Set category
Adds all of the memory facilities in FEAT_MOPS to the emitter.
2025-04-18 10:53:38 -04:00
Tony Wasserka 63035fd5f3 FEXConfig: Fix layout issues on Qt 6.9 2025-04-18 11:17:51 +02:00
Tony Wasserka 0fe28129a0 JIT: Fix warning about unused variable 2025-04-18 11:16:46 +02:00
Ryan Houdek c44d8eed35 LinuxEmulation: Fixes remaining memory leak on pthread teardown
With the previous stack leak fix, RUINER reduced its memory leaking down
to around 50MB/s. The remaining memory leaks come from the pthread stack
that we are required to allocate (128KB per thread) and some internal
DTV tracking structures.

The problems come in the fact that glibc/pthread only tears down its
internal state for these if the pthread function actually returns! We
can **technically** switch the initial stack over to a "user" stack but
that introduces more problems around internal dtv tracking that we
already fixed months ago, so we can't actually do that in practice.

This leaves us no choice, we effectively are mandated to return from the
pthread function in order to free the memory from glibc. The only way we
can safely do this is with a long jump and deferring some data structure
management until that case.

This is all incredibly sucky but it's necessary to work. With these
changes, RUINER is no longer leaking memory (Hovering at around 3GB used
while in-game) and even Steam is consuming less memory.

It doesn't solve the problem that thread creation and teardown could
likely be faster, but not many applications are creating 720
threads/second.
2025-04-17 18:16:32 -07:00
Ryan Houdek 6f5588f71e Merge pull request #4504 from pmatos/NoStrictAliasing
Enable -fno-strict-aliasing
2025-04-17 17:27:33 -07:00
Ryan Houdek 68f8c244b4 Merge pull request #4516 from lioncash/radd
CodeEmitter/ASIMDOps: Support RADDHN{2}/RSUBHN{2}
2025-04-17 17:27:22 -07:00
Lioncache d7a08fc7d4 CodeEmitter/ASIMDOps: Support RADDHN{2}/RSUBHN{2}
These are trivial enough to just drop right in.
2025-04-17 20:13:30 -04:00
LC 572d6e0395 Merge pull request #4509 from Sonicadvance1/in_the_twilight_of_the_pale_blue_moon
A couple of barrier and timing fixes.
2025-04-17 17:30:03 -04:00
LC 156b6745a2 Merge pull request #4513 from Sonicadvance1/fix_stack_leak
LinuxSyscalls: Fixes a major stack memory leak
2025-04-17 17:27:15 -04:00
Ryan Houdek b4eb38e2b7 Merge pull request #4515 from lioncash/scalar
CodeEmitter/ScalarOps: Add two more instruction categories
2025-04-17 14:10:04 -07:00
Lioncache bafcc8762b CodeEmitter/ScalarOps: Handle ASIMD scalar x indexed element group 2025-04-17 14:47:38 -04:00
Lioncache ec6ff7899d CodeEmitter/ScalarOps: Remove implemented TODO
These scalar ops are already implemented and this was accidentally left in.
2025-04-17 08:14:10 -04:00
Lioncache 6330b398a5 CodeEmitter/ScalarOps: Handle ASIMD scalar three same extra group
These are trivial enough to drop in.
2025-04-17 08:09:42 -04:00
Ryan Houdek 86c492da13 LinuxSyscalls: Fixes a major stack memory leak
Fixes an issue where a thread that exits with the `exit` syscall never
actually frees its pivot stack. This is common practice and it was
missed when I was fixing the previous stack pivot leak.

This was uncovered when looking at the game
[RUINER](https://store.steampowered.com/agecheck/app/464060/?curator_clanid=4777282)
for timing bugs. Turns out the Linux build of the game creates and
destroys a VLC object every tick of the engine. Creating this VLC
object creates six threads behind the scenes. At what I assume the
default tick of the engine is of 120Hz(?) this would mean it is
attempting to create and destroy 720 threads per second, quickly
leading to memory exhaustion under FEX.

While this hits the biggest memory leak we have around thread creation,
this game is still hitting thread creation so hard that I can see other
leaks that I need to track down still.
2025-04-16 19:47:56 -07:00
Ryan Houdek 8ff7497d59 Merge pull request #4501 from JunChi1022/mask_signal_at_defer
LinuxSyscalls: Update signal mask at deferring time
2025-04-16 10:56:40 -07:00
Ryan Houdek 73593ad4f1 Merge pull request #4499 from bylaws/x87fix
OpcodeDispatcher: Mark NZCV as dirty when clobbering in FCOMIF64
2025-04-16 10:32:13 -07:00
Ryan Houdek 572a58ed27 Merge pull request #4512 from lioncash/internal
Arm64: Mark several functions as internally linked
2025-04-16 10:31:06 -07:00
Ryan Houdek d7df4be379 Merge pull request #4511 from lioncash/test
ASIMD_Tests: Re-enable tests disabled due to dissassembly bugs
2025-04-16 10:30:25 -07:00
Ryan Houdek f227722898 Merge pull request #4503 from pmatos/UBSANflags
Add no-sanitize flags to UBSAN
2025-04-16 10:29:02 -07:00
JustinChi f371bf2bc4 LinuxSyscalls: Update signal mask at deferring time
In regular x86 programs, when a signal occurs, the signal will not be handled within the signal handler. However, under FEX's defer signal mechanism, the signal is not immediately masked when it is deferred. When returning to the location that receives the signal and continues processing, the signal might be received again, causing inconsistency between the emulation and the actual program.

Here is an unit test for this patch from ltp:
https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/syscalls/timer_settime/timer_settime03.c
2025-04-16 21:39:53 +08:00
Lioncache e6841ea46d Arm64_stubs: Remove non-public functions
These aren't exposed in the public interface anymore, so they can be removed.
2025-04-16 09:21:51 -04:00
Lioncache edbfe45a20 Arm64: Mark several functions as internally linked
These aren't directly used outside of the translation unit.
2025-04-16 09:16:44 -04:00
Lioncache 5f83e89be5 ASIMD_Tests: Re-enable tests disabled due to dissassembly bugs
These issues seem to be resolved now.
2025-04-16 08:47:18 -04:00
Billy Laws 5850b26de5 Update InstCountCI 2025-04-16 13:16:07 +01:00
Billy Laws 416267a238 OpcodeDispatcher: Safely clobber NZCV in FCOMIF64
Also fix a small typo that broke the !flagm2 path.
2025-04-16 13:06:37 +01:00
LC 94499ed8fc Merge pull request #4510 from Sonicadvance1/fix_space
CodeEmitter: Fixes misaligned function
2025-04-15 19:58:40 -04:00
Ryan Houdek da1868288d CodeEmitter: Fixes misaligned function 2025-04-15 15:47:31 -07:00
Ryan Houdek c8d6b39585 EmulatedFiles: Match bogomips calculation
Instead of hardcoding the bogomips calculation, more closely match what
the Linux kernel does for bogomips. There are some applications out
there that use bogomips for silly timing calculations, so this is a
better version.
2025-04-15 15:42:52 -07:00
Ryan Houdek 00f8181c3a OpcodeDispatcher: Fix CPUID being a instruction fence
It is common practice for games to use CPUID as an instruction barrier
for various reasons. Ensure that we respect this by adding support for
an instruction barrier.
2025-04-15 15:42:43 -07:00
Ryan Houdek 60565cc2ef OpcodeDispatcher: Implement support for self-synchronizing cycle counter
FEAT_ECV added a new synchronizing cycle counter instruction that
restrict speculation across the cycle counter access. Because it
restricts speculation, it effectively acts like an isb and load dsb.

Luckily for us, this actually matches behaviour for what rdtscp does, so
we can take advantage of it if the host supports FEAT_ECV.
2025-04-15 15:42:43 -07:00
Ryan Houdek bfee39ae70 HostFeatures: Passthrough if the host supports ECV 2025-04-15 15:38:00 -07:00
Ryan Houdek 753c5e72be CodeEmitter: Implement support for CNTVCTSS_EL0 2025-04-15 15:37:02 -07:00
Ryan Houdek 571a533543 Merge pull request #4508 from lioncash/crypto
Emitter: Add missing ASIMD crypto operations
2025-04-15 11:24:51 -07:00
Lioncache 222fcfd70e CodeEmitter/ASIMDOps: Add crypto two-register SHA512 category 2025-04-15 11:48:30 -04:00
Lioncache 95c8a70ce6 CodeEmitter/ASIMDOps: Add crypto four-reg category 2025-04-15 11:41:07 -04:00
Lioncache 0f6bec4d5e CodeEmitter/ASIMDOps: Add crypto three-reg SHA512 category 2025-04-15 11:31:00 -04:00
Lioncache 5fb0522c78 CodeEmitter/ASIMDOps: Add crypto three-reg imm2 category 2025-04-15 11:15:21 -04:00
Tony Wasserka 7c0bc2d972 Merge pull request #4506 from pmatos/CastFixValidateCode
Fix cast in ValidateCode impl
2025-04-15 14:07:51 +02:00
Paulo Matos cb972e165b Fix cast in ValidateCode impl 2025-04-15 09:35:46 +02:00
Paulo Matos fb319a352a Enable -fno-strict-aliasing 2025-04-14 09:29:15 +02:00
Paulo Matos e10646b436 Add no-sanitize flags to UBSAN
See discussion in https://github.com/FEX-Emu/FEX/pull/4494 for context.
2025-04-14 09:10:28 +02:00
Ryan Houdek 211bec65d2 Merge pull request #4492 from bylaws/badencodings
Frontend: Be more tolerant of bad instruction encodings
2025-04-13 19:09:07 -07:00
Ryan Houdek 66807539ce Merge pull request #4500 from bylaws/arm64ec-fixes-etc
Windows: Misc cleanups and fixes
2025-04-13 19:03:17 -07:00
Ryan Houdek 2919c32a8f Merge pull request #4497 from pmatos/CleanupGCCTests
Cleanup info test files for 32bits
2025-04-11 17:43:09 -07:00
Billy Laws 05be9e9c37 Windows: Process cross-process notifications before code compilation 2025-04-11 12:07:16 +01:00
Billy Laws 039aaae041 FEXCore: Add a pre-compilation frontend callback to SyscallHandler 2025-04-11 12:07:16 +01:00
Billy Laws 890653ae39 WOW64: Handle the image mapped callback 2025-04-11 12:07:16 +01:00
Billy Laws 8644fdd595 ARM64EC: Slight cleanup 2025-04-11 12:07:16 +01:00
Billy Laws 4fca8fb8e6 AllocatorHooks: Correctly restore the region protection in VirtualDontNeed 2025-04-11 12:07:16 +01:00
Billy Laws 9625201cbf Frontend: Remove asserts on invalid instruction encodings 2025-04-11 12:06:01 +01:00
Paulo Matos e7dea56adf NFC: Cleanup info test files for 32bits
No point in duplicating test info between Disabled and Known Failures.

Leaving tests known to fail in Known_Failures. Flakes / Unreliable tests go into Disabled_Tests.
2025-04-11 10:50:41 +02:00
LC 5b802d17d1 Merge pull request #4498 from Sonicadvance1/fix_llseek
LinuxSyscalls: Fixes 32-bit llseek result
2025-04-10 16:18:54 -04:00
Ryan Houdek a43ddba87c LinuxSyscalls: Fixes 32-bit llseek result
`llseek` returns only ever 0 or errno in the return register. This is in
contrast to `lseek` which returns the result (or errno) in the return
register.

We were accidentally returning the result on non-error conditions which
could freak out some software. Thanks to
[OFFTKP](https://github.com/OFFTKP) for pointing out this issue
2025-04-10 13:08:06 -07:00
Tony Wasserka f1d007ca4c Merge pull request #4495 from pmatos/CleanupCompWarn
Cleanup compile warnings
2025-04-10 16:40:46 +02:00
Paulo Matos c2111e7384 Cleanup compile warnings 2025-04-10 08:49:51 +02:00
Billy Laws a1c6378317 Frontend: Only accept POP opcodes with a 0 ModRM.reg field 2025-04-09 23:04:11 +01:00
Billy Laws f064013b9a Frontend: Enforce FLAGS_SF_MOD_MEM_ONLY/FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:02:09 +01:00
Billy Laws d4480c3566 X86Tables: Mark VMOVNTDQ as FLAGS_SF_MOD_MEM_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws b3a4de7aa8 X86Tables: Mark MASKMOVQ as FLAGS_SF_MOD_REG_ONLY 2025-04-09 23:01:26 +01:00
Billy Laws 60d8131a4d X86Tables: Drop FLAGS_SF_MOD_MEM_ONLY from (V)MOV(L/H)PS
These encodings are shared with MOVLHPS/MOVHLPS and FEX handles both variants.
2025-04-09 23:01:26 +01:00
Billy Laws 6d78aefa13 Frontend: Ignore REX register extension for MMX registers 2025-04-09 23:01:26 +01:00
Billy Laws badd953600 unittests: Test for REX.B being ignored with MMX registers 2025-04-09 23:01:26 +01:00
LC 5d385c540d Merge pull request #4489 from Sonicadvance1/reduce_stack_cpu_count
FEXCore: Reduce stack usage in CalculateNumberOfCPUs
2025-04-09 12:40:52 -04:00
Ryan Houdek 52351902ee Merge pull request #4488 from pmatos/EnableUBSAN
Add option to enable UBSAN
2025-04-09 00:13:29 -07:00
Paulo Matos 6e697e5ec1 Add option to enable UBSAN 2025-04-09 08:58:24 +02:00
Ryan Houdek c55c8979f1 AOTGen: Only calculate number of CPU cores once
Removes some per-iteration string processing and file IO.
2025-04-08 22:55:55 -07:00
Ryan Houdek 7ca757bb6d FEXCore: Move CPUInfo to FEX
This is only ever used in the frontend now.
2025-04-08 22:54:43 -07:00
Ryan Houdek c2d59b02bd FEXCore: Reduce stack usage in CalculateNumberOfCPUs
Nothing crazy, just recalculate the maximum string length rather than
use PATH_MAX.
2025-04-08 22:54:43 -07:00
LC 45b638fc45 Merge pull request #4491 from Sonicadvance1/push_pop_callee
FEXCore/Emitter: Stop creating a vector on the heap
2025-04-09 01:46:34 -04:00
Ryan Houdek 5767c61a91 FEXCore/Emitter: Stop creating a vector on the heap
In the Push/Pop CalleeSavedRegisters these vectors were getting created
on the heap, allocating memory and then just iterating them.

Just use a std::array which makes it stop allocating memory and saves
the number of instructions.
2025-04-08 18:36:47 -07:00
LC e3f8a817e0 Merge pull request #4490 from Sonicadvance1/remove_old_workaround
FEXCore/Allocator: Removes old workaround for kernel 4.17
2025-04-08 21:34:47 -04:00
Ryan Houdek 4377603d7e FEXCore/Allocator: Removes old workaround for kernel 4.17
This was only used for working around our old CI machines and now that
our minimum kernel requirement is 5.15 this isn't required anymore.
2025-04-08 18:22:39 -07:00
LC bc72239181 Merge pull request #4485 from Sonicadvance1/softfloat-3e_potara_cephes
cephes: Rewrite to use softfloat-3e 128-bit
2025-04-08 11:26:01 -04:00
Ryan Houdek 2d7f37386e Softfloat: Remove warnings
These precision warnings are no longer true!
2025-04-07 15:22:24 -07:00
Ryan Houdek dffba2d83c unittests/ASM: Enable x87 tests under simulator
Since the x86 simulator is now handling these instructions at 128-bit
precision, these now just work.
2025-04-07 15:13:04 -07:00
Ryan Houdek 0fbb6aa02f cephes: Rewrite to use softfloat-3e 128-bit
And also use it at the same time, since the function signatures changed.

Instead of relying on the host libc math libraries for `long double` ALU
operations, rewrite the entire thing to use softfloat-3e fixed width
float128_t types.

This is a very invasive change in cephes but is a necessary requirement
for getting the precision we require in environments that map `long
  double` to be the same as `double`, like Win32 and MacOS.

This fixes the precision issue in transcendental operations when running
under WINE.
2025-04-07 15:13:04 -07:00
Ryan Houdek 7364affd69 Softfloat-3e: Adds some missing files and passes state through 2025-04-07 14:57:48 -07:00
LC 09e622d5a0 Merge pull request #4482 from Sonicadvance1/softfloat_3e
Softfloat-3e: Add support for f128
2025-04-06 18:03:35 -04:00
Ryan Houdek 21ccaf56e6 Softfloat-3e: Add support for f128
This is going to be necessary soon.
2025-04-05 17:48:57 -07:00
Ryan Houdek d8cd807520 Softfloat-3e: Moves to Externals 2025-04-05 17:26:15 -07:00
LC 6069d05aa9 Merge pull request #4481 from Sonicadvance1/instcountci_ensure_consistent_rip
InstcountCI: Ensure RIP of blocks is consistent
2025-04-05 19:04:39 -04:00
Ryan Houdek 09a3d4851a InstcountCI: Update 2025-04-05 15:55:08 -07:00
Ryan Houdek d41374c55a InstcountCI: Ensure RIP of blocks is consistent
This changes the instcountCI code to consistently load test data in to
RIP 0x1'0000 so we don't have any spurious changes due to json changes.

This has been a minor annoyance where if a test was added, it had the
potential to shift the rest of the data in the tests. This now ensures
it is consistent.
2025-04-05 15:55:08 -07:00
LC dd8a1f00fc Merge pull request #4468 from Sonicadvance1/thats_the_softfloat_guarantee
FEXCore/Softfloat: Wire up cephes math library for transcendental operations
2025-04-04 19:11:55 -04:00
Ryan Houdek ffca27cbde FEXCore/Softfloat: Wire up cephes math library for transcendental operations
This is solving a different problem than what #4411 is specifically
trying to solve.

For our transcendental operations, we can't currently guarantee that
these functions will actually operate at the 128-bit softfloat
precision. While this is true with glibc, this is /not/ true for musl
and likely more libraries.

Instead of relying on our libc implementation to implement these,
instead include the cephes math library directly which is what most
people use for this. Including musl even, but not for all operations.

With this we are no longer beholden to the standard libraries for
providing a correct implementation.
2025-04-04 16:03:09 -07:00
Ryan Houdek 82edd901fe FEXCore/Common: Adds cephes math library
Only the few transcendental functions that FEX needs.
Disabled when building on x86-64.
2025-04-04 16:03:09 -07:00
388 changed files with 52563 additions and 179483 deletions

No files matched your search

-2
View File
@@ -3,8 +3,6 @@
# Ignore all files in the External directory
External/*
# SoftFloat-3e code doesn't belong to us
FEXCore/Source/Common/SoftFloat-3e/*
Source/Common/cpp-optparse/*
# Files with human-indented tables for readability - don't mess with these
+1 -1
View File
@@ -78,7 +78,7 @@ jobs:
# Note the current convention is to use the -S and -B options here to specify source
# and build directories, but this is only available with CMake 3.13 and higher.
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/Data/CMake/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
- name: Build
working-directory: ${{runner.workspace}}/build
+26 -11
View File
@@ -13,6 +13,7 @@ option(ENABLE_IWYU "Enables include what you use program" FALSE)
option(ENABLE_LTO "Enable LTO with compilation" TRUE)
option(ENABLE_XRAY "Enable building with LLVM X-Ray" FALSE)
set(USE_LINKER "" CACHE STRING "Allow overriding the linker path directly")
option(ENABLE_UBSAN "Enables Clang UBSAN" FALSE)
option(ENABLE_ASAN "Enables Clang ASAN" FALSE)
option(ENABLE_TSAN "Enables Clang TSAN" FALSE)
option(ENABLE_COVERAGE "Enables Coverage" FALSE)
@@ -34,10 +35,13 @@ set (FEXCORE_PROFILER_BACKEND "gpuvis" CACHE STRING "Set which backend to use fo
option(ENABLE_GLIBC_ALLOCATOR_HOOK_FAULT "Enables glibc memory allocation hooking with fault for CI testing")
option(USE_PDB_DEBUGINFO "Builds debug info in PDB format" FALSE)
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_32_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_32.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting i686")
set (X86_64_TOOLCHAIN_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/toolchain_x86_64.cmake" CACHE FILEPATH "Toolchain file for the (cross-)compiler targeting x86_64")
set (X86_DEV_ROOTFS "/" CACHE FILEPATH "Path to the sysroot used for cross-compiling for i686 and x86_64")
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu" CACHE PATH "global data directory")
set (DATA_DIRECTORY "" CACHE PATH "Global data directory (override)")
if (NOT DATA_DIRECTORY)
set (DATA_DIRECTORY "${CMAKE_INSTALL_PREFIX}/share/fex-emu")
endif()
string(FIND ${CMAKE_BASE_NAME} mingw CONTAINS_MINGW)
if (NOT CONTAINS_MINGW EQUAL -1)
@@ -93,7 +97,7 @@ endif()
# uninstall target
if(NOT TARGET uninstall)
configure_file(
"${CMAKE_CURRENT_SOURCE_DIR}/CMakeFiles/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/cmake_uninstall.cmake.in"
"${CMAKE_CURRENT_BINARY_DIR}/CMakeFiles/cmake_uninstall.cmake"
IMMEDIATE @ONLY)
@@ -228,6 +232,19 @@ if (NOT ENABLE_OFFLINE_TELEMETRY)
add_definitions(-DFEX_DISABLE_TELEMETRY=1)
endif()
if (ENABLE_UBSAN)
# See https://github.com/FEX-Emu/FEX/pull/4494#issuecomment-2800608944
# and related discussion for the use of -fno-sanitize=alignment -fno-sanitize=function
# with UBSAN.
# alignment: we don't follow a strict alignment policy, for example IR uses packed structs
# that are regularly access unaligned.
# function: syscalls cast function pointers to void (*)(unsigned long...), causing warnings
# related to this access.
add_definitions(-DENABLE_UBSAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
link_libraries(-fno-omit-frame-pointer -fsanitize=undefined -fno-sanitize=alignment -fno-sanitize=function -fno-sanitize-recover=undefined)
endif()
if (ENABLE_ASAN)
add_definitions(-DENABLE_ASAN=1)
add_compile_options(-fno-omit-frame-pointer -fsanitize=address -fsanitize-address-use-after-scope)
@@ -400,7 +417,7 @@ if (TUNE_CPU STREQUAL "native")
list(APPEND FEX_TUNE_COMPILE_FLAGS "-march=native")
endif()
endif()
else()
elseif (NOT TUNE_CPU STREQUAL "none")
check_cxx_compiler_flag("-mcpu=${TUNE_CPU}" COMPILER_SUPPORTS_CPU_TYPE)
if(COMPILER_SUPPORTS_CPU_TYPE)
list(APPEND FEX_TUNE_COMPILE_FLAGS "-mcpu=${TUNE_CPU}")
@@ -419,10 +436,6 @@ endif()
add_compile_options(-Wall)
configure_file(
${CMAKE_CURRENT_SOURCE_DIR}/include/Config.h.in
${CMAKE_BINARY_DIR}/generated/ConfigDefines.h)
include(CTest)
if (BUILD_TESTS)
message(STATUS "Unit tests are enabled")
@@ -440,6 +453,8 @@ if (BUILD_TESTS)
set(TEST_JOB_FLAG "-j${TEST_JOB_COUNT}")
endif()
add_subdirectory(External/SoftFloat-3e/)
add_subdirectory(External/cephes/)
add_subdirectory(FEXHeaderUtils/)
add_subdirectory(CodeEmitter/)
add_subdirectory(FEXCore/)
@@ -603,12 +618,12 @@ set (CPACK_PACKAGE_CONTACT "FEX-Emu Maintainers <team@fex-emu.com>")
set (CPACK_PACKAGE_VERSION_MAJOR "${FEX_VERSION_MAJOR}")
set (CPACK_PACKAGE_VERSION_MINOR "${FEX_VERSION_MINOR}")
set (CPACK_PACKAGE_VERSION_PATCH "${FEX_VERSION_PATCH}")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/CPack/Description.txt")
set (CPACK_PACKAGE_DESCRIPTION_FILE "${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/Description.txt")
# Debian defines
set (CPACK_DEBIAN_PACKAGE_DEPENDS "libc6, libstdc++6, libepoxy0, libsdl2-2.0-0, libegl1, libx11-6, squashfuse")
set (CPACK_DEBIAN_PACKAGE_CONTROL_EXTRA
"${CMAKE_CURRENT_SOURCE_DIR}/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/CPack/triggers")
"${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/postinst;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/prerm;${CMAKE_CURRENT_SOURCE_DIR}/Data/CMake/CPack/triggers")
if (CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64")
# binfmt_misc conflicts with qemu-user-static
# We also only install binfmt_misc on aarch64 hosts
-3
View File
@@ -1,3 +0,0 @@
x86 and x86-64 Linux emulator
FEX is very much work in progress, so expect things to change.
File diff suppressed because it is too large. Load diff
+4 -4
View File
@@ -47,13 +47,13 @@ public:
CurrentOffset += StringLength;
}
void Align() {
// Align the buffer to instruction size
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & 0b11;
void Align(size_t Size = 4) {
// Align the buffer to provided size.
auto CurrentAlignment = reinterpret_cast<uint64_t>(CurrentOffset) & (Size - 1);
if (!CurrentAlignment) {
return;
}
CurrentOffset += 4 - CurrentAlignment;
CurrentOffset += Size - CurrentAlignment;
}
template<typename T>
+5
View File
@@ -355,6 +355,7 @@ enum class SystemRegister : uint32_t {
TPIDRRO_EL0 = GenSystemReg<0b11, 0b011, 0b1101, 0b0000, 0b011>,
CNTFRQ_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b000>,
CNTVCT_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b010>,
CNTVCTSS_EL0 = GenSystemReg<0b11, 0b011, 0b1110, 0b0000, 0b110>,
};
template<uint32_t op1, uint32_t CRm, uint32_t op2>
@@ -572,6 +573,10 @@ enum class Rotation : uint32_t {
template<typename T>
concept IsXOrWRegister = std::is_same_v<T, XRegister> || std::is_same_v<T, WRegister>;
// Concept for contraining some instructions to accept only a QRegister or DRegister.
template<typename T>
concept IsQOrDRegister = std::is_same_v<T, QRegister> || std::is_same_v<T, DRegister>;
// Whether or not a given set of vector registers are sequential
// in increasing order as far as the register file is concerned (modulo its size)
//
+414 -12
View File
@@ -416,8 +416,7 @@ public:
0;
ASIMDSTLD<size, true, 1>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld1r(T rt, Register rn) {
constexpr uint32_t Op = 0b0000'1101'000 << 21;
constexpr uint32_t Opcode = 0b110;
@@ -435,8 +434,7 @@ public:
0;
ASIMDSTLD<size, true, 2>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld2r(T rt, T rt2, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2), "rt and rt2 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -455,8 +453,7 @@ public:
0;
ASIMDSTLD<size, true, 3>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld3r(T rt, T rt2, T rt3, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2, rt3), "rt, rt2, and rt3 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -475,8 +472,7 @@ public:
0;
ASIMDSTLD<size, true, 4>(Op, Opcode, rt, Index, rn, Reg::r0);
}
template<SubRegSize size, typename T>
requires (std::is_same_v<QRegister, T> || std::is_same_v<DRegister, T>)
template<SubRegSize size, IsQOrDRegister T>
void ld4r(T rt, T rt2, T rt3, T rt4, Register rn) {
LOGMAN_THROW_A_FMT(AreVectorsSequential(rt, rt2, rt3, rt4), "rt, rt2, rt3, and rt4 must be sequential");
constexpr uint32_t Op = 0b0000'1101'000 << 21;
@@ -1652,8 +1648,8 @@ public:
template<typename T>
void ASIMDLoadStoreSinglePost(uint32_t Op, uint32_t Q, uint32_t L, uint32_t R, uint32_t opcode, uint32_t S, uint32_t size,
ARMEmitter::Register rm, ARMEmitter::Register rn, T rt) {
LOGMAN_THROW_A_FMT(std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>, "Only supports 128-bit and "
"64-bit vector registers.");
LOGMAN_THROW_A_FMT((std::is_same_v<ARMEmitter::QRegister, T> || std::is_same_v<ARMEmitter::DRegister, T>), "Only supports 128-bit and "
"64-bit vector registers.");
uint32_t Instr = Op;
Instr |= Q << 30;
@@ -2136,7 +2132,370 @@ public:
}
// Memory copy/set
// TODO
void cpyfp(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0000, rs, rn, rd);
}
void cpyfm(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0000, rs, rn, rd);
}
void cpyfe(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0000, rs, rn, rd);
}
void cpyfpwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0001, rs, rn, rd);
}
void cpyfmwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0001, rs, rn, rd);
}
void cpyfewt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0001, rs, rn, rd);
}
void cpyfprt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0010, rs, rn, rd);
}
void cpyfmrt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0010, rs, rn, rd);
}
void cpyfert(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0010, rs, rn, rd);
}
void cpyfpt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0011, rs, rn, rd);
}
void cpyfmt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0011, rs, rn, rd);
}
void cpyfet(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0011, rs, rn, rd);
}
void cpyfpwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0100, rs, rn, rd);
}
void cpyfmwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0100, rs, rn, rd);
}
void cpyfewn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0100, rs, rn, rd);
}
void cpyfpwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0101, rs, rn, rd);
}
void cpyfmwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0101, rs, rn, rd);
}
void cpyfewtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0101, rs, rn, rd);
}
void cpyfprtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0110, rs, rn, rd);
}
void cpyfmrtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0110, rs, rn, rd);
}
void cpyfertwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0110, rs, rn, rd);
}
void cpyfptwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b0111, rs, rn, rd);
}
void cpyfmtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b0111, rs, rn, rd);
}
void cpyfetwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b0111, rs, rn, rd);
}
void cpyfprn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1000, rs, rn, rd);
}
void cpyfmrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1000, rs, rn, rd);
}
void cpyfern(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1000, rs, rn, rd);
}
void cpyfpwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1001, rs, rn, rd);
}
void cpyfmwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1001, rs, rn, rd);
}
void cpyfewtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1001, rs, rn, rd);
}
void cpyfprtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1010, rs, rn, rd);
}
void cpyfmrtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1010, rs, rn, rd);
}
void cpyfertrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1010, rs, rn, rd);
}
void cpyfptrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1011, rs, rn, rd);
}
void cpyfmtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1011, rs, rn, rd);
}
void cpyfetrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1011, rs, rn, rd);
}
void cpyfpn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1100, rs, rn, rd);
}
void cpyfmn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1100, rs, rn, rd);
}
void cpyfen(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1100, rs, rn, rd);
}
void cpyfpwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1101, rs, rn, rd);
}
void cpyfmwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1101, rs, rn, rd);
}
void cpyfewtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1101, rs, rn, rd);
}
void cpyfprtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1110, rs, rn, rd);
}
void cpyfmrtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1110, rs, rn, rd);
}
void cpyfertn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1110, rs, rn, rd);
}
void cpyfptn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b00, 0b1111, rs, rn, rd);
}
void cpyfmtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b01, 0b1111, rs, rn, rd);
}
void cpyfetn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 0, 0b10, 0b1111, rs, rn, rd);
}
void setp(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0000, rs, rn, rd);
}
void setm(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0100, rs, rn, rd);
}
void sete(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1000, rs, rn, rd);
}
void setpt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0001, rs, rn, rd);
}
void setmt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0101, rs, rn, rd);
}
void setet(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1001, rs, rn, rd);
}
void setpn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0010, rs, rn, rd);
}
void setmn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0110, rs, rn, rd);
}
void seten(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1010, rs, rn, rd);
}
void setptn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0011, rs, rn, rd);
}
void setmtn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b0111, rs, rn, rd);
}
void setetn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 0, 0b11, 0b1011, rs, rn, rd);
}
void cpyp(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0000, rs, rn, rd);
}
void cpym(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0000, rs, rn, rd);
}
void cpye(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0000, rs, rn, rd);
}
void cpypwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0001, rs, rn, rd);
}
void cpymwt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0001, rs, rn, rd);
}
void cpyewt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0001, rs, rn, rd);
}
void cpyprt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0010, rs, rn, rd);
}
void cpymrt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0010, rs, rn, rd);
}
void cpyert(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0010, rs, rn, rd);
}
void cpypt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0011, rs, rn, rd);
}
void cpymt(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0011, rs, rn, rd);
}
void cpyet(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0011, rs, rn, rd);
}
void cpypwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0100, rs, rn, rd);
}
void cpymwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0100, rs, rn, rd);
}
void cpyewn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0100, rs, rn, rd);
}
void cpypwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0101, rs, rn, rd);
}
void cpymwtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0101, rs, rn, rd);
}
void cpyewtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0101, rs, rn, rd);
}
void cpyprtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0110, rs, rn, rd);
}
void cpymrtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0110, rs, rn, rd);
}
void cpyertwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0110, rs, rn, rd);
}
void cpyptwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b0111, rs, rn, rd);
}
void cpymtwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b0111, rs, rn, rd);
}
void cpyetwn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b0111, rs, rn, rd);
}
void cpyprn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1000, rs, rn, rd);
}
void cpymrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1000, rs, rn, rd);
}
void cpyern(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1000, rs, rn, rd);
}
void cpypwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1001, rs, rn, rd);
}
void cpymwtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1001, rs, rn, rd);
}
void cpyewtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1001, rs, rn, rd);
}
void cpyprtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1010, rs, rn, rd);
}
void cpymrtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1010, rs, rn, rd);
}
void cpyertrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1010, rs, rn, rd);
}
void cpyptrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1011, rs, rn, rd);
}
void cpymtrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1011, rs, rn, rd);
}
void cpyetrn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1011, rs, rn, rd);
}
void cpypn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1100, rs, rn, rd);
}
void cpymn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1100, rs, rn, rd);
}
void cpyen(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1100, rs, rn, rd);
}
void cpypwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1101, rs, rn, rd);
}
void cpymwtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1101, rs, rn, rd);
}
void cpyewtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1101, rs, rn, rd);
}
void cpyprtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1110, rs, rn, rd);
}
void cpymrtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1110, rs, rn, rd);
}
void cpyertn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1110, rs, rn, rd);
}
void cpyptn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b00, 0b1111, rs, rn, rd);
}
void cpymtn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b01, 0b1111, rs, rn, rd);
}
void cpyetn(Register rd, Register rs, Register rn) {
MemoryCopyAndMemorySet(0, 1, 0b10, 0b1111, rs, rn, rd);
}
void setgp(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0000, rs, rn, rd);
}
void setgm(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0100, rs, rn, rd);
}
void setge(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1000, rs, rn, rd);
}
void setgpt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0001, rs, rn, rd);
}
void setgmt(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0101, rs, rn, rd);
}
void setget(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1001, rs, rn, rd);
}
void setgpn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0010, rs, rn, rd);
}
void setgmn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0110, rs, rn, rd);
}
void setgen(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1010, rs, rn, rd);
}
void setgptn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0011, rs, rn, rd);
}
void setgmtn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b0111, rs, rn, rd);
}
void setgetn(Register rd, Register rn, Register rs) {
MemoryCopyAndMemorySet(0, 1, 0b11, 0b1011, rs, rn, rd);
}
// Loadstore no-allocate pair
void stnp(ARMEmitter::WRegister rt, ARMEmitter::WRegister rt2, ARMEmitter::Register rn, int32_t Imm) {
LOGMAN_THROW_A_FMT(Imm >= -256 && Imm <= 252 && ((Imm & 0b11) == 0), "Unscaled offset too large");
@@ -3819,7 +4178,12 @@ public:
}
// Loadstore PAC
// TODO
void ldraa(XRegister rt, XRegister rn, IndexType type, int32_t offset = 0) {
LoadStorePAC(0b11, 0, 0, offset, type, rn, rt);
}
void ldrab(XRegister rt, XRegister rn, IndexType type, int32_t offset = 0) {
LoadStorePAC(0b11, 0, 1, offset, type, rn, rt);
}
// Loadstore unsigned immediate
// Maximum values of unsigned immediate offsets for particular data sizes.
@@ -3968,6 +4332,20 @@ private:
dc32(Instr);
}
void MemoryCopyAndMemorySet(uint32_t sz, uint32_t o0, uint32_t op1, uint32_t op2, Register rs, Register rn, Register rd) {
uint32_t Instr = 0b0001'1001'0000'0000'0000'0100'0000'0000;
Instr |= sz << 30;
Instr |= o0 << 26;
Instr |= op1 << 22;
Instr |= rs.Idx() << 16;
Instr |= op2 << 12;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Loadstore no-allocate pair
template<typename T>
void LoadStoreNoAllocate(uint32_t Op, T rt, T rt2, ARMEmitter::Register rn, uint32_t Imm) {
@@ -4034,6 +4412,30 @@ private:
dc32(Instr);
}
void LoadStorePAC(uint32_t size, uint32_t VR, uint32_t M, int32_t imm, IndexType type, Register rn, Register rt) {
LOGMAN_THROW_A_FMT((imm % 8) == 0, "Immediate ({}) must be divisible by 8", imm);
LOGMAN_THROW_A_FMT(imm >= -4096 && imm <= 4088, "Immediate ({}) must be within [-4096, 4088]", imm);
LOGMAN_THROW_A_FMT(type == IndexType::OFFSET || type == IndexType::PRE, "PAC may only use offset or pre-indexed values");
// The immediate is scaled down in order to fit within the available 10 immediate bits.
const auto scaled_imm = static_cast<uint32_t>(imm / 8);
const auto imm9 = scaled_imm & 0b1'1111'1111;
const auto S = (scaled_imm >> 9) & 1;
const auto W = type == IndexType::OFFSET ? 0U : 1U;
uint32_t Instr = 0b0011'1000'0010'0000'0000'0100'0000'0000;
Instr |= size << 30;
Instr |= VR << 26;
Instr |= M << 23;
Instr |= S << 22;
Instr |= imm9 << 12;
Instr |= W << 11;
Instr |= rn.Idx() << 5;
Instr |= rt.Idx();
dc32(Instr);
}
// Loadstore unsigned immediate
template<typename T>
void LoadStoreUnsigned(uint32_t size, uint32_t V, uint32_t opc, T rt, Register rn, uint32_t Imm) {
+106 -7
View File
@@ -137,7 +137,15 @@ public:
}
// Advanced SIMD scalar three same extra
// XXX:
void sqrdmlah(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i32Bit, "Only supports 16/32-bit");
ASIMDScalarThreeSameExtra(1, size, 0b0000, rm, rn, rd);
}
void sqrdmlsh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm) {
LOGMAN_THROW_A_FMT(size == ScalarRegSize::i16Bit || size == ScalarRegSize::i32Bit, "Only supports 16/32-bit");
ASIMDScalarThreeSameExtra(1, size, 0b0001, rm, rn, rd);
}
// Advanced SIMD scalar two-register miscellaneous
void suqadd(ScalarRegSize size, VRegister rd, VRegister rn) {
ASIMDScalar2RegMisc(0, 0, size, 0b00011, rd, rn);
@@ -744,10 +752,51 @@ public:
const uint32_t immb = InvertedShift & 0b111;
ASIMDScalarShiftByImm(1, immh, immb, 0b10011, rd, rn);
}
// TODO: UCVTF, FCVTZU
// Advanced SIMD scalar x indexed element
// XXX:
//
void sqdmlal(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b0011, rm, rn, rd, index);
}
void sqdmlsl(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b0111, rm, rn, rd, index);
}
void sqdmull(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1011, rm, rn, rd, index);
}
void sqdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1100, rm, rn, rd, index);
}
void sqrdmulh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(0, size, 0b1101, rm, rn, rd, index);
}
void fmla(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b0001, rm, rn, rd, index);
}
void fmls(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b0101, rm, rn, rd, index);
}
void fmul(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(0, size, 0b1001, rm, rn, rd, index);
}
void sqrdmlah(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(1, size, 0b1101, rm, rn, rd, index);
}
void sqrdmlsh(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i64Bit, "Scalar size must not be 64-bit");
ASIMDScalarXIndexedElement(1, size, 0b1111, rm, rn, rd, index);
}
void fmulx(ScalarRegSize size, VRegister rd, VRegister rn, VRegister rm, uint32_t index) {
ASIMDScalarXIndexedElement(1, size, 0b1001, rm, rn, rd, index);
}
// Floating-point data-processing (1 source)
void fmov(ScalarRegSize size, VRegister rd, VRegister rn) {
Float1Source(size, 0, 0, 0b000000, rd, rn);
@@ -1269,7 +1318,17 @@ private:
}
// Advanced SIMD scalar three same extra
// XXX:
void ASIMDScalarThreeSameExtra(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd) {
uint32_t Instr = 0b0101'1110'0000'0000'1000'0100'0000'0000;
Instr |= U << 29;
Instr |= FEXCore::ToUnderlying(size) << 22;
Instr |= rm.Idx() << 16;
Instr |= opcode << 11;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Advanced SIMD scalar two-register miscellaneous
void ASIMDScalar2RegMisc(uint32_t b20, uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'1000'0000'0000;
@@ -1283,8 +1342,6 @@ private:
dc32(Instr);
}
// Advanced SIMD scalar pairwise
// XXX:
// Advanced SIMD scalar three different
void ASIMD3RegDifferent(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rd, VRegister rn, VRegister rm) {
uint32_t Instr = 0b0101'1110'0010'0000'0000'0000'0000'0000;
@@ -1321,8 +1378,50 @@ private:
Instr |= Encode_rd(rd);
dc32(Instr);
}
// Advanced SIMD scalar x indexed element
// XXX:
void ASIMDScalarXIndexedElement(uint32_t U, ScalarRegSize size, uint32_t opcode, VRegister rm, VRegister rn, VRegister rd, uint32_t index) {
LOGMAN_THROW_A_FMT(size != ScalarRegSize::i8Bit, "Scalar size must not be 8-bit");
[[maybe_unused]] const auto invalid_bound = 16U >> FEXCore::ToUnderlying(size);
LOGMAN_THROW_A_FMT(index < invalid_bound, "Index ({}) must be within [0-{}]", index, invalid_bound - 1);
uint32_t Instr = 0b0101'1111'0000'0000'0000'0000'0000'0000;
// FMUL/FMLA/FMLS indexed variants deal with size differently.
if (opcode == 0b0001 || opcode == 0b0101 || opcode == 0b1001) {
// Unlike other instructions in the group, 16-bit is encoded as zero
// and 32/64-bit are encoded with the top bit always set to one.
if (size != ScalarRegSize::i16Bit) {
Instr |= (0b10 | (FEXCore::ToUnderlying(size) & 1)) << 22;
}
} else {
Instr |= FEXCore::ToUnderlying(size) << 22;
}
uint32_t H = 0;
uint32_t LM = 0;
if (size == ScalarRegSize::i16Bit) {
LOGMAN_THROW_A_FMT(rm <= VReg::v15, "rm ({}) must be within [v0-v15]", rm.Idx());
H = (index >> 2) & 1;
LM = index & 0b11;
} else if (size == ScalarRegSize::i32Bit) {
H = (index >> 1) & 1;
LM = (index & 0b01) << 1;
} else {
H = index & 1;
}
Instr |= U << 29;
Instr |= LM << 20;
Instr |= rm.Idx() << 16;
Instr |= opcode << 12;
Instr |= H << 11;
Instr |= rn.Idx() << 5;
Instr |= rd.Idx();
dc32(Instr);
}
// Floating-point data-processing (1 source)
void Float1Source(uint32_t M, uint32_t S, uint32_t ptype, uint32_t opcode, VRegister rd, VRegister rn) {
uint32_t Instr = 0b0001'1110'0010'0000'0100'0000'0000'0000;
+6
View File
@@ -13,6 +13,12 @@ struct EmitterOps : Emitter {
#endif
public:
// Reserved
void udf(uint32_t Imm) {
LOGMAN_THROW_A_FMT(Imm < 0x1'0000, "Immediate needs to be 16-bit");
dc32(Imm);
}
// System with result
// TODO: SYSL
// System Instruction
File renamed without changes.
File renamed without changes.
File renamed without changes.
+3
View File
@@ -0,0 +1,3 @@
x86 and x86-64 Linux emulator
FEX allows you to run x86 applications on ARM64 Linux devices. It offers broad compatibility with both 32-bit and 64-bit binaries, and it can be used alongside Wine/Proton to play Windows games.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
@@ -1,6 +1,6 @@
# This is a reference AArch64 cross compile script
# Pass in to cmake when building:
# eg: cmake -DCMAKE_TOOLCHAIN_FILE=../CMakeToolchains/AArch64.cmake ..
# eg: cmake --toolchain ../Data/CMake/toolchain_aarch64.cmake ..
if (NOT DEFINED ENV{SYSROOT})
message(FATAL_ERROR "Need to have SYSROOT environment variable set")
endif()
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
+31
View File
@@ -0,0 +1,31 @@
# --- Stage 1: Builder ---
FROM ubuntu:22.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-13 llvm-13 nasm ninja-build pkg-config \
libcap-dev libglfw3-dev libepoxy-dev python3-dev libsdl2-dev \
python3 linux-headers-generic \
git qtbase5-dev qtdeclarative5-dev lld
RUN git clone --recurse-submodules https://github.com/FEX-Emu/FEX.git
WORKDIR /FEX
RUN mkdir build
ARG CC=clang-13
ARG CXX=clang++-13
RUN cmake -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_BUILD_TYPE=Release -DUSE_LINKER=lld -DENABLE_LTO=True -DBUILD_TESTS=False -DENABLE_ASSERTIONS=False -G Ninja .
RUN ninja
WORKDIR /FEX/build
# --- Stage 2: Runner ---
FROM builder as runner
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y \
libcap-dev libglfw3-dev libepoxy-dev
COPY --from=builder /FEX/Bin/* /usr/bin/
WORKDIR /
+35
View File
@@ -0,0 +1,35 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
gcc32 = pkgs.writeText "toolchain_nix_gcc_x86_32.txt" ''
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.gcc}/bin/i686-unknown-linux-gnu-g++)
'';
gcc64 = pkgs.writeText "toolchain_nix_gcc_x86_64.txt" ''
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-gcc)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.gcc}/bin/x86_64-unknown-linux-gnu-g++)
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "toolchain32: ${gcc32}"
echo "toolchain64: ${gcc64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${gcc32} -DX86_64_TOOLCHAIN_FILE=${gcc64}";
}
+83
View File
@@ -0,0 +1,83 @@
{ pkgs ? import <nixpkgs> { } }:
let
pkgsCross32 = pkgs.pkgsCross.gnu32;
pkgsCross64 = pkgs.pkgsCross.gnu64;
devRootFS = pkgs.buildEnv {
name = "fex-dev-rootfs";
paths = [
pkgsCross64.stdenv.cc.libc_dev
pkgsCross32.stdenv.cc.libc_dev
pkgsCross64.stdenv.cc.cc
pkgsCross32.stdenv.cc.cc
pkgs.alsa-lib.dev
pkgs.libdrm.dev
pkgs.libGL.dev
pkgs.wayland.dev
pkgs.xorg.libX11.dev
pkgs.xorg.libxcb.dev
pkgs.xorg.libXrandr.dev
pkgs.xorg.libXrender.dev
pkgs.xorg.xorgproto
];
ignoreCollisions = true;
pathsToLink = [
"/include"
"/lib"
];
postBuild = ''
mkdir -p $out/usr
ln -s $out/include $out/usr/
'';
};
toolchain32 = pkgs.writeText "toolchain_nix_x86_32.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR i686)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross32.buildPackages.clang}/bin/i686-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target i686-linux-gnu -msse2 -mfpmath=sse --sysroot=${devRootFS} -iwithsysroot/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
toolchain64 = pkgs.writeText "toolchain_nix_x86_64.txt" ''
set(CMAKE_EXE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_MODULE_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SHARED_LINKER_FLAGS_INIT "-fuse-ld=lld")
set(CMAKE_SYSTEM_PROCESSOR x86_64)
set(CMAKE_C_COMPILER clang)
set(CMAKE_CXX_COMPILER clang++)
set(CMAKE_C_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang)
set(CMAKE_CXX_COMPILER ${pkgsCross64.buildPackages.clang}/bin/x86_64-unknown-linux-gnu-clang++)
set(CLANG_FLAGS "-nodefaultlibs -nostartfiles -lstdc++ -target x86_64-linux-gnu --sysroot=${devRootFS} -iwithsysroot/usr/include")
set(CMAKE_C_FLAGS "''${CMAKE_C_FLAGS} ''${CLANG_FLAGS}")
set(CMAKE_CXX_FLAGS "''${CMAKE_CXX_FLAGS} ''${CLANG_FLAGS}")
'';
in
pkgs.mkShell {
buildInputs = [
pkgsCross64.buildPackages.clang
pkgsCross32.buildPackages.clang
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "Set up dev RootFS at ${devRootFS}"
echo "toolchain32: ${toolchain32}"
echo "toolchain64: ${toolchain64}"
echo ""
echo "Use \$FEX_CMAKE_TOOLCHAINS to configure CMake."
fi
'';
FEX_CMAKE_TOOLCHAINS = "-DX86_32_TOOLCHAIN_FILE=${toolchain32} -DX86_64_TOOLCHAIN_FILE=${toolchain64} -DX86_DEV_ROOTFS=${devRootFS}";
ROOTFS = "${devRootFS}";
}
+52
View File
@@ -0,0 +1,52 @@
{ pkgs ? import <nixpkgs> { } }:
let
toolchain = pkgs.fetchzip {
url = "https://github.com/bylaws/llvm-mingw/releases/download/20250305/llvm-mingw-20250305-ucrt-ubuntu-20.04-aarch64.tar.xz";
sha256 = "sha256-cA03/ab9O61eO9+S2JzIXD4V0HzTXK5/AYyxW2d73Po=";
};
cmakeToolchainFile = pkgs.substitute {
# Use absolute paths that are discoverable outside of the nix shell
src = ../../CMake/toolchain_mingw.cmake;
substitutions = ["--replace-fail" "\${MINGW_TRIPLE}-" "${toolchain}/bin/\${MINGW_TRIPLE}-"];
};
mesonCrossFile = pkgs.writeText "crossfile_llvm_mingw.txt" ''
[binaries]
ar = '${toolchain}/bin/arm64ec-w64-mingw32-ar'
c = '${toolchain}/bin/arm64ec-w64-mingw32-gcc'
cpp = '${toolchain}/bin/arm64ec-w64-mingw32-g++'
ld = '${toolchain}/bin/arm64ec-w64-mingw32-ld'
windres = '${toolchain}/bin/arm64ec-w64-mingw32-windres'
strip = '${toolchain}/bin/strip'
widl = '${toolchain}/bin/arm64ec-w64-mingw32-widl'
pkgconfig = 'aarch64-linux-gnu-pkg-config'
[host_machine]
system = 'windows'
cpu_family = 'aarch64'
cpu = 'aarch64'
endian = 'little'
'';
in
pkgs.mkShell {
buildInputs = [
toolchain
];
shellHook = ''
if [[ $- == *i* ]]; then
echo "llvm-mingw set up at ${toolchain}."
echo ""
echo "To configure DXVK/vkd3d-proton: meson setup \$FEX_MESON_CROSSFILE"
echo ""
echo "To configure 32-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_WOW64"
echo "To configure 64-bit FEX build: cmake \$FEX_CMAKE_TOOLCHAIN_ARM64EC"
fi
'';
# E.g. cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False
FEX_CMAKE_TOOLCHAIN_ARM64EC = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=arm64ec-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_CMAKE_TOOLCHAIN_WOW64 = "--toolchain ${cmakeToolchainFile} -DMINGW_TRIPLE=aarch64-w64-mingw32 -DCMAKE_INSTALL_LIBDIR=/usr/lib/wine/aarch64-windows";
FEX_MESON_CROSSFILE = "--cross-file ${mesonCrossFile}";
}
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 32-bit applications in Wine/Proton.
# The required cross-toolchains will be set up and managed by nix.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_WOW64 -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+21
View File
@@ -0,0 +1,21 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash WineOnArm/shell.nix
# Helper script to configure CMake for building FEX as library for emulation
# of 64-bit applications in Wine/Proton
# Nix is used to install and manage the required cross-toolchains.
if [ $# -eq 0 ]
then
echo "Expected CMake argument list"
exit 1
fi
if [ -f CMakeCache.txt ]
then
echo "Expected empty build folder"
exit 1
fi
set -o xtrace
cmake $FEX_CMAKE_TOOLCHAIN_ARM64EC -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=/usr -DENABLE_LTO=False -DBUILD_TESTS=False $@
+17
View File
@@ -0,0 +1,17 @@
#! /usr/bin/env nix-shell
#! nix-shell -i bash FEXLinuxTests/shell.nix
# Helper script to configure CMake for building FEXLinuxTests.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf unittests/FEXLinuxTests
set -o xtrace
cmake . $FEX_CMAKE_TOOLCHAINS -DBUILD_TESTS=ON -DBUILD_FEX_LINUX_TESTS=ON
+22
View File
@@ -0,0 +1,22 @@
# Helper script to configure CMake for library forwarding in FEX.
# Nix is used to install and manage the required cross-toolchains.
if [ ! -f CMakeCache.txt ]
then
echo "Must be run from a pre-configured CMake build folder"
exit 1
fi
# Remove previous build to ensure the new toolchain is applied
rm -rf guest-libs guest-libs-32 Guest Guest_32
# Set clang executable path manually since the one from the nix store
# will be picked up otherwise
CLANG_EXEC_PATH=""
if ! grep -q CLANG_EXEC_PATH CMakeCache.txt
then
CLANG_EXEC_PATH="-DCLANG_EXEC_PATH=`which clang`"
fi
nix-shell `dirname -- "$0"`/LibraryForwarding/shell.nix \
--run "set -o xtrace; cmake . \$FEX_CMAKE_TOOLCHAINS -DBUILD_THUNKS=ON $CLANG_EXEC_PATH; set +o xtrace"
-31
View File
@@ -1,31 +0,0 @@
# --- Stage 1: Builder ---
FROM ubuntu:20.04 as builder
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y cmake \
clang-10 llvm-10 nasm ninja-build pkg-config \
libcap-dev libglfw3-dev libepoxy-dev python3-dev libsdl2-dev \
python3 linux-headers-generic \
git
RUN git clone --recurse-submodules https://github.com/FEX-Emu/FEX.git
CMD [ "mkdir /opt/FEX/build" ]
WORKDIR /opt/FEX/build
ARG CC=clang-10
ARG CXX=clang++-10
RUN cmake -G Ninja .. -DCMAKE_BUILD_TYPE=Release
RUN ninja
# --- Stage 2: Runner ---
FROM ubuntu:20.04
RUN DEBIAN_FRONTEND="noninteractive" apt-get update
RUN DEBIAN_FRONTEND="noninteractive" apt install -y \
libcap-dev libglfw3-dev libepoxy-dev
COPY --from=builder /opt/FEX/build/Bin/* /usr/bin/
WORKDIR /root
+98
View File
@@ -0,0 +1,98 @@
set (SRCS
# F80 support
src/extF80_add.c
src/extF80_div.c
src/extF80_sub.c
src/extF80_mul.c
src/extF80_rem.c
src/extF80_sqrt.c
src/extF80_le.c
src/extF80_to_i32.c
src/extF80_to_i64.c
src/extF80_to_ui64.c
src/extF80_to_f32.c
src/extF80_to_f64.c
src/i32_to_extF80.c
src/ui64_to_extF80.c
src/extF80_to_f128.c
src/f128_to_extF80.c
# F128 support
src/f128_add.c
src/f128_div.c
src/f128_eq.c
src/f128_eq_signaling.c
src/f128_isSignalingNaN.c
src/f128_le.c
src/f128_le_quiet.c
src/f128_lt.c
src/f128_lt_quiet.c
src/f128_mulAdd.c
src/f128_mul.c
src/f128_rem.c
src/f128_sqrt.c
src/f128_sub.c
src/f128_to_f16.c
src/f128_to_f32.c
src/f128_to_f64.c
src/f128_to_i32.c
src/f128_to_i64.c
src/f128_to_ui32.c
src/f128_to_ui64.c
src/s_addMagsF128.c
src/s_subMagsF128.c
src/s_normRoundPackToF128.c
src/s_roundPackToF128.c
src/s_propagateNaNF128UI.c
# Conversion
src/f32_to_f128.c
src/i32_to_f128.c
src/s_roundToUI64.c
src/s_f128UIToCommonNaN.c
src/s_commonNaNToF128UI.c
src/s_normSubnormalF128Sig.c
src/s_roundToI32.c
src/s_roundToI64.c
src/s_roundPackToF32.c
src/s_addMagsExtF80.c
src/s_extF80UIToCommonNaN.c
src/s_commonNaNToF32UI.c
src/s_commonNaNToF64UI.c
src/s_roundPackToF64.c
src/s_propagateNaNExtF80UI.c
src/s_roundPackToExtF80.c
src/s_normSubnormalExtF80Sig.c
src/s_subMagsExtF80.c
src/s_shiftRightJam128.c
src/s_shiftRightJam128Extra.c
src/s_normRoundPackToExtF80.c
src/s_approxRecip_1Ks.c
src/s_approxRecipSqrt32_1.c
src/s_approxRecipSqrt_1Ks.c
src/softfloat_raiseFlags.c
src/f64_to_extF80.c
src/s_commonNaNToExtF80UI.c
src/s_normSubnormalF64Sig.c
src/s_f64UIToCommonNaN.c
src/extF80_roundToInt.c
src/extF80_eq.c
src/extF80_lt.c
src/f32_to_extF80.c
src/s_normSubnormalF32Sig.c
src/s_f32UIToCommonNaN.c)
if (_M_ARM_64 AND HAS_CLANG_PRESERVE_ALL)
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=__attribute__((preserve_all));-DFEXCORE_HAS_PRESERVE_ALL_ATTR=1")
else()
list(APPEND DEFINES "-DFEXCORE_PRESERVE_ALL_ATTR=;-DFEXCORE_HAS_PRESERVE_ALL_ATTR=0")
endif()
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ=1;-DINLINE=static inline;-DINLINE_LEVEL=4;-DSOFTFLOAT_FAST_INT64=1;-DSOFTFLOAT_FAST_DIV32TO16=1;-DSOFTFLOAT_FAST_DIV64TO32=1")
add_library(softfloat_3e STATIC ${SRCS})
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/)
target_include_directories(softfloat_3e PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}/include/SoftFloat-3e/)
target_compile_definitions(softfloat_3e PUBLIC ${DEFINES})
@@ -149,7 +149,7 @@ float64_t f32_to_f64( float32_t );
#ifdef SOFTFLOAT_FAST_INT64
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f32_to_extF80( struct softfloat_state *, float32_t );
float128_t f32_to_f128( float32_t );
float128_t f32_to_f128( struct softfloat_state *, float32_t );
#endif
void f32_to_extF80M( float32_t, extFloat80_t * );
void f32_to_f128M( float32_t, float128_t * );
@@ -243,13 +243,17 @@ FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t extF80_sqrt( struct softfloat_state *, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_eq( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_le( extFloat80_t, extFloat80_t );
bool extF80_le( struct softfloat_state *, extFloat80_t, extFloat80_t );
FEXCORE_PRESERVE_ALL_ATTR
bool extF80_lt( struct softfloat_state *, extFloat80_t, extFloat80_t );
bool extF80_eq_signaling( extFloat80_t, extFloat80_t );
bool extF80_le_quiet( extFloat80_t, extFloat80_t );
bool extF80_lt_quiet( extFloat80_t, extFloat80_t );
bool extF80_isSignalingNaN( extFloat80_t );
static inline extFloat80_t extF80_complement_sign(extFloat80_t a) {
a.signExp ^= 1ULL << 15;
return a;
}
#endif
uint_fast32_t extF80M_to_ui32( const extFloat80_t *, uint_fast8_t, bool );
uint_fast64_t extF80M_to_ui64( const extFloat80_t *, uint_fast8_t, bool );
@@ -284,34 +288,38 @@ bool extF80M_isSignalingNaN( const extFloat80_t * );
| 128-bit (quadruple-precision) floating-point operations.
*----------------------------------------------------------------------------*/
#ifdef SOFTFLOAT_FAST_INT64
uint_fast32_t f128_to_ui32( float128_t, uint_fast8_t, bool );
uint_fast64_t f128_to_ui64( float128_t, uint_fast8_t, bool );
int_fast32_t f128_to_i32( float128_t, uint_fast8_t, bool );
int_fast64_t f128_to_i64( float128_t, uint_fast8_t, bool );
uint_fast32_t f128_to_ui32( struct softfloat_state *, float128_t, uint_fast8_t, bool );
uint_fast64_t f128_to_ui64( struct softfloat_state *, float128_t, uint_fast8_t, bool );
int_fast32_t f128_to_i32( struct softfloat_state *, float128_t, uint_fast8_t, bool );
int_fast64_t f128_to_i64( struct softfloat_state *, float128_t, uint_fast8_t, bool );
uint_fast32_t f128_to_ui32_r_minMag( float128_t, bool );
uint_fast64_t f128_to_ui64_r_minMag( float128_t, bool );
int_fast32_t f128_to_i32_r_minMag( float128_t, bool );
int_fast64_t f128_to_i64_r_minMag( float128_t, bool );
float16_t f128_to_f16( float128_t );
float32_t f128_to_f32( float128_t );
float64_t f128_to_f64( float128_t );
float16_t f128_to_f16( struct softfloat_state *, float128_t );
float32_t f128_to_f32( struct softfloat_state *, float128_t );
float64_t f128_to_f64( struct softfloat_state *, float128_t );
FEXCORE_PRESERVE_ALL_ATTR
extFloat80_t f128_to_extF80( struct softfloat_state *, float128_t );
float128_t f128_roundToInt( float128_t, uint_fast8_t, bool );
float128_t f128_add( float128_t, float128_t );
float128_t f128_sub( float128_t, float128_t );
float128_t f128_mul( float128_t, float128_t );
float128_t f128_mulAdd( float128_t, float128_t, float128_t );
float128_t f128_div( float128_t, float128_t );
float128_t f128_rem( float128_t, float128_t );
float128_t f128_sqrt( float128_t );
bool f128_eq( float128_t, float128_t );
bool f128_le( float128_t, float128_t );
bool f128_lt( float128_t, float128_t );
bool f128_eq_signaling( float128_t, float128_t );
bool f128_le_quiet( float128_t, float128_t );
bool f128_lt_quiet( float128_t, float128_t );
float128_t f128_add( struct softfloat_state *, float128_t, float128_t );
float128_t f128_sub( struct softfloat_state *, float128_t, float128_t );
float128_t f128_mul( struct softfloat_state *, float128_t, float128_t );
float128_t f128_mulAdd( struct softfloat_state *, float128_t, float128_t, float128_t );
float128_t f128_div( struct softfloat_state *, float128_t, float128_t );
float128_t f128_rem( struct softfloat_state *, float128_t, float128_t );
float128_t f128_sqrt( struct softfloat_state *, float128_t );
bool f128_eq( struct softfloat_state *, float128_t, float128_t );
bool f128_le( struct softfloat_state *, float128_t, float128_t );
bool f128_lt( struct softfloat_state *, float128_t, float128_t );
bool f128_eq_signaling( struct softfloat_state *, float128_t, float128_t );
bool f128_le_quiet( struct softfloat_state *, float128_t, float128_t );
bool f128_lt_quiet( struct softfloat_state *, float128_t, float128_t );
bool f128_isSignalingNaN( float128_t );
static inline float128_t f128_complement_sign(float128_t a) {
a.v[1] ^= 1ULL << 63;
return a;
}
#endif
uint_fast32_t f128M_to_ui32( const float128_t *, uint_fast8_t, bool );
uint_fast64_t f128M_to_ui64( const float128_t *, uint_fast8_t, bool );
File renamed without changes.
File renamed without changes.
File renamed without changes.
+73
View File
@@ -0,0 +1,73 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool extF80_le( struct softfloat_state *state, extFloat80_t a, extFloat80_t b )
{
union { struct extFloat80M s; extFloat80_t f; } uA;
uint_fast16_t uiA64;
uint_fast64_t uiA0;
union { struct extFloat80M s; extFloat80_t f; } uB;
uint_fast16_t uiB64;
uint_fast64_t uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.s.signExp;
uiA0 = uA.s.signif;
uB.f = b;
uiB64 = uB.s.signExp;
uiB0 = uB.s.signif;
if ( isNaNExtF80UI( uiA64, uiA0 ) || isNaNExtF80UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signExtF80UI64( uiA64 );
signB = signExtF80UI64( uiB64 );
return
(signA != signB)
? signA || ! (((uiA64 | uiB64) & 0x7FFF) | uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_add( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
float128_t
(*magsFuncPtr)(
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
#endif
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_addMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_subMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_addMagsF128 : softfloat_subMagsF128;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
+199
View File
@@ -0,0 +1,199 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_div( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
int_fast32_t expB;
struct uint128 sigB;
bool signZ;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
struct uint128 rem;
uint_fast32_t recip32;
int ix;
uint_fast64_t q64;
uint_fast32_t q;
struct uint128 term;
uint_fast32_t qs[3];
uint_fast64_t sigZExtra;
struct uint128 sigZ, uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
signZ = signA ^ signB;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) goto propagateNaN;
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
goto invalid;
}
goto infinity;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
goto zero;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) {
if ( ! (expA | sigA.v64 | sigA.v0) ) goto invalid;
softfloat_raiseFlags( state, softfloat_flag_infinite );
goto infinity;
}
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
expZ = expA - expB + 0x3FFE;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB.v64 |= UINT64_C( 0x0001000000000000 );
rem = sigA;
if ( softfloat_lt128( sigA.v64, sigA.v0, sigB.v64, sigB.v0 ) ) {
--expZ;
rem = softfloat_add128( sigA.v64, sigA.v0, sigA.v64, sigA.v0 );
}
recip32 = softfloat_approxRecip32_1( sigB.v64>>17 );
ix = 3;
for (;;) {
q64 = (uint_fast64_t) (uint32_t) (rem.v64>>19) * recip32;
q = (q64 + 0x80000000)>>32;
--ix;
if ( ix < 0 ) break;
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
--q;
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
qs[ix] = q;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ((q + 1) & 7) < 2 ) {
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
--q;
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
} else if ( softfloat_le128( sigB.v64, sigB.v0, rem.v64, rem.v0 ) ) {
++q;
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
if ( rem.v64 | rem.v0 ) q |= 1;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
sigZExtra = (uint64_t) ((uint_fast64_t) q<<60);
term = softfloat_shortShiftLeft128( 0, qs[1], 54 );
sigZ =
softfloat_add128(
(uint_fast64_t) qs[2]<<19, ((uint_fast64_t) qs[0]<<25) + (q>>4),
term.v64, term.v0
);
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
infinity:
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
goto uiZ0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
zero:
uiZ.v64 = packToF128UI64( signZ, 0, 0 );
uiZ0:
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+73
View File
@@ -0,0 +1,73 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_eq( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
return
(uiA0 == uiB0)
&& ( (uiA64 == uiB64)
|| (! uiA0 && ! ((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF )))
);
}
+67
View File
@@ -0,0 +1,67 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_eq_signaling( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
return
(uiA0 == uiB0)
&& ( (uiA64 == uiB64)
|| (! uiA0 && ! ((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF )))
);
}
+51
View File
@@ -0,0 +1,51 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_isSignalingNaN( float128_t a )
{
union ui128_f128 uA;
uA.f = a;
return softfloat_isSigNaNF128UI( uA.ui.v64, uA.ui.v0 );
}
+72
View File
@@ -0,0 +1,72 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_le( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
|| ! (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_le_quiet( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
|| ! (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 == uiB64) && (uiA0 == uiB0))
|| (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+72
View File
@@ -0,0 +1,72 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
bool f128_lt( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
&& (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 != uiB64) || (uiA0 != uiB0))
&& (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
bool f128_lt_quiet( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signA, signB;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
if ( isNaNF128UI( uiA64, uiA0 ) || isNaNF128UI( uiB64, uiB0 ) ) {
if (
softfloat_isSigNaNF128UI( uiA64, uiA0 )
|| softfloat_isSigNaNF128UI( uiB64, uiB0 )
) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
}
return false;
}
signA = signF128UI64( uiA64 );
signB = signF128UI64( uiB64 );
return
(signA != signB)
? signA
&& (((uiA64 | uiB64) & UINT64_C( 0x7FFFFFFFFFFFFFFF ))
| uiA0 | uiB0)
: ((uiA64 != uiB64) || (uiA0 != uiB0))
&& (signA ^ softfloat_lt128( uiA64, uiA0, uiB64, uiB0 ));
}
+163
View File
@@ -0,0 +1,163 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_mul( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
int_fast32_t expB;
struct uint128 sigB;
bool signZ;
uint_fast64_t magBits;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
uint64_t sig256Z[4];
uint_fast64_t sigZExtra;
struct uint128 sigZ;
struct uint128_extra sig128Extra;
struct uint128 uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
signZ = signA ^ signB;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if (
(sigA.v64 | sigA.v0) || ((expB == 0x7FFF) && (sigB.v64 | sigB.v0))
) {
goto propagateNaN;
}
magBits = expB | sigB.v64 | sigB.v0;
goto infArg;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
magBits = expA | sigA.v64 | sigA.v0;
goto infArg;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) goto zero;
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
expZ = expA + expB - 0x4000;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB = softfloat_shortShiftLeft128( sigB.v64, sigB.v0, 16 );
softfloat_mul128To256M( sigA.v64, sigA.v0, sigB.v64, sigB.v0, sig256Z );
sigZExtra = sig256Z[indexWord( 4, 1 )] | (sig256Z[indexWord( 4, 0 )] != 0);
sigZ =
softfloat_add128(
sig256Z[indexWord( 4, 3 )], sig256Z[indexWord( 4, 2 )],
sigA.v64, sigA.v0
);
if ( UINT64_C( 0x0002000000000000 ) <= sigZ.v64 ) {
++expZ;
sig128Extra =
softfloat_shortShiftRightJam128Extra(
sigZ.v64, sigZ.v0, sigZExtra, 1 );
sigZ = sig128Extra.v;
sigZExtra = sig128Extra.extra;
}
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
infArg:
if ( ! magBits ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
goto uiZ;
}
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
goto uiZ0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
zero:
uiZ.v64 = packToF128UI64( signZ, 0, 0 );
uiZ0:
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+63
View File
@@ -0,0 +1,63 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_mulAdd( struct softfloat_state *state, float128_t a, float128_t b, float128_t c )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
union ui128_f128 uC;
uint_fast64_t uiC64, uiC0;
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
uC.f = c;
uiC64 = uC.ui.v64;
uiC0 = uC.ui.v0;
return softfloat_mulAddF128( uiA64, uiA0, uiB64, uiB0, uiC64, uiC0, 0 );
}
+190
View File
@@ -0,0 +1,190 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_rem( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
int_fast32_t expB;
struct uint128 sigB;
struct exp32_sig128 normExpSig;
struct uint128 rem;
int_fast32_t expDiff;
uint_fast32_t q, recip32;
uint_fast64_t q64;
struct uint128 term, altRem, meanRem;
bool signRem;
struct uint128 uiZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if (
(sigA.v64 | sigA.v0) || ((expB == 0x7FFF) && (sigB.v64 | sigB.v0))
) {
goto propagateNaN;
}
goto invalid;
}
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
return a;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expB ) {
if ( ! (sigB.v64 | sigB.v0) ) goto invalid;
normExpSig = softfloat_normSubnormalF128Sig( sigB.v64, sigB.v0 );
expB = normExpSig.exp;
sigB = normExpSig.sig;
}
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) return a;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sigB.v64 |= UINT64_C( 0x0001000000000000 );
rem = sigA;
expDiff = expA - expB;
if ( expDiff < 1 ) {
if ( expDiff < -1 ) return a;
if ( expDiff ) {
--expB;
sigB = softfloat_add128( sigB.v64, sigB.v0, sigB.v64, sigB.v0 );
q = 0;
} else {
q = softfloat_le128( sigB.v64, sigB.v0, rem.v64, rem.v0 );
if ( q ) {
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
}
} else {
recip32 = softfloat_approxRecip32_1( sigB.v64>>17 );
expDiff -= 30;
for (;;) {
q64 = (uint_fast64_t) (uint32_t) (rem.v64>>19) * recip32;
if ( expDiff < 0 ) break;
q = (q64 + 0x80000000)>>32;
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
rem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
}
expDiff -= 29;
}
/*--------------------------------------------------------------------
| (`expDiff' cannot be less than -29 here.)
*--------------------------------------------------------------------*/
q = (uint32_t) (q64>>32)>>(~expDiff & 31);
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, expDiff + 30 );
term = softfloat_mul128By32( sigB.v64, sigB.v0, q );
rem = softfloat_sub128( rem.v64, rem.v0, term.v64, term.v0 );
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
altRem = softfloat_add128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
goto selectRem;
}
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
do {
altRem = rem;
++q;
rem = softfloat_sub128( rem.v64, rem.v0, sigB.v64, sigB.v0 );
} while ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) );
selectRem:
meanRem = softfloat_add128( rem.v64, rem.v0, altRem.v64, altRem.v0 );
if (
(meanRem.v64 & UINT64_C( 0x8000000000000000 ))
|| (! (meanRem.v64 | meanRem.v0) && (q & 1))
) {
rem = altRem;
}
signRem = signA;
if ( rem.v64 & UINT64_C( 0x8000000000000000 ) ) {
signRem = ! signRem;
rem = softfloat_sub128( 0, 0, rem.v64, rem.v0 );
}
return softfloat_normRoundPackToF128( state, signRem, expB - 1, rem.v64, rem.v0 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
goto uiZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+201
View File
@@ -0,0 +1,201 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f128_sqrt( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
int_fast32_t expA;
struct uint128 sigA, uiZ;
struct exp32_sig128 normExpSig;
int_fast32_t expZ;
uint_fast32_t sig32A, recipSqrt32, sig32Z;
struct uint128 rem;
uint32_t qs[3];
uint_fast32_t q;
uint_fast64_t x64, sig64Z;
struct uint128 y, term;
uint_fast64_t sigZExtra;
struct uint128 sigZ;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) {
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, 0, 0 );
goto uiZ;
}
if ( ! signA ) return a;
goto invalid;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( signA ) {
if ( ! (expA | sigA.v64 | sigA.v0) ) return a;
goto invalid;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! expA ) {
if ( ! (sigA.v64 | sigA.v0) ) return a;
normExpSig = softfloat_normSubnormalF128Sig( sigA.v64, sigA.v0 );
expA = normExpSig.exp;
sigA = normExpSig.sig;
}
/*------------------------------------------------------------------------
| (`sig32Z' is guaranteed to be a lower bound on the square root of
| `sig32A', which makes `sig32Z' also a lower bound on the square root of
| `sigA'.)
*------------------------------------------------------------------------*/
expZ = ((expA - 0x3FFF)>>1) + 0x3FFE;
expA &= 1;
sigA.v64 |= UINT64_C( 0x0001000000000000 );
sig32A = sigA.v64>>17;
recipSqrt32 = softfloat_approxRecipSqrt32_1( expA, sig32A );
sig32Z = ((uint_fast64_t) sig32A * recipSqrt32)>>32;
if ( expA ) {
sig32Z >>= 1;
rem = softfloat_shortShiftLeft128( sigA.v64, sigA.v0, 12 );
} else {
rem = softfloat_shortShiftLeft128( sigA.v64, sigA.v0, 13 );
}
qs[2] = sig32Z;
rem.v64 -= (uint_fast64_t) sig32Z * sig32Z;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = ((uint32_t) (rem.v64>>2) * (uint_fast64_t) recipSqrt32)>>32;
x64 = (uint_fast64_t) sig32Z<<32;
sig64Z = x64 + ((uint_fast64_t) q<<3);
y = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
/*------------------------------------------------------------------------
| (Repeating this loop is a rare occurrence.)
*------------------------------------------------------------------------*/
for (;;) {
term = softfloat_mul64ByShifted32To128( x64 + sig64Z, q );
rem = softfloat_sub128( y.v64, y.v0, term.v64, term.v0 );
if ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) ) break;
--q;
sig64Z -= 1<<3;
}
qs[1] = q;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = ((rem.v64>>2) * recipSqrt32)>>32;
y = softfloat_shortShiftLeft128( rem.v64, rem.v0, 29 );
sig64Z <<= 1;
/*------------------------------------------------------------------------
| (Repeating this loop is a rare occurrence.)
*------------------------------------------------------------------------*/
for (;;) {
term = softfloat_shortShiftLeft128( 0, sig64Z, 32 );
term = softfloat_add128( term.v64, term.v0, 0, (uint_fast64_t) q<<6 );
term = softfloat_mul128By32( term.v64, term.v0, q );
rem = softfloat_sub128( y.v64, y.v0, term.v64, term.v0 );
if ( ! (rem.v64 & UINT64_C( 0x8000000000000000 )) ) break;
--q;
}
qs[0] = q;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
q = (((rem.v64>>2) * recipSqrt32)>>32) + 2;
sigZExtra = (uint64_t) ((uint_fast64_t) q<<59);
term = softfloat_shortShiftLeft128( 0, qs[1], 53 );
sigZ =
softfloat_add128(
(uint_fast64_t) qs[2]<<18, ((uint_fast64_t) qs[0]<<24) + (q>>5),
term.v64, term.v0
);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( (q & 0xF) <= 2 ) {
q &= ~3;
sigZExtra = (uint64_t) ((uint_fast64_t) q<<59);
y = softfloat_shortShiftLeft128( sigZ.v64, sigZ.v0, 6 );
y.v0 |= sigZExtra>>58;
term = softfloat_sub128( y.v64, y.v0, 0, q );
y = softfloat_mul64ByShifted32To128( term.v0, q );
term = softfloat_mul64ByShifted32To128( term.v64, q );
term = softfloat_add128( term.v64, term.v0, 0, y.v64 );
rem = softfloat_shortShiftLeft128( rem.v64, rem.v0, 20 );
term = softfloat_sub128( term.v64, term.v0, rem.v64, rem.v0 );
/*--------------------------------------------------------------------
| The concatenation of `term' and `y.v0' is now the negative remainder
| (3 words altogether).
*--------------------------------------------------------------------*/
if ( term.v64 & UINT64_C( 0x8000000000000000 ) ) {
sigZExtra |= 1;
} else {
if ( term.v64 | term.v0 | y.v0 ) {
if ( sigZExtra ) {
--sigZExtra;
} else {
sigZ = softfloat_sub128( sigZ.v64, sigZ.v0, 0, 1 );
sigZExtra = ~0;
}
}
}
}
return softfloat_roundPackToF128( state, 0, expZ, sigZ.v64, sigZ.v0, sigZExtra );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
invalid:
softfloat_raiseFlags( state, softfloat_flag_invalid );
uiZ.v64 = defaultNaNF128UI64;
uiZ.v0 = defaultNaNF128UI0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+78
View File
@@ -0,0 +1,78 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t f128_sub( struct softfloat_state *state, float128_t a, float128_t b )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool signA;
union ui128_f128 uB;
uint_fast64_t uiB64, uiB0;
bool signB;
#if ! defined INLINE_LEVEL || (INLINE_LEVEL < 2)
float128_t
(*magsFuncPtr)(
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
#endif
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
signA = signF128UI64( uiA64 );
uB.f = b;
uiB64 = uB.ui.v64;
uiB0 = uB.ui.v0;
signB = signF128UI64( uiB64 );
#if defined INLINE_LEVEL && (2 <= INLINE_LEVEL)
if ( signA == signB ) {
return softfloat_subMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
} else {
return softfloat_addMagsF128( state, uiA64, uiA0, uiB64, uiB0, signA );
}
#else
magsFuncPtr =
(signA == signB) ? softfloat_subMagsF128 : softfloat_addMagsF128;
return (*magsFuncPtr)( uiA64, uiA0, uiB64, uiB0, signA );
#endif
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float16_t f128_to_f16( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64;
struct commonNaN commonNaN;
uint_fast16_t uiZ, frac16;
union ui16_f16 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF16UI( &commonNaN );
} else {
uiZ = packToF16UI( sign, 0x1F, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac16 = softfloat_shortShiftRightJam64( frac64, 34 );
if ( ! (exp | frac16) ) {
uiZ = packToF16UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3FF1;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x40 ) exp = -0x40;
}
return softfloat_roundPackToF16( sign, exp, frac16 | 0x4000 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float32_t f128_to_f32( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64;
struct commonNaN commonNaN;
uint_fast32_t uiZ, frac32;
union ui32_f32 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF32UI( &commonNaN );
} else {
uiZ = packToF32UI( sign, 0xFF, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac32 = softfloat_shortShiftRightJam64( frac64, 18 );
if ( ! (exp | frac32) ) {
uiZ = packToF32UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3F81;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return softfloat_roundPackToF32( state, sign, exp, frac32 | 0x40000000 );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+100
View File
@@ -0,0 +1,100 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float64_t f128_to_f64( struct softfloat_state *state, float128_t a )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t frac64, frac0;
struct commonNaN commonNaN;
uint_fast64_t uiZ;
struct uint128 frac128;
union ui64_f64 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
frac64 = fracF128UI64( uiA64 );
frac0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0x7FFF ) {
if ( frac64 | frac0 ) {
softfloat_f128UIToCommonNaN( state, uiA64, uiA0, &commonNaN );
uiZ = softfloat_commonNaNToF64UI( &commonNaN );
} else {
uiZ = packToF64UI( sign, 0x7FF, 0 );
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
frac128 = softfloat_shortShiftLeft128( frac64, frac0, 14 );
frac64 = frac128.v64 | (frac128.v0 != 0);
if ( ! (exp | frac64) ) {
uiZ = packToF64UI( sign, 0, 0 );
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
exp -= 0x3C01;
if ( sizeof (int_fast16_t) < sizeof (int_fast32_t) ) {
if ( exp < -0x1000 ) exp = -0x1000;
}
return
softfloat_roundPackToF64(
state, sign, exp, frac64 | UINT64_C( 0x4000000000000000 ) );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+85
View File
@@ -0,0 +1,85 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
int_fast32_t f128_to_i32( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#if (i32_fromNaN != i32_fromPosOverflow) || (i32_fromNaN != i32_fromNegOverflow)
if ( (exp == 0x7FFF) && (sig64 | sig0) ) {
#if (i32_fromNaN == i32_fromPosOverflow)
sign = 0;
#elif (i32_fromNaN == i32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
return i32_fromNaN;
#endif
}
#endif
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sig64 |= (sig0 != 0);
shiftDist = 0x4023 - exp;
if ( 0 < shiftDist ) sig64 = softfloat_shiftRightJam64( sig64, shiftDist );
return softfloat_roundToI32( state, sign, sig64, roundingMode, exact );
}
+95
View File
@@ -0,0 +1,95 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
int_fast64_t f128_to_i64( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
struct uint128 sig128;
struct uint64_extra sigExtra;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
shiftDist = 0x402F - exp;
if ( shiftDist <= 0 ) {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist < -15 ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig64 | sig0) ? i64_fromNaN
: sign ? i64_fromNegOverflow : i64_fromPosOverflow;
}
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
sig64 |= UINT64_C( 0x0001000000000000 );
if ( shiftDist ) {
sig128 = softfloat_shortShiftLeft128( sig64, sig0, -shiftDist );
sig64 = sig128.v64;
sig0 = sig128.v0;
}
} else {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sigExtra = softfloat_shiftRightJam64Extra( sig64, sig0, shiftDist );
sig64 = sigExtra.v;
sig0 = sigExtra.extra;
}
return softfloat_roundToI64( state, sign, sig64, sig0, roundingMode, exact );
}
+86
View File
@@ -0,0 +1,86 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
uint_fast32_t
f128_to_ui32( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64;
int_fast32_t shiftDist;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 ) | (uiA0 != 0);
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
#if (ui32_fromNaN != ui32_fromPosOverflow) || (ui32_fromNaN != ui32_fromNegOverflow)
if ( (exp == 0x7FFF) && sig64 ) {
#if (ui32_fromNaN == ui32_fromPosOverflow)
sign = 0;
#elif (ui32_fromNaN == ui32_fromNegOverflow)
sign = 1;
#else
softfloat_raiseFlags( softfloat_flag_invalid );
return ui32_fromNaN;
#endif
}
#endif
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
shiftDist = 0x4023 - exp;
if ( 0 < shiftDist ) {
sig64 = softfloat_shiftRightJam64( sig64, shiftDist );
}
return softfloat_roundToUI32( sign, sig64, roundingMode, exact );
}
+96
View File
@@ -0,0 +1,96 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016, 2017 The Regents of the
University of California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
uint_fast64_t
f128_to_ui64( struct softfloat_state *state, float128_t a, uint_fast8_t roundingMode, bool exact )
{
union ui128_f128 uA;
uint_fast64_t uiA64, uiA0;
bool sign;
int_fast32_t exp;
uint_fast64_t sig64, sig0;
int_fast32_t shiftDist;
struct uint128 sig128;
struct uint64_extra sigExtra;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA64 = uA.ui.v64;
uiA0 = uA.ui.v0;
sign = signF128UI64( uiA64 );
exp = expF128UI64( uiA64 );
sig64 = fracF128UI64( uiA64 );
sig0 = uiA0;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
shiftDist = 0x402F - exp;
if ( shiftDist <= 0 ) {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( shiftDist < -15 ) {
softfloat_raiseFlags( state, softfloat_flag_invalid );
return
(exp == 0x7FFF) && (sig64 | sig0) ? ui64_fromNaN
: sign ? ui64_fromNegOverflow : ui64_fromPosOverflow;
}
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
sig64 |= UINT64_C( 0x0001000000000000 );
if ( shiftDist ) {
sig128 = softfloat_shortShiftLeft128( sig64, sig0, -shiftDist );
sig64 = sig128.v64;
sig0 = sig128.v0;
}
} else {
/*--------------------------------------------------------------------
*--------------------------------------------------------------------*/
if ( exp ) sig64 |= UINT64_C( 0x0001000000000000 );
sigExtra = softfloat_shiftRightJam64Extra( sig64, sig0, shiftDist );
sig64 = sigExtra.v;
sig0 = sigExtra.extra;
}
return softfloat_roundToUI64( state, sign, sig64, sig0, roundingMode, exact );
}
+96
View File
@@ -0,0 +1,96 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
#include "softfloat.h"
float128_t f32_to_f128( struct softfloat_state *state, float32_t a )
{
union ui32_f32 uA;
uint_fast32_t uiA;
bool sign;
int_fast16_t exp;
uint_fast32_t frac;
struct commonNaN commonNaN;
struct uint128 uiZ;
struct exp16_sig32 normExpSig;
union ui128_f128 uZ;
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uA.f = a;
uiA = uA.ui;
sign = signF32UI( uiA );
exp = expF32UI( uiA );
frac = fracF32UI( uiA );
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( exp == 0xFF ) {
if ( frac ) {
softfloat_f32UIToCommonNaN( state, uiA, &commonNaN );
uiZ = softfloat_commonNaNToF128UI( &commonNaN );
} else {
uiZ.v64 = packToF128UI64( sign, 0x7FFF, 0 );
uiZ.v0 = 0;
}
goto uiZ;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
if ( ! exp ) {
if ( ! frac ) {
uiZ.v64 = packToF128UI64( sign, 0, 0 );
uiZ.v0 = 0;
goto uiZ;
}
normExpSig = softfloat_normSubnormalF32Sig( frac );
exp = normExpSig.exp - 1;
frac = normExpSig.sig;
}
/*------------------------------------------------------------------------
*------------------------------------------------------------------------*/
uiZ.v64 = packToF128UI64( sign, exp + 0x3F80, (uint_fast64_t) frac<<25 );
uiZ.v0 = 0;
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
+64
View File
@@ -0,0 +1,64 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014, 2015, 2016 The Regents of the University of
California. All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "softfloat.h"
float128_t i32_to_f128( int32_t a )
{
uint_fast64_t uiZ64;
bool sign;
uint_fast32_t absA;
int_fast8_t shiftDist;
union ui128_f128 uZ;
uiZ64 = 0;
if ( a ) {
sign = (a < 0);
absA = sign ? -(uint_fast32_t) a : (uint_fast32_t) a;
shiftDist = softfloat_countLeadingZeros32( absA ) + 17;
uiZ64 =
packToF128UI64(
sign, 0x402E - shiftDist, (uint_fast64_t) absA<<shiftDist );
}
uZ.ui.v64 = uiZ64;
uZ.ui.v0 = 0;
return uZ.f;
}
@@ -196,16 +196,20 @@ struct exp32_sig128
float128_t
softfloat_roundPackToF128(
struct softfloat_state *,
bool, int_fast32_t, uint_fast64_t, uint_fast64_t, uint_fast64_t );
float128_t
softfloat_normRoundPackToF128(
struct softfloat_state *,
bool, int_fast32_t, uint_fast64_t, uint_fast64_t );
float128_t
softfloat_addMagsF128(
struct softfloat_state *,
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
float128_t
softfloat_subMagsF128(
struct softfloat_state *,
uint_fast64_t, uint_fast64_t, uint_fast64_t, uint_fast64_t, bool );
float128_t
softfloat_mulAddF128(
File renamed without changes.
File renamed without changes.
+155
View File
@@ -0,0 +1,155 @@
/*============================================================================
This C source file is part of the SoftFloat IEEE Floating-Point Arithmetic
Package, Release 3e, by John R. Hauser.
Copyright 2011, 2012, 2013, 2014 The Regents of the University of California.
All rights reserved.
Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:
1. Redistributions of source code must retain the above copyright notice,
this list of conditions, and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
this list of conditions, and the following disclaimer in the documentation
and/or other materials provided with the distribution.
3. Neither the name of the University nor the names of its contributors may
be used to endorse or promote products derived from this software without
specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE REGENTS AND CONTRIBUTORS "AS IS", AND ANY
EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE, ARE
DISCLAIMED. IN NO EVENT SHALL THE REGENTS OR CONTRIBUTORS BE LIABLE FOR ANY
DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES
(INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES;
LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND
ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT
(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS
SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
=============================================================================*/
#include <stdbool.h>
#include <stdint.h>
#include "platform.h"
#include "internals.h"
#include "specialize.h"
float128_t
softfloat_addMagsF128(
struct softfloat_state *state,
uint_fast64_t uiA64,
uint_fast64_t uiA0,
uint_fast64_t uiB64,
uint_fast64_t uiB0,
bool signZ
)
{
int_fast32_t expA;
struct uint128 sigA;
int_fast32_t expB;
struct uint128 sigB;
int_fast32_t expDiff;
struct uint128 uiZ, sigZ;
int_fast32_t expZ;
uint_fast64_t sigZExtra;
struct uint128_extra sig128Extra;
union ui128_f128 uZ;
expA = expF128UI64( uiA64 );
sigA.v64 = fracF128UI64( uiA64 );
sigA.v0 = uiA0;
expB = expF128UI64( uiB64 );
sigB.v64 = fracF128UI64( uiB64 );
sigB.v0 = uiB0;
expDiff = expA - expB;
if ( ! expDiff ) {
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 | sigB.v64 | sigB.v0 ) goto propagateNaN;
uiZ.v64 = uiA64;
uiZ.v0 = uiA0;
goto uiZ;
}
sigZ = softfloat_add128( sigA.v64, sigA.v0, sigB.v64, sigB.v0 );
if ( ! expA ) {
uiZ.v64 = packToF128UI64( signZ, 0, sigZ.v64 );
uiZ.v0 = sigZ.v0;
goto uiZ;
}
expZ = expA;
sigZ.v64 |= UINT64_C( 0x0002000000000000 );
sigZExtra = 0;
goto shiftRight1;
}
if ( expDiff < 0 ) {
if ( expB == 0x7FFF ) {
if ( sigB.v64 | sigB.v0 ) goto propagateNaN;
uiZ.v64 = packToF128UI64( signZ, 0x7FFF, 0 );
uiZ.v0 = 0;
goto uiZ;
}
expZ = expB;
if ( expA ) {
sigA.v64 |= UINT64_C( 0x0001000000000000 );
} else {
++expDiff;
sigZExtra = 0;
if ( ! expDiff ) goto newlyAligned;
}
sig128Extra =
softfloat_shiftRightJam128Extra( sigA.v64, sigA.v0, 0, -expDiff );
sigA = sig128Extra.v;
sigZExtra = sig128Extra.extra;
} else {
if ( expA == 0x7FFF ) {
if ( sigA.v64 | sigA.v0 ) goto propagateNaN;
uiZ.v64 = uiA64;
uiZ.v0 = uiA0;
goto uiZ;
}
expZ = expA;
if ( expB ) {
sigB.v64 |= UINT64_C( 0x0001000000000000 );
} else {
--expDiff;
sigZExtra = 0;
if ( ! expDiff ) goto newlyAligned;
}
sig128Extra =
softfloat_shiftRightJam128Extra( sigB.v64, sigB.v0, 0, expDiff );
sigB = sig128Extra.v;
sigZExtra = sig128Extra.extra;
}
newlyAligned:
sigZ =
softfloat_add128(
sigA.v64 | UINT64_C( 0x0001000000000000 ),
sigA.v0,
sigB.v64,
sigB.v0
);
--expZ;
if ( sigZ.v64 < UINT64_C( 0x0002000000000000 ) ) goto roundAndPack;
++expZ;
shiftRight1:
sig128Extra =
softfloat_shortShiftRightJam128Extra(
sigZ.v64, sigZ.v0, sigZExtra, 1 );
sigZ = sig128Extra.v;
sigZExtra = sig128Extra.extra;
roundAndPack:
return
softfloat_roundPackToF128( state, signZ, expZ, sigZ.v64, sigZ.v0, sigZExtra );
propagateNaN:
uiZ = softfloat_propagateNaNF128UI( state, uiA64, uiA0, uiB64, uiB0 );
uiZ:
uZ.ui = uiZ;
return uZ.f;
}
Loaded 100 of 388 files, more files were not shown because too many files have changed in this diff. Show more